Profiling & the Roofline
Nsight is your microscope; the roofline is your map.
The two Nsight tools answer different questions, and people conflate them constantly. Systems shows the timeline, where the GPU sits idle between kernels — the silent killer in inference. Compute shows the inside of one kernel; learn to read the SOL section and the memory charts, because that screenshot is the shared language of every performance discussion. The roofline drill is the most reliable interview filter in the field: given a workload's arithmetic intensity, say which side of the ridge it lands on and what that implies. Prefill and decode land on opposite sides. That's the entire field in one picture.
CUDA graphs close the loop on the overhead bottleneck from the very first quest. A decode step launches hundreds of tiny kernels, launch overhead swamps them, so engines capture the step once and replay it. And do the memory-snapshot task for real; the caching allocator and fragmentation OOMs are genuine on-call work, not curriculum filler.
- PyTorch Developer PodcastEdward Z. Yangpodcast↗
- Modal GPU GlossaryModalreference↗
Find the gaps between kernels — idle GPU time is the silent killer. Do it on a real model on your fleet.
SOL section, memory charts, occupancy. Learn to read what everyone screenshots.
Coalescing, occupancy, ILP — the optimization playbook, in one lecture.
watch+50 XPRidge points, arithmetic intensity of prefill vs decode, why batching works. 75% to pass.
drill+80 XPauto-verifiedLaunch-overhead elimination — this is why vLLM/TRT-LLM capture the decode step.
Record a snapshot of a real inference run, open it in memory_viz, and explain one allocation spike. Also learn the caching allocator and expandable_segments — fragmentation OOMs are real on-call work.