The Compiler Stack
torch.compile is load-bearing in vLLM V1 — stop treating it as magic.
This quest exists because vLLM V1 made torch.compile part of the engine. The model graph is compiled piecewise, split at the attention ops, and the pieces get captured into CUDA graphs, so “the compiler is magic” stops being a tenable position for anyone who wants to work on the engine. Trigger a recompile on purpose and watch the logs. Unexpected recompiles from dynamic shapes are the production footgun: a latency spike with no visible cause until you know where to look.
The Inductor task is the best demystification trick in the stack. TORCH_LOGS=output_code shows you the Triton the compiler writes, with its fusion and tiling decisions laid bare. Reading generated kernels right before the phase that asks you to write your own is intentional sequencing; you arrive with a working example of what good pointer math looks like.
- PyTorch Developer PodcastEdward Z. Yangpodcast↗
- Modal GPU GlossaryModalreference↗
Compile a decode loop with a static KV cache and measure the speedup; then trigger a recompile on purpose (dynamic shape) and watch it in the logs.
TORCH_LOGS=output_code on a small fused op. Find the fusion decisions, the tiling, the pointer math — compare with the kernels you'll write in Phase 4.
build+100 XPPiecewise compilation split at attention ops, captured into CUDA graphs — the concrete architecture your Phase 6 deployment runs on.