The Triton Track
Python-first kernels are how 2026 writes them.
Leaning on the official tutorials isn't laziness. They're maintained by the compiler's own authors and sequenced exactly right: vector add teaches the programming model, fused softmax teaches the fusion win, matmul teaches block-level tiling, and fused attention previews the next quest. Triton's bargain is that you think in blocks and pointer arithmetic while the compiler handles the warp-level details. The stride math from Under the Tensor comes back here as code you literally write.
The grader's two failure modes come from real life: odd shapes, where masking has to be right when the row doesn't divide the block, and large logits, where the max-subtraction trick from your from-scratch attention never stops mattering. Landing within 1.25× of torch.softmax proves you actually fused the passes rather than transliterated numpy.
- PyTorch Developer PodcastEdward Z. Yangpodcast↗
Vector add → fused softmax → matmul → layernorm → fused attention. This IS the curriculum.
Row-wise softmax in Triton: correct on odd shapes and large logits, and measured within 25% of torch.softmax on (4096, 4096).
kernel+200 XPauto-verified