Matmul Mastery
The rite of passage: chase cuBLAS.
Boehm's worklog teaches the method, not just the kernel: change one thing, profile, explain the delta, repeat. Each rung has a name you'll reuse forever (coalescing, shared-memory tiling, vectorized loads, warp tiling), and the biggest single jump is coalescing. That's why memory-access patterns, not FLOPs, are the first thing to check in any slow kernel. The flash-attention quest asks you to produce a worklog of your own in this genre, and worklogs like it have opened real doors.
The 40%-of-cuBLAS bar is set where it is on purpose: reachable with tiling done right, unreachable by accident. Afterward, don't skip the reduction, scan, and histogram patterns. Reductions are the skeleton of softmax, layernorm, and RMSNorm, which is most of what an inference kernel engineer actually ships.
- PyTorch Developer PodcastEdward Z. Yangpodcast↗
- Modal GPU GlossaryModalreference↗
Naive → coalesced → tiled → vectorized → warp-tiled, with measured speedups. The de facto kernel-interview prep (its author got hired to Anthropic's perf team).
Shared-memory tiling in CUDA C++ or Triton. Graded on correctness plus measured TFLOPS as a fraction of cuBLAS on 4096³ — GPU-agnostic by construction.
kernel+250 XPauto-verifiedChapters 9–11 (4th ed. numbering) + GPU MODE lecture 9. Reductions underlie softmax and every norm.
read+80 XP