InferQuest
← Quest map
Foundations Phase 3 · The Metal

Matmul Mastery

The rite of passage: chase cuBLAS.

0/3 · 430 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

Boehm's worklog teaches the method, not just the kernel: change one thing, profile, explain the delta, repeat. Each rung has a name you'll reuse forever (coalescing, shared-memory tiling, vectorized loads, warp tiling), and the biggest single jump is coalescing. That's why memory-access patterns, not FLOPs, are the first thing to check in any slow kernel. The flash-attention quest asks you to produce a worklog of your own in this genre, and worklogs like it have opened real doors.

The 40%-of-cuBLAS bar is set where it is on purpose: reachable with tiling done right, unreachable by accident. Afterward, don't skip the reduction, scan, and histogram patterns. Reductions are the skeleton of softmax, layernorm, and RMSNorm, which is most of what an inference kernel engineer actually ships.

From the library
Tasks
  • Naive → coalesced → tiled → vectorized → warp-tiled, with measured speedups. The de facto kernel-interview prep (its author got hired to Anthropic's perf team).

    read+100 XPresource ↗
  • Shared-memory tiling in CUDA C++ or Triton. Graded on correctness plus measured TFLOPS as a fraction of cuBLAS on 4096³ — GPU-agnostic by construction.

    kernel+250 XPauto-verified
  • Chapters 9–11 (4th ed. numbering) + GPU MODE lecture 9. Reductions underlie softmax and every norm.

    read+80 XP