InferQuest
← Quest map
Foundations Phase 3 · The Metal

CUDA Foundations

PMPP + GPU MODE: the canonical on-ramp.

0/5 · 450 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

PMPP is the textbook, and chapters 1–6 are the ones that matter most: the execution model (grids, blocks, warps) and the memory hierarchy are the two mental models every kernel you write will lean on. GPU MODE rides alongside as the closest thing this niche has to a guild — practitioner lectures that recap PMPP with war stories, and a Discord that posts jobs. Don't skip the Stephen Jones talks. Nobody explains better why GPUs are shaped the way they are (latency hiding through massive oversubscription), and that “why” is what survives after the API details fade.

Watch for the occupancy trap early. Occupancy is a means to an end, enough warps in flight to hide latency, not a score to maximize, and chasing it produces slow kernels with beautiful occupancy numbers. And do the humble vector-add task; the load_inline workflow it teaches is how you'll iterate on every kernel through the Kernel Engineering phase.

From the library
Tasks
  • 5th edition (2026) if buying fresh; 4th is fine. These six chapters are the core.

    read+120 XPresource ↗
  • The de facto open curriculum for this job. Join the Discord while you're there — it posts jobs.

    watch+80 XPresource ↗
  • Thread-indexing drills, zero setup on Colab.

    kernel+100 XPresource ↗
  • Stephen Jones' talks — the best conceptual account of why GPUs are shaped this way.

    watch+50 XPresource ↗
  • CUDA C++ via torch.utils.cpp_extension.load_inline, per GPU MODE lecture 1.

    kernel+100 XP