InferQuest
← Quest map
Foundations Phase 6 · Many GPUs

Parallelism

TP, PP, EP — and the math of when each wins.

0/8 · 860 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

“How To Scale Your Model” is the best thing written on this subject, and the inference chapter is why this quest exists. It's TPU-flavored, but the communication math transfers to NVLink unchanged. Megatron is the concrete instantiation: column-parallel then row-parallel means one all-reduce per attention block and one per MLP, and knowing exactly where those land is what the drill (and interviews) test. TP versus PP is an interconnect-bandwidth question with a numeric answer, not a matter of taste.

The ring all-reduce build looks like a toy and isn't. Reduce-scatter plus all-gather over point-to-point sends is the algorithm inside NCCL, and implementing it once is how bandwidth-optimal collectives stop being folklore. The scaling measurement completes the argument: TP=2 will not give you 2×, and working out why (per-layer all-reduce cost against your actual interconnect) is the lesson rather than the disappointment.

Tasks
  • DeepMind's systems-math text. TPU-flavored, universally applicable. The best thing written on this.

    read+120 XPresource ↗
  • Column/row-parallel splits and where the all-reduces land. Pair with GPU MODE lecture 17 (NCCL).

    paper+60 XPresource ↗
  • Ultra-Scale Playbook appendices A0 and A3 — the analytical grounding under everything in this phase.

    read+60 XPresource ↗
  • The fabric under the collectives: InfiniBand vs RoCE as a real design decision, rail-optimized topologies, and where GPUDirect RDMA fits. Familiarity-level on purpose — but it's the vocabulary training-infra job posts list by name.

    read+40 XPresource ↗
  • Reduce-scatter + all-gather over point-to-point send/recv across 4 processes (CPU, gloo — no fleet needed). Collectives are monkeypatched to raise, so the ring is yours. Three peer courses grade exactly this before letting students near NCCL.

    build+200 XPauto-verified
  • Shard your Phase 1 GPT's attention + MLP column/row-parallel across 2+ processes, placing the f/g all-reduces yourself (torch.distributed, CPU is fine). Logits must match single-process.

    build+150 XP
  • It will not be 2×. Explaining exactly why (NVLink vs PCIe, all-reduce cost per layer) is the lesson.

    bench+150 XP
  • What gets communicated when, TP vs PP vs EP tradeoffs, interconnect math. 75% to pass.

    drill+80 XPauto-verified