InferQuest
← Quest map
Foundations Phase 2 · Foundations

The Architecture Zoo

MHA → MQA → GQA → MLA, RoPE, and MoE — through the serving lens.

0/5 · 380 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

Read every architecture here through one lens: what it does to KV bytes per token. MQA and GQA exist because the cache, not the weights, is what caps batch size, so actually do the arithmetic for each variant. The arithmetic is the content, not a chore attached to it. Slow down on MLA: latent compression changes both the cache math and what an attention kernel has to do, and DeepSeek-shaped models have made it table stakes.

EleutherAI's RoPE explainer is assigned alongside RoFormer because the paper's notation is heavier than the idea deserves. YaRN is here because context extension is a serving-time concern, not a training curiosity. Read the MoE material as systems papers: routing is an all-to-all communication problem, and it comes back hard in the disaggregation quest. The build task is the checkpoint. Verifying logits against a real Llama-family model catches the head-interleaving and rotation details that prose lets you gloss over.

From the library
Tasks
  • Shazeer 2019 (one write-head) + GQA. Work out the KV-cache arithmetic for each — that IS the point of these designs.

    paper+60 XPresource ↗
  • Latent KV compression — mandatory 2026 knowledge. It changes cache math AND kernel design.

    paper+70 XPresource ↗
  • Then skim YaRN (arxiv.org/abs/2309.00071) — long-context serving depends on it.

    paper+50 XPresource ↗
  • Modify your implementation, verify logits against a HF reference model that uses them (e.g. a small Llama).

    build+120 XP
  • Routing, capacity factors, fine-grained + shared experts. Then skim the DeepSeek-V3 report (arxiv.org/abs/2412.19437) — it doubles as a systems paper.

    paper+80 XPresource ↗