InferQuest
← Quest map
Inference Phase 1 · Engine Core

The KV Cache

The single most important idea in LLM serving.

0/3 · 410 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

kipply's post is the highest-leverage read on the whole map. Half a dozen small formulas — KV bytes per token, FLOPs per token, bandwidth-implied latency floors — turn “inference is slow” from vibes into arithmetic. The task says do the math by hand because the derived numbers are exactly what the sizing drill asks, and what a surprising share of real interviews ask: how many concurrent sequences fit next to the weights on an 80 GB card, and why decode saturates bandwidth long before compute.

The cached-decoder build makes the core asymmetry concrete: prefill is compute-bound and parallel, decode is bandwidth-bound and one token at a time. Every scheduling paper in the next quest exists because of that asymmetry.

From the library
Tasks
  • KV bytes/token, FLOPs, latency floors. Do it once by hand for a model you serve and you'll never be fooled by a benchmark again.

    read+80 XPresource ↗
  • Implement a cached decoder for the harness's reference GPT: prefill + O(1)-per-token decode. Graded on exact logit match AND ≥2× measured speedup over naive recompute.

    build+250 XPauto-verified
  • The interview staple: cache-per-token math, GQA arithmetic, quantized-KV capacity. 75% to pass.

    drill+80 XPauto-verified