The KV Cache
The single most important idea in LLM serving.
kipply's post is the highest-leverage read on the whole map. Half a dozen small formulas — KV bytes per token, FLOPs per token, bandwidth-implied latency floors — turn “inference is slow” from vibes into arithmetic. The task says do the math by hand because the derived numbers are exactly what the sizing drill asks, and what a surprising share of real interviews ask: how many concurrent sequences fit next to the weights on an 80 GB card, and why decode saturates bandwidth long before compute.
The cached-decoder build makes the core asymmetry concrete: prefill is compute-bound and parallel, decode is bandwidth-bound and one token at a time. Every scheduling paper in the next quest exists because of that asymmetry.
- Inference EngineeringPhilip Kielybook↗
KV bytes/token, FLOPs, latency floors. Do it once by hand for a model you serve and you'll never be fooled by a benchmark again.
Implement a cached decoder for the harness's reference GPT: prefill + O(1)-per-token decode. Graded on exact logit match AND ≥2× measured speedup over naive recompute.
build+250 XPauto-verifiedThe interview staple: cache-per-token math, GQA arithmetic, quantized-KV capacity. 75% to pass.
drill+80 XPauto-verified