Disaggregation & MoE at Scale
The 2025-26 production frontier.
DistServe supplies the frame the industry adopted. Colocated prefill and decode contaminate each other's latency, so split them and optimize each against its own SLO; “goodput per GPU” is the phrase that survived. Mooncake is the proof at production scale, with the further idea that the KV cache is the center of the architecture, tiered across DRAM and SSD. The LMSYS expert-parallelism writeup is the single best account of modern MoE serving, and it's what makes the TP-vs-EP build meaningful: implement both, find the crossover, and the all-to-all pattern from the MoE papers stops being abstract.
The infrastructure survey (Dynamo, NIXL, LMCache, llm-d) is here because KV movement grew into a subsystem with job requisitions attached. Be able to say what each piece does in one sentence. Same logic for the Kubernetes task: KV-aware load balancing is a GA gateway primitive now, and the K8s layer has stopped being someone else's problem.
- Inference EngineeringPhilip Kielybook↗
Prefill/decode disaggregation and per-request goodput — the framing the industry adopted.
Cache tiering across DRAM/SSD, SLO-aware scheduling, real production numbers.
KV movement is a subsystem with job reqs attached now. Know what each piece does and why it exists.
DeepEP, EPLB, two-batch overlap — the single best writeup of modern MoE serving.
UCSD CSE 234-style: a sharded-linear expert layer both ways — TP (shard every expert) vs EP (all-to-all token routing) — and measure where each wins. CPU processes are fine; the communication pattern is the lesson.
build+180 XPThe fourth parallelism axis — vLLM's decode context parallelism is production-real for long-context agents.
KV-aware load balancing is a K8s primitive now (GA). DRA replaced device-plugin-only thinking in 1.34.