InferQuest
← Quest map
Inference Phase 4 · Datacenter Scale

Disaggregation & MoE at Scale

The 2025-26 production frontier.

0/7 · 620 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

DistServe supplies the frame the industry adopted. Colocated prefill and decode contaminate each other's latency, so split them and optimize each against its own SLO; “goodput per GPU” is the phrase that survived. Mooncake is the proof at production scale, with the further idea that the KV cache is the center of the architecture, tiered across DRAM and SSD. The LMSYS expert-parallelism writeup is the single best account of modern MoE serving, and it's what makes the TP-vs-EP build meaningful: implement both, find the crossover, and the all-to-all pattern from the MoE papers stops being abstract.

The infrastructure survey (Dynamo, NIXL, LMCache, llm-d) is here because KV movement grew into a subsystem with job requisitions attached. Be able to say what each piece does in one sentence. Same logic for the Kubernetes task: KV-aware load balancing is a GA gateway primitive now, and the K8s layer has stopped being someone else's problem.

From the library
Tasks
  • Prefill/decode disaggregation and per-request goodput — the framing the industry adopted.

    paper+70 XPresource ↗
  • Cache tiering across DRAM/SSD, SLO-aware scheduling, real production numbers.

    paper+70 XPresource ↗
  • KV movement is a subsystem with job reqs attached now. Know what each piece does and why it exists.

    read+80 XPresource ↗
  • DeepEP, EPLB, two-batch overlap — the single best writeup of modern MoE serving.

    read+80 XPresource ↗
  • UCSD CSE 234-style: a sharded-linear expert layer both ways — TP (shard every expert) vs EP (all-to-all token routing) — and measure where each wins. CPU processes are fine; the communication pattern is the lesson.

    build+180 XP
  • The fourth parallelism axis — vLLM's decode context parallelism is production-real for long-context agents.

    paper+60 XPresource ↗
  • KV-aware load balancing is a K8s primitive now (GA). DRA replaced device-plugin-only thinking in 1.34.

    read+80 XPresource ↗