InferQuest
← Quest map
Inference Phase 1 · Engine Core

Batching & Scheduling

Continuous batching is why serving companies exist.

0/4 · 260 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

Read these in order; each one answers a problem the previous one created. Orca schedules at iteration granularity instead of request granularity, so a finished sequence leaves the batch immediately instead of holding its slot (the Anyscale explainer has the diagrams the paper lacks). PagedAttention applies the OS virtual-memory playbook to the KV cache: fixed-size blocks and a block table kill the fragmentation that Orca-style batching creates, and the same indirection later buys prefix caching almost for free. Sarathi chunks long prefills so one fat prompt can't stall everyone else's decode, and is now default behavior in vLLM.

The thread running through all three is the tension between TTFT and ITL. Every scheduling choice trades time-to-first-token for someone against inter-token latency for someone else. The drill leans on that tension hard.

From the library
Tasks
  • Pair with Anyscale's continuous-batching explainer (anyscale.com/blog/continuous-batching-llm-inference).

    paper+60 XPresource ↗
  • The founding document of modern serving: KV fragmentation → paged blocks. Non-negotiable.

    paper+70 XPresource ↗
  • Stall-free scheduling — now default in vLLM V1.

    paper+50 XPresource ↗
  • TTFT vs ITL, continuous batching, chunked prefill, prefix caching. 75% to pass.

    drill+80 XPauto-verified