Batching & Scheduling
Continuous batching is why serving companies exist.
Read these in order; each one answers a problem the previous one created. Orca schedules at iteration granularity instead of request granularity, so a finished sequence leaves the batch immediately instead of holding its slot (the Anyscale explainer has the diagrams the paper lacks). PagedAttention applies the OS virtual-memory playbook to the KV cache: fixed-size blocks and a block table kill the fragmentation that Orca-style batching creates, and the same indirection later buys prefix caching almost for free. Sarathi chunks long prefills so one fat prompt can't stall everyone else's decode, and is now default behavior in vLLM.
The thread running through all three is the tension between TTFT and ITL. Every scheduling choice trades time-to-first-token for someone against inter-token latency for someone else. The drill leans on that tension hard.
- Inference EngineeringPhilip Kielybook↗
Pair with Anyscale's continuous-batching explainer (anyscale.com/blog/continuous-batching-llm-inference).
The founding document of modern serving: KV fragmentation → paged blocks. Non-negotiable.
Stall-free scheduling — now default in vLLM V1.
TTFT vs ITL, continuous batching, chunked prefill, prefix caching. 75% to pass.
drill+80 XPauto-verified