The Anatomy post is the closest thing to a textbook chapter on a production engine, and the task's V1 warning matters: the internet is thick with V0-era posts describing a scheduler that no longer exists, and repeating them in an interview is a tell. The request trace is the real work of this quest. Writing down every file and class one request touches is the difference between “has used vLLM” and “has read it,” and it's exactly the preparation for landing PRs in the Arena phase.
When you tune your deployment, one knob at a time with a results table isn't pedantry. Several knobs interact (chunked-prefill budget against prefix-cache hit rate, quantized KV against batch capacity), and the table of what each did on your hardware is both how you'll actually understand them and a solid interview artifact in its own right.
- Inference EngineeringPhilip Kielybook↗
The closest thing to a textbook chapter on engine internals. Study the V1 architecture — V0 posts describe a dead engine.
Arrival → scheduler → block manager → model runner → sampler → stream. Write down every file and class it touches.
read+150 XPmax-num-batched-tokens, prefix caching, chunked prefill, quantized weights + FP8 KV. One knob at a time, keep a results table.
Expose it (with an API key) and let InferQuest probe it live — same suite your toy engine passed, now against production infrastructure.
build+200 XPauto-verified