Build Your Own Engine
The capstone: a real serving engine, probed live by InferQuest.
nano-vllm is the blueprint because it's the rare codebase small enough to read end to end and honest enough to contain the real ideas: prefix caching, tensor parallelism, CUDA graphs, in about 1.2k lines. Read it before writing a line of yours. Then resist copying it — the point of the build is discovering why each piece exists by needing it.
The details the probe grades were picked because that's where real engines leak: SSE framing with [DONE], usage accounting, max_tokens cutoffs, error shapes. Give per-request cancellation particular respect. A client disconnecting mid-stream has to free its KV blocks and leave the batch cleanly, and “streaming token generator with cancellation” shows up as a literal coding exercise at serving companies. This capstone also matters later: the flash-attention quest asks you to swap your own kernel into this engine and keep the probe green.
- Inference EngineeringPhilip Kielybook↗
~1.2k lines: prefix caching, TP, CUDA graphs. The nanoGPT of inference engines — your blueprint.
Requests join/leave the batch at token boundaries; block-allocated KV; SSE streaming out with per-request cancellation and timeouts (the 'streaming token generator with cancellation' is a recurring Fireworks/Together coding exercise). Any small model.
build+300 XPOutlines-style FSM masking, or integrate xgrammar. Read the vLLM structured-decoding intro first.
Expose your engine as an OpenAI-compatible API and point InferQuest's prober at it: /v1/models, chat completions, usage accounting, max_tokens cutoff, SSE streaming framing with [DONE], and error shapes — all live-tested.
build+400 XPauto-verified