InferQuest
← Quest map
Inference Phase 1 · Engine Core

Build Your Own Engine

The capstone: a real serving engine, probed live by InferQuest.

0/4 · 900 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

nano-vllm is the blueprint because it's the rare codebase small enough to read end to end and honest enough to contain the real ideas: prefix caching, tensor parallelism, CUDA graphs, in about 1.2k lines. Read it before writing a line of yours. Then resist copying it — the point of the build is discovering why each piece exists by needing it.

The details the probe grades were picked because that's where real engines leak: SSE framing with [DONE], usage accounting, max_tokens cutoffs, error shapes. Give per-request cancellation particular respect. A client disconnecting mid-stream has to free its KV blocks and leave the batch cleanly, and “streaming token generator with cancellation” shows up as a literal coding exercise at serving companies. This capstone also matters later: the flash-attention quest asks you to swap your own kernel into this engine and keep the probe green.

From the library
Tasks
  • ~1.2k lines: prefix caching, TP, CUDA graphs. The nanoGPT of inference engines — your blueprint.

    read+80 XPresource ↗
  • Requests join/leave the batch at token boundaries; block-allocated KV; SSE streaming out with per-request cancellation and timeouts (the 'streaming token generator with cancellation' is a recurring Fireworks/Together coding exercise). Any small model.

    build+300 XP
  • Outlines-style FSM masking, or integrate xgrammar. Read the vLLM structured-decoding intro first.

    build+120 XPresource ↗
  • Expose your engine as an OpenAI-compatible API and point InferQuest's prober at it: /v1/models, chat completions, usage accounting, max_tokens cutoff, SSE streaming framing with [DONE], and error shapes — all live-tested.

    build+400 XPauto-verified