InferQuest
← Quest map
Inference Phase 3 · Production

vLLM, Deeply

Not 'can use it' — 'has read it'.

0/4 · 600 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

The Anatomy post is the closest thing to a textbook chapter on a production engine, and the task's V1 warning matters: the internet is thick with V0-era posts describing a scheduler that no longer exists, and repeating them in an interview is a tell. The request trace is the real work of this quest. Writing down every file and class one request touches is the difference between “has used vLLM” and “has read it,” and it's exactly the preparation for landing PRs in the Arena phase.

When you tune your deployment, one knob at a time with a results table isn't pedantry. Several knobs interact (chunked-prefill budget against prefix-cache hit rate, quantized KV against batch capacity), and the table of what each did on your hardware is both how you'll actually understand them and a solid interview artifact in its own right.

From the library
Tasks
  • The closest thing to a textbook chapter on engine internals. Study the V1 architecture — V0 posts describe a dead engine.

    read+100 XPresource ↗
  • Arrival → scheduler → block manager → model runner → sampler → stream. Write down every file and class it touches.

    read+150 XP
  • max-num-batched-tokens, prefix caching, chunked prefill, quantized weights + FP8 KV. One knob at a time, keep a results table.

    build+150 XPresource ↗
  • Expose it (with an API key) and let InferQuest probe it live — same suite your toy engine passed, now against production infrastructure.

    build+200 XPauto-verified