InferQuest
← Quest map
Inference Phase 5 · SLOs & Cost

Eyes on the Fleet

If it isn't graphed, it isn't served.

0/3 · 310 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

A short quest, because the skill is a discipline rather than a literature. The vLLM metrics doc tells you what the engine already exports (KV utilization, queue depth, TTFT histograms), and the work is wiring it up so your fleet alerts before OOM, not after. The OTel GenAI conventions are a quick read with a career-shaped reason attached: the standard is young enough that knowing it puts you ahead of most incumbents, and shared naming conventions are how your observability work stays legible to other teams.

Over-invest in the SLO task. Define TTFT and ITL targets, run load, report the percentage of requests meeting both. That's goodput, the DistServe idea turned into an operational number, and it appears nearly verbatim in serving-team job postings, Anthropic's included.

From the library
Tasks
  • KV utilization, queue depth, TTFT histograms, per-GPU DCGM. Alert before OOM, not after.

    build+150 XPresource ↗
  • The emerging standard for LLM trace/metric naming.

    read+40 XPresource ↗
  • Pick TTFT/ITL targets, run load, report the % of requests meeting both. This is the Anthropic-posting skill verbatim.

    bench+120 XP