Eyes on the Fleet
If it isn't graphed, it isn't served.
A short quest, because the skill is a discipline rather than a literature. The vLLM metrics doc tells you what the engine already exports (KV utilization, queue depth, TTFT histograms), and the work is wiring it up so your fleet alerts before OOM, not after. The OTel GenAI conventions are a quick read with a career-shaped reason attached: the standard is young enough that knowing it puts you ahead of most incumbents, and shared naming conventions are how your observability work stays legible to other teams.
Over-invest in the SLO task. Define TTFT and ITL targets, run load, report the percentage of requests meeting both. That's goodput, the DistServe idea turned into an operational number, and it appears nearly verbatim in serving-team job postings, Anthropic's included.
- Inference EngineeringPhilip Kielybook↗
KV utilization, queue depth, TTFT histograms, per-GPU DCGM. Alert before OOM, not after.
The emerging standard for LLM trace/metric naming.
Pick TTFT/ITL targets, run load, report the % of requests meeting both. This is the Anthropic-posting skill verbatim.
bench+120 XP