InferQuest
← Quest map
Inference Phase 3 · Production

SGLang & the Benchmark

Two engines, one methodology, publishable numbers.

0/6 · 890 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

SGLang gets equal billing because job postings list it as vLLM's equal, and because RadixAttention is a genuinely different idea: the KV cache as a prefix tree shared across requests, which turns multi-turn and agentic traffic (everyone's traffic now) into a cache-hit problem. The methodology reading sits before the head-to-head on purpose. Most published LLM benchmarks are subtly wrong, usually through unrealistic length distributions or a single conflated latency number, and the goal of this quest is to make you incapable of producing one.

Report the head-to-head as Pareto curves with TTFT and ITL separated, and goodput at a stated SLO. A bare tokens/sec claim with no latency constraint is the field's most common lie. That framing carries straight into the observability phase, and the published post at the end is an artifact a laptop-bound candidate can't fake: your numbers, your hardware, reproducible configs.

From the library
Tasks
  • Prefix-tree KV reuse. SGLang is co-equal with vLLM in job postings — know both.

    paper+60 XPresource ↗
  • TTFT, ITL/TPOT, goodput, and the measurement pitfalls. Realistic length distributions or it doesn't count.

    read+80 XPresource ↗
  • Latency-throughput Pareto curves, TTFT/ITL separated, goodput at a stated SLO, cost/M tokens.

    bench+200 XP
  • InferQuest streams real completions against your endpoint and takes medians of 3 runs: TTFT ≤ 800ms and ≥ 25 tok/s single-stream. Tune until it passes.

    bench+250 XPauto-verified
  • The post no laptop-bound candidate can write — real hardware, real methodology, reproducible configs.

    write+200 XPauto-verified
  • build+100 XPresource ↗