← Quest map
Phase 5 · Compression
Precision Games
GPTQ → AWQ → FP8 default → NVFP4/MXFP4 frontier.
0/6 · 430 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Tasks
Hessian-based compensation vs salient-channel scaling vs outlier migration. Interviews ask you to compare them.
FP8 W8A8 is the boring production default; block-scaled 4-bit is the Blackwell-era shift (GPT-OSS ships in MXFP4).
Why keys quantize per-channel and values per-token. FP8 KV is a near-free 2× cache-capacity win.
Wanda-style pruning, 2:4 semi-structured sparsity on sparse tensor cores (TRT-LLM ships it), and Minitron-style distillation — which also trains spec-decode drafts.
Formats, methods, quality tradeoffs, KV quantization. 75% to pass.
drill+80 XPauto-verified