Precision Games
GPTQ → AWQ → FP8 default → NVFP4/MXFP4 frontier.
GPTQ, AWQ, and SmoothQuant are assigned together because the field's real content is the comparison: three answers to the same enemy, activation outliers. GPTQ repairs the damage with second-order information. AWQ protects the channels that activations say are salient. SmoothQuant migrates the difficulty from activations into weights. Seen that way, three algorithms to memorize collapse into one argument, and interviews ask exactly this compare-and-contrast.
Hold onto the two-tier frame from the formats task. FP8 W8A8 is the boring production default; block-scaled 4-bit (NVFP4, MXFP4) is where the frontier moved with Blackwell, and frontier models now ship natively in it. KIVI's asymmetry is the memorable detail: keys have channel-wise outlier structure and values don't, hence per-channel keys and per-token values. FP8 KV is the nearest-to-free 2× capacity win in serving, which plugs this quest straight back into the cache arithmetic.
Hessian-based compensation vs salient-channel scaling vs outlier migration. Interviews ask you to compare them.
FP8 W8A8 is the boring production default; block-scaled 4-bit is the Blackwell-era shift (GPT-OSS ships in MXFP4).
Why keys quantize per-channel and values per-token. FP8 KV is a near-free 2× cache-capacity win.
Wanda-style pruning, 2:4 semi-structured sparsity on sparse tensor cores (TRT-LLM ships it), and Minitron-style distillation — which also trains spec-decode drafts.
Formats, methods, quality tradeoffs, KV quantization. 75% to pass.
drill+80 XPauto-verified