InferQuest
← Quest map
Inference Phase 2 · Compression

Compress & Prove It

Quantization without evals is vandalism.

0/5 · 740 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

The from-scratch quantizer comes first because scale-and-zero-point per group is trivial to describe and instructive to get exactly right. The grader's outlier-heavy weights show you precisely why per-tensor fails, which is the whole motivation of the theory quest made concrete. The sensitivity scan in the middle is the professional habit: measure where the model is fragile and spend precision there, before reaching for anyone's uniform recipe.

llm-compressor and lm-eval-harness are the production pairing. One produces the checkpoint vLLM actually loads; the other tells you what it cost you. The discipline this quest teaches is right there in the tagline: a speedup number without a quality delta next to it is not a result. Publish the regressions too. Honest numbers are rarer than good ones, and hiring managers can tell the difference.

From the library
Tasks
  • Scale/zero-point derivation, per-group granularity, dequantize — no quantization libraries. Graded on 4-bit validity, beating the per-tensor baseline ≥3× on outlier-heavy weights, and keeping the reference GPT's predictions intact end-to-end.

    build+200 XPauto-verified
  • MIT 6.5940-style: quantize one layer (or group) at a time, measure the damage, and let the scan choose where precision goes — before reaching for uniform recipes.

    bench+90 XP
  • The production pipeline for vLLM. Produce both checkpoints from the same base model.

    build+150 XPresource ↗
  • Quality deltas on real tasks, plus throughput deltas from your benchmark harness. Red Hat's half-million-eval study is your methodology template.

    bench+150 XPresource ↗
  • Quality × speed × memory, real numbers, honest about regressions.

    write+150 XPauto-verified