Prove It
Evals are the training world's observability.
Every claim this path produces — the capstone, the adapters, the RL lift — is only as good as its eval, and this quest is the discipline that makes the rest credible. Run both harnesses on the same model and same benchmark (lm-eval-harness is the de facto standard; lighteval is what HF's own projects use) and let the discrepancies teach you how much implementation details move scores. That lesson generalizes to every leaderboard you'll ever read.
GSM1k is the contamination result to internalize: up to 8-point drops when models meet genuinely fresh problems. OLMES exists because small base models are especially easy to mis-measure; cloze versus multiple-choice formulation changes rankings. The finale is the artifact that matters — a model on the Hub with an honest card, and a writeup whose numbers you can defend, regressions included. Same rule as the quantization path, same reason.
Same benchmark, both harnesses, note every discrepancy and find its cause (prompt format, few-shot selection, normalization). The discrepancy IS the lesson.
Fresh GSM8K-style problems, up to 8-point drops. Then look at how OLMo 3 ships decontamination tooling as part of the release.
Cloze vs multiple-choice formulations, curated few-shots — the standardization that makes sub-7B comparisons meaningful. Pair with “The Leaderboard Illusion” (2504.20879).
Model on the HF Hub with a real card, and a public writeup with the numbers, the methodology, and the regressions. The writeup is URL-verified now; Hub-checking becomes its own verifier later.
write+150 XPauto-verified