Feeding the Beast
Data quality is a research problem — the job postings say so verbatim.
The FineWeb blogpost is the central text, and not for the dataset — for the method. Every filtering decision is ablated with real training runs, which is what “data quality as a research problem” (OpenAI's posting language) looks like in practice. The counterintuitive findings are the value: per-snapshot dedup beat global dedup, and more aggressive cleaning is not monotonically better. DCLM matters as the benchmark formulation and the origin of the CORE metric you'll meet again in evals.
The pipeline build is the employable skill: run a real Common Crawl slice through datatrove's extract, filter, and dedup stages, and report what each stage kills. TinyStories is the cheapest profound result in the field. Restrict the data distribution and coherent English emerges in models a thousandth the size — the ancestor of the whole synthetic-data arc (Cosmopedia through SYNTH) now feeding production small models.
The methodology is the content: how each filter and dedup decision was validated with training runs, and which intuitive cleanups turned out to hurt.
Fixed token pool, fixed training recipe, curation as the only variable — plus the low-noise CORE eval used by nanochat and the speedruns.
datatrove: extraction → Gopher/C4 quality filters → MinHash dedup. Report document survival rates per stage and inspect what died — the inspection is the skill.
Train your loop on TinyStories vs a same-budget raw-web sample and compare generations. Then skim the synthetic-data arc: Cosmopedia → SYNTH.
Dedup, filtering rules, FineWeb-Edu's classifier, CORE, synthetic data. 75% to pass.
drill+80 XPauto-verified