Go Faster
The modded-nanoGPT lineage: 45 minutes to 74 seconds, one trick at a time.
The NanoGPT speedrun is this path's version of Boehm's worklog: 89 records that took GPT-2 training from 45 minutes to 74 seconds, each one a named, measured technique. Its biggest export, Muon, went from speedrun trick to torch.optim to training Kimi K2 at a trillion parameters. Read the record history as a method, then earn the ideas by A/B-ing Muon against your own AdamW. The speedrun's discipline — one change, measured, statistically defended — is the actual curriculum.
The skeptic's paper is assigned right next to the hype on purpose. Under fair tuning, most claimed 2× optimizer speedups shrink to about 1.1× at even modest scale. Hold both facts at once. The capstone milestone points the same discipline at your own loop: make it measurably faster without losing loss, graded as a same-device A/B so your hardware doesn't matter.
All 89 records and what each changed: Muon, QK-norm, ReLU², untied embeddings, FP8 matmuls, window schedules. The method matters more than any single trick.
Newton-Schulz orthogonalization on the momentum of 2D weights; embeddings/head stay on AdamW. It ships in PyTorch 2.13 as torch.optim.Muon — implement it first, then check yours against the real one.
Fused RMSNorm/RoPE/CE Triton kernels (~20% throughput, big memory cuts) and FlashAttention-2 document-masked packing. Measure, don't trust the README.
float8 matmuls compose with torch.compile and FSDP2 (1.3–1.5× on real pretraining). On consumer GPUs support varies — knowing where it does and doesn't apply is the point.
The reality check: under fair tuning, matrix-preconditioner gains shrink toward 1.1× as models grow. Speedrun claims need this next to them.
The harness trains its own plain fp32 AdamW baseline on a fixed 1M-token workload, then times YOUR loop on an identical model and budget — same device, so any GPU can pass. ≥1.5× wall-clock with final val loss within 0.05 nats. compile, bf16, fused optimizers, Muon — earn it however you like.
build+200 XPauto-verified