InferQuest
← Quest map
Training Phase 5 · Receipts

Open Weights, Open Doors

Merged PRs into the training stack, and the interview drilled boring.

0/6 · 1150 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

The allowlist here is the training stack's hiring-signal list: TRL, Unsloth, Axolotl, torchtitan, OLMo-core, datatrove, the eval harnesses, verl, nanochat. The documented precedent is stronger on this path than anywhere: Keller Jordan was hired onto OpenAI pretraining off the NanoGPT speedrun, Maxime Labonne went from open finetunes to shipping Liquid's production models, and training postings at Anthropic, Liquid, and Prime Intellect list open-source contributions as preferred quals or acceptable in lieu of publications. The honest counterweight: another record-holder writes “I don't work in AI” — the artifact opens doors; the interview leg still has to carry you through them.

The drills mirror the reported rounds: backprop live in Python, scaling arithmetic on demand, debugging a diverging run, designing a post-training pipeline (data → SFT → preference → RL → evals). Anthropic interviews in Colab with references allowed and states half its technical staff had no prior ML experience — the bar is what you can do, which is exactly what this path spent five phases building receipts for.

Tasks
  • TRL, Unsloth, Axolotl, torchtitan, OLMo-core, datatrove, lm-eval — pick the repo whose internals you already walked in this path.

    oss+100 XPresource ↗
  • Checked live against the GitHub API: exists, merged, non-trivial.

    oss+300 XPauto-verified
  • A training-quality or performance change with measured evidence in the description. Same allowlist, same live verification.

    oss+350 XPauto-verified
  • Backprop from scratch in Python against a clock, 6ND/scaling arithmetic on demand, “this run is diverging — debug it”, and the post-training pipeline design (data → SFT → preference → RL → evals). Timed, alone, spoken.

    build+100 XP
  • A serious attempt at modded-nanogpt's optimization track (hardware-agnostic, step-count-scored) or a nanochat time-to-GPT-2 entry. The artifact channel with the strongest documented precedent.

    build+150 XPresource ↗
  • From the researched archetypes: RL/post-training first (largest category, lowest experience bar — xAI says “relevant experience is not required”), then pretraining and data roles; frontier labs, Liquid/Zyphra/Prime Intellect, AI2, Apple Foundation Models.

    build+150 XP