InferQuest
← Quest map
Foundations Phase 2 · Foundations

The Forward Pass

Build GPT-2 from nothing, load real weights, sample text.

0/6 · 630 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

Karpathy doesn't present a transformer, he derives one, starting from a bigram model, so every block exists to solve a problem you've already felt. When you build yours, two details sink most re-implementations: layernorm placement (GPT-2 is pre-LN, with a final layernorm after the last block) and softmax stability (subtract the row max; the attention grader feeds you large logits on purpose). If real weights produce rambling text, check transposes before anything else. The HF GPT-2 checkpoint stores its linear layers Conv1D-style.

The tokenizer video looks optional and isn't. BPE explains half of all “weird LLM behavior,” and streaming detokenization (why an engine can't just emit one string per token) is a real serving problem you'll meet again when you build an engine. The sampling zoo is verbatim interview material at several serving companies.

From the library
Tasks
  • watch+40 XPresource ↗
  • From scratch in PyTorch — embeddings, attention, MLP, layernorm placement. No nn.Transformer anything.

    build+150 XP
  • Implement scaled dot-product attention (causal + non-causal, numerically stable) and pass the harness — checked against torch SDPA including a large-logit stability case.

    build+150 XPauto-verified
  • Pull the HF checkpoint into your implementation. If it rambles, your shapes are wrong somewhere.

    build+100 XP
  • BPE from scratch. Explains half of all 'weird LLM behavior' — and why streaming detokenization is subtle.

    build+100 XPresource ↗
  • Min-p (ICLR 2025) ships in every engine now — read arxiv.org/abs/2407.01082 alongside. Include beam search: Perplexity hands candidates its exact signature and unit tests, and Mistral asks top-k/top-p from scratch with no libraries.

    build+90 XPresource ↗