Quest map
One Foundations trunk, two specialist paths. Foundations quests count for both paths — progress and XP are a single pool — and each specialist quest unlocks when its prerequisites are at least half done, wherever they live.
Foundations
The fundamentals both jobs are built on — performance mental models, transformer internals, GPU architecture, kernels, quantization theory, parallelism. Whichever path you’re headed for, this is where it starts; finish a quest here once and it counts everywhere.
Bedrock
The performance mental models and PyTorch fluency everything else builds on. Skimmable if you're strong here — but don't skip Horace He.
Transformer Internals
Build GPT from scratch, then learn the architecture zoo through the serving lens: attention variants, positional encodings, tokenizers, sampling, MoE.
GPU Architecture & CUDA
PMPP, the memory hierarchy, occupancy, rooflines, and profiling with Nsight — learn why kernels are fast, and prove yours are.
PMPP + GPU MODE: the canonical on-ramp.
5 tasks
The rite of passage: chase cuBLAS.
3 tasks
Nsight is your microscope; the roofline is your map.
6 tasks
torch.compile is load-bearing in vLLM V1 — stop treating it as magic.
3 tasks
Kernel Engineering
Triton from puzzles to a working flash attention, the FlashAttention lineage, and the modern kernel-library landscape (FlashInfer, CUTLASS, ThunderKittens).
Quantization Theory
GPTQ, AWQ and SmoothQuant as one argument about outliers; FP8 as the production default; NVFP4/MXFP4 as the frontier; KV quantization. The theory both serving and training compress with.
Parallelism
Collectives from scratch, Megatron TP by hand, and the communication math — the shared foundation under both distributed serving and distributed training.
Inference Engineering
Serve LLMs fast and cheap — from KV caches and kernels to production fleets.
The Inference Engine
KV caching, continuous batching, paged attention, speculative decoding — then the capstone: build your own engine and pass a live OpenAI-conformance probe against it.
The single most important idea in LLM serving.
3 tasks
Continuous batching is why serving companies exist.
4 tasks
Free tokens, provably distribution-preserving.
4 tasks
When the cache can't hold everything, something has to give.
3 tasks
The capstone: a real serving engine, probed live by InferQuest.
4 tasks
Quantization
Quantization put to work: compress real models with llm-compressor, prove quality with lm-eval, and publish the study. The theory lives in Foundations.
Production Serving
vLLM V1 and SGLang on your own hardware: deploy, tune, benchmark with real methodology — and pass live conformance and latency probes against your endpoints.
Distributed Inference
Disaggregated prefill-decode, KV-cache tiering and transfer, MoE serving at scale, and the Kubernetes layer (llm-d, Gateway API, DRA) — built on the Parallelism foundation.
Observability & Economics
Metrics, SLOs, goodput, capacity planning, and cost-per-token from first principles — the part that makes you dangerous in a business conversation.
The Arena
Merged PRs into the engines everyone runs, public benchmarks nobody can argue with, interview gauntlets, and the offer.
Model Training
Build LLMs and make them as good as possible on the cheapest hardware available.
Learning to Learn
The other half of the transformer: gradients, optimizers, and a loop that converges — plus data curation and the scaling-laws math that decides every training run.
The Speedrun
The modded-nanoGPT lineage: Muon, FP8, fused kernels, multi-GPU training — then the capstone: pretrain a real GPT-2-class model on your own hardware or fifty dollars.
Post-Training
From base model to assistant to reasoner: SFT, LoRA, DPO, then GRPO/RLVR on a single consumer GPU — the skills the largest training-side hiring category wants.
Proof
Evals are the training world's observability: harnesses, contamination, small-model pitfalls, and publishing models with honest numbers.
The Open Ladder
Merged PRs into the training stack, the training-flavored interview drilled boring, and the target list for pretraining, post-training, and RL roles.