InferQuest

Quest map

One Foundations trunk, two specialist paths. Foundations quests count for both paths — progress and XP are a single pool — and each specialist quest unlocks when its prerequisites are at least half done, wherever they live.

The Trunk

Foundations

The fundamentals both jobs are built on — performance mental models, transformer internals, GPU architecture, kernels, quantization theory, parallelism. Whichever path you’re headed for, this is where it starts; finish a quest here once and it counts everywhere.

Phase 1 · Mental Models

Bedrock

The performance mental models and PyTorch fluency everything else builds on. Skimmable if you're strong here — but don't skip Horace He.

Phase 2 · Foundations

Transformer Internals

Build GPT from scratch, then learn the architecture zoo through the serving lens: attention variants, positional encodings, tokenizers, sampling, MoE.

Phase 3 · The Metal

GPU Architecture & CUDA

PMPP, the memory hierarchy, occupancy, rooflines, and profiling with Nsight — learn why kernels are fast, and prove yours are.

Phase 4 · Kernelcraft

Kernel Engineering

Triton from puzzles to a working flash attention, the FlashAttention lineage, and the modern kernel-library landscape (FlashInfer, CUTLASS, ThunderKittens).

Phase 5 · Precision

Quantization Theory

GPTQ, AWQ and SmoothQuant as one argument about outliers; FP8 as the production default; NVFP4/MXFP4 as the frontier; KV quantization. The theory both serving and training compress with.

Phase 6 · Many GPUs

Parallelism

Collectives from scratch, Megatron TP by hand, and the communication math — the shared foundation under both distributed serving and distributed training.

Path 01

Inference Engineering

Serve LLMs fast and cheap — from KV caches and kernels to production fleets.

Phase 1 · Engine Core

The Inference Engine

KV caching, continuous batching, paged attention, speculative decoding — then the capstone: build your own engine and pass a live OpenAI-conformance probe against it.

Phase 2 · Compression

Quantization

Quantization put to work: compress real models with llm-compressor, prove quality with lm-eval, and publish the study. The theory lives in Foundations.

Phase 3 · Production

Production Serving

vLLM V1 and SGLang on your own hardware: deploy, tune, benchmark with real methodology — and pass live conformance and latency probes against your endpoints.

Phase 4 · Datacenter Scale

Distributed Inference

Disaggregated prefill-decode, KV-cache tiering and transfer, MoE serving at scale, and the Kubernetes layer (llm-d, Gateway API, DRA) — built on the Parallelism foundation.

Phase 5 · SLOs & Cost

Observability & Economics

Metrics, SLOs, goodput, capacity planning, and cost-per-token from first principles — the part that makes you dangerous in a business conversation.

Phase 6 · Proof of Work

The Arena

Merged PRs into the engines everyone runs, public benchmarks nobody can argue with, interview gauntlets, and the offer.

Path 02

Model Training

Build LLMs and make them as good as possible on the cheapest hardware available.

Phase 1 · Optimization

Learning to Learn

The other half of the transformer: gradients, optimizers, and a loop that converges — plus data curation and the scaling-laws math that decides every training run.

Phase 2 · Efficiency

The Speedrun

The modded-nanoGPT lineage: Muon, FP8, fused kernels, multi-GPU training — then the capstone: pretrain a real GPT-2-class model on your own hardware or fifty dollars.

Phase 3 · SFT & RL

Post-Training

From base model to assistant to reasoner: SFT, LoRA, DPO, then GRPO/RLVR on a single consumer GPU — the skills the largest training-side hiring category wants.

Phase 4 · Evals

Proof

Evals are the training world's observability: harnesses, contamination, small-model pitfalls, and publishing models with honest numbers.

Phase 5 · Receipts

The Open Ladder

Merged PRs into the training stack, the training-flavored interview drilled boring, and the target list for pretraining, post-training, and RL roles.