Go Brrrr
The three-bottleneck worldview: compute, memory, overhead.
Nearly every performance question in this field comes down to which of three bottlenecks you're hitting: compute, memory bandwidth, or overhead. Horace He's post is where that taxonomy comes from, and his factory-and-warehouse framing turns arithmetic intensity into something you can picture instead of a formula to memorize. The operator-fusion section deserves a slow read; it explains why a chain of pointwise ops can be nearly free, and why eager-mode PyTorch sometimes isn't. The post does assume you've trained something before and know words like training loss and overfitting. If you don't yet, take the on-ramp task first.
Keep the taxonomy in your head for the rest of the map. The central fact of inference is that decode is memory-bandwidth-bound: generating one token streams every weight, plus the entire KV cache, through HBM to do a comparatively tiny amount of math. Batching, quantization, paged KV, and speculative decoding are each an attack on one of the three bottlenecks. For any technique on this map you should be able to name which one.
- Build a Large Language Model (From Scratch)Sebastian Raschkabook↗
- Neural Networks: Zero to HeroAndrej Karpathycourse↗
- PyTorch internalsEdward Z. Yangpost↗
- Let's talk about the PyTorch dispatcherEdward Z. Yangpost↗
- Making Deep Learning Go Brrrr From First PrinciplesHorace Hepost↗
If training loss, gradients, and overfitting aren't familiar words yet, start here: Karpathy builds neural nets from scratch, ending at a small GPT. Skip it if you've trained anything before — the reading below assumes that background.
Horace He's compute/memory/overhead taxonomy. Every perf conversation in this field silently assumes this post.
SMs, warps, occupancy, HBM, TMA, NVLink — get the vocabulary loaded before it's needed.
Only if hand-multiplying two small matrices feels shaky. Do the multiplication exercises in the matrix unit until they're boring — an hour or two of instant-feedback reps. Skip freely if the mechanics are already second nature.
In a fresh notebook, write matmul three times: (1) three nested loops; (2) each output entry as a dot product of a row of A with a column of B; (3) the output as a weighted sum of A's columns. Check all three against torch.matmul on random shapes. Then make Q and K with torch.randn(4, 128, 64) and — before running each line — write down the shape of Q @ K.transpose(-2, -1), of the scores divided by √64, and of their softmax. Done when the shape predictions stop being guesses. Tensor Puzzles in the next quest continues this in NumPy.
build+60 XPIntuition only, by design — matmul as composition of transformations, dot products as projection. It lands hardest AFTER the drill above, not instead of it. The determinant and eigen chapters can wait until the map actually needs them (low-rank ideas show up with LoRA, much later).