Under the Tensor
What a tensor actually is, and where Python time goes.
ezyang's post packs years of PyTorch core knowledge into one read. The idea to hold onto is strides: a tensor is a flat buffer plus indexing math, and view/permute/slice never copy anything. The same idea shows up later as paged-KV block tables, Triton pointer arithmetic, and coalesced-access analysis, which is why the strided-tensor build is here even though it feels like a detour.
py-spy is here because the third bottleneck, overhead, is usually Python. In a real engine the model runs on GPU while the scheduler, API server, and detokenizer are host code, and profiling is how you find out when they're the actual ceiling. Run it against something real you own, not a toy.
- Neural Networks: Zero to HeroAndrej Karpathycourse↗
- PyTorch internalsEdward Z. Yangpost↗
- Let's talk about the PyTorch dispatcherEdward Z. Yangpost↗
- Making Deep Learning Go Brrrr From First PrinciplesHorace Hepost↗
- PyTorch Developer PodcastEdward Z. Yangpodcast↗
Tensor vs Storage, strides, dispatch, autograd — from a PyTorch core dev.
Sasha Rush's drills: broadcasting fluency without loops. Do all of them.
Engine host overhead (scheduler, API server) is Python. Find a real bottleneck in any project you own.
CMU 10-714 hw3-style: views share one flat buffer (zero copy), plus a compact() that materializes. This indexing math underlies paged-KV layouts, Triton pointer arithmetic, and coalescing analysis.