The Cost of a Token
GPU-hour → tokens → margin, from first principles.
The tensoreconomics piece is the best published derivation of cost per token from first principles. It connects the bandwidth arithmetic from the KV-cache quest to dollars, a translation most engineers never learn to make. InferenceMAX is the methodology reference: cost claims as Pareto frontiers across hardware with stated assumptions, no single-number cherry-picking. Read it as a template for making a cost claim you'd defend under cross-examination.
Then the capstone: a cost model for the hardware you actually run, from amortized hardware and power, through measured throughput, to $/M tokens per model and quantization. Almost nobody walks into an interview with a cost model built from real bills and real measured throughput. That's what makes the write-up at the end one of the strongest artifacts on this roadmap.
- Inference EngineeringPhilip Kielybook↗
The best cost-per-token derivation published. Also skim Epoch AI's inference-economics piece.
Pareto frontiers and $/M-token across hardware — the reference for how cost claims get made honestly.
Amortized hardware + power → $/M tokens at measured throughput, per model and quantization. This spreadsheet is interview gold.
build+150 XPThe cost story with real numbers from your own hardware. Your single best piece of evidence.
write+250 XPauto-verified