InferQuest
← Quest map
Inference Phase 5 · SLOs & Cost

The Cost of a Token

GPU-hour → tokens → margin, from first principles.

0/4 · 530 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

The tensoreconomics piece is the best published derivation of cost per token from first principles. It connects the bandwidth arithmetic from the KV-cache quest to dollars, a translation most engineers never learn to make. InferenceMAX is the methodology reference: cost claims as Pareto frontiers across hardware with stated assumptions, no single-number cherry-picking. Read it as a template for making a cost claim you'd defend under cross-examination.

Then the capstone: a cost model for the hardware you actually run, from amortized hardware and power, through measured throughput, to $/M tokens per model and quantization. Almost nobody walks into an interview with a cost model built from real bills and real measured throughput. That's what makes the write-up at the end one of the strongest artifacts on this roadmap.

From the library
Tasks
  • The best cost-per-token derivation published. Also skim Epoch AI's inference-economics piece.

    read+80 XPresource ↗
  • Pareto frontiers and $/M-token across hardware — the reference for how cost claims get made honestly.

    read+50 XPresource ↗
  • Amortized hardware + power → $/M tokens at measured throughput, per model and quantization. This spreadsheet is interview gold.

    build+150 XP
  • The cost story with real numbers from your own hardware. Your single best piece of evidence.

    write+250 XPauto-verified