1. What is it?#
The economics of inference: where the money goes, why costs grow the way they do, and which engineering levers actually move the number.
The headline: an LLM token costs roughly a million times more to produce than a database row costs to fetch. At scale, that ratio makes inference efficiency a first-order business concern rather than an engineering nicety.
2. Why does it exist?#
Three compounding factors.
Factor 1 — the hardware is expensive and must be dedicated. An H100 costs $25-40k to buy or $2-5/hour to rent. Unlike a CPU, it cannot be meaningfully time-shared across unrelated workloads at fine granularity: the model weights occupy its memory. So the GPU is yours whether you use it or not.
Factor 2 — the work per user interaction is enormous. A single chat turn with a 70B model generating 500 tokens costs ~70 TFLOPs of arithmetic and reads ~70 TB of memory traffic. That is more compute than a typical web server does in a day.
Factor 3 — utilization is hard. Traffic is bursty; models are memory-bound at low batch; long requests block short ones. A GPU running at 20% effective utilization costs five times as much per token as one at 100%.
Multiply those together and you get the industry’s defining cost problem.
3. Simple analogy#
Chartering a private jet versus taking a bus.
The jet (GPU) costs $5,000/hour whether it carries 1 passenger or 12. Your job as the operator is to fill the seats. One passenger → $5,000/seat. Twelve → $417/seat.
Now add the wrinkles that make it hard:
- Passengers arrive at random times, and you can’t wait forever (latency SLO).
- Some passengers want to fly 100 miles, others 4,000 (variable output length).
- The jet takes 8 minutes to spin up (cold start), so you can’t summon one per passenger.
- Empty seats can’t be sold later — capacity is perishable.
Every technique in this curriculum is a way to sell more seats per flight without making anyone wait too long.
4. Tiny example#
Compute cost per million tokens from first principles.
Setup: Llama-3-70B, FP16, 8×H100 node, TP=8
Node cost: $8/hour on-demand-ish (use your real number)
Step 1 — throughput per node
Per-GPU weights: 140/8 = 17.5 GB
Decode step time at batch 64: bytes ≈ 17.5 GB + KV(64 seq × 4k × 40 KB/8) ≈ 18.8 GB
T = 18.8/3350 = 5.6 ms, plus ~1.5 ms comms/overhead ≈ 7.1 ms
Output tokens/sec = 64 / 0.0071 = 9,014 tok/s
Step 2 — cost per output token
$8/hour ÷ 3600 = $0.00222 /sec
$0.00222 / 9014 tok/s = $2.46e-7 per token
→ $0.246 per million output tokens ← at 100% utilization, batch 64
Step 3 — reality adjustments
Average utilization 40% → ÷0.40 → $0.62 / M tokens
Prefill cost (prompt tokens, typically 5-20% of total GPU time) → ×1.15 → $0.71
Redundancy / spare capacity (N+1) → ×1.3 → $0.93
Failed & abandoned generations (~10%) → ×1.1 → $1.02 / M output tokensNote the shape of that calculation. The raw hardware cost was $0.25/M tokens. Everything after Step 2 — utilization, redundancy, waste — multiplied it by 4x. Most cost optimization is not about making the GPU faster; it is about making sure the GPU is doing useful work.
Now the sensitivity analysis, which is the actually useful output:
| Change | New cost/M | Multiplier |
|---|---|---|
| Baseline | $1.02 | 1.00x |
| Batch 64 → 8 (bad batching) | $6.80 | 6.7x worse |
| Batch 64 → 256 | $0.42 | 2.4x better |
| FP16 → FP8 | $0.55 | 1.9x better |
| Utilization 40% → 70% | $0.58 | 1.8x better |
| 70B → 8B model | $0.14 | 7x better |
| + prefix caching (60% hit rate on prompts) | $0.85 | 1.2x better |
| All good changes stacked | ~$0.09 | 11x better |
An 11x cost difference between a naive and a well-engineered deployment of the same model. That number is the entire justification for this field.
5. Technical explanation#
The cost model#
cost_per_token = (hourly_cost_of_hardware)
÷ (tokens_per_hour_produced)
tokens_per_hour = 3600 × batch_size / step_time_seconds × utilizationExpanding, the levers are exactly:
hardware_$/hr
cost/token = ─────────────────────────────────────────
3600 × B/T_step × U × (1 − waste)
lever what it changes typical range
────────────────────────────────────────────────────────────
hardware $/hr per unit of bandwidth 1-3x
B batch size (scheduler quality) 2-20x
T_step precision, kernels, model size 2-8x
U utilization (traffic shaping) 1.5-4x
waste cancellations, retries, padding 1.1-1.5xMultiplying the achievable ranges: 10-100x total spread. This is unusually large for a software engineering discipline, and it is why the same model is offered by different providers at prices that differ by an order of magnitude.
Where the money actually goes#
For a typical production LLM deployment:
GPU compute 70-85%
├─ decode 50-70% of GPU time
├─ prefill 15-35%
└─ idle / bubbles 5-30% ← the recoverable part
Networking (esp. multi-node) 2-8%
Storage (weights, caches, logs) 2-5%
CPU hosts (tokenization, gateway) 3-8%
Observability & logging 1-4%The “idle / bubbles” line is where optimization effort pays. Sources of bubbles:
- Batch not full (traffic too low, or scheduler too conservative)
- Head-of-line blocking by long prefills
- Preemption/swapping when KV memory is exhausted
- Launch overhead and CPU stalls (no CUDA graphs, slow Python)
- Load imbalance across replicas
- Draining before deployments
Why cost scales superlinearly with context length#
This surprises people. Doubling the context does not double the cost:
1. KV cache doubles → concurrent sequences halve → batch halves
→ per-token cost roughly doubles
2. Attention FLOPs per token double (linear in S during decode)
3. Prefill FLOPs quadruple (quadratic in S)
4. Longer sequences occupy slots longer → more head-of-line blockingEmpirically, going from 4k to 32k context often costs 4-8x per token, not 8x-in-prompt-length. This is why long-context pricing is disproportionate and why context compression, prefix caching, and MLA-style KV reduction are commercially important, not just academically interesting.
Prefill vs decode economics#
Prompt-heavy workload (RAG, summarization, code review):
10,000 input : 500 output
Prefill dominates GPU time. Optimize: prefix caching, chunked prefill,
compute-efficient kernels, larger prefill batches.
Output-heavy workload (creative writing, long agent chains):
100 input : 4,000 output
Decode dominates. Optimize: quantization, batching, speculative decoding,
GQA/MLA, MoE.These want different hardware, different configurations, and sometimes different clusters. Providers price input and output tokens differently for exactly this reason — typically output tokens cost 3-5x input tokens, because decode is memory-bound and prefill is not.
6. Under the hood#
Where a “wasted” GPU-hour physically goes:
Timeline of one GPU-second under a mediocre scheduler:
|■■■■ decode batch 12 ■■■■|░░ idle ░░|■ prefill ■|░ preempt/swap ░|■■■ decode ■■■|
420 ms 180 ms 150 ms 90 ms 160 ms
Useful: 420+150+160 = 730 ms
Wasted: 270 ms (27%)
And the "useful" decode ran at batch 12 when memory allowed 48
→ effective utilization ≈ 0.73 × (12/48) ≈ 18%That second line is the one people miss. Running is not the same as running efficiently. A GPU that is 100% busy at batch 4 when it could run batch 64 is wasting 94% of its economic potential, and every conventional monitoring dashboard will show it as perfectly healthy.
This is why KV occupancy and average running batch size belong on your primary dashboard (Section XI.08).
7. Performance implications#
The ranked list of cost levers, by typical impact per unit of engineering effort:
| Rank | Lever | Typical gain | Effort |
|---|---|---|---|
| 1 | Use an engine with continuous batching | 3-10x | low (adopt vLLM/SGLang) |
| 2 | Prefix caching (if prompts share prefixes) | 1.3-5x | low-medium |
| 3 | Quantize weights to FP8/INT8 | 1.5-2x | low-medium |
| 4 | Right-size the model (do you need 70B?) | 2-10x | medium (needs eval) |
| 5 | Cancel abandoned generations | 1.1-1.4x | low |
| 6 | Raise utilization via traffic consolidation | 1.5-3x | medium (platform work) |
| 7 | Chunked prefill / better scheduling | 1.2-1.5x | low (config) |
| 8 | Speculative decoding | 1.3-2.5x | medium-high |
| 9 | CUDA graphs / kernel fusion | 1.1-1.4x | medium |
| 10 | Custom kernels | 1.05-1.3x | very high |
Note that the top of the list is mostly configuration and product decisions, not deep systems work. Teams routinely skip straight to #10 while leaving #1-#5 on the table. Do them in order.
8. Production implications#
- Instrument cost per million tokens as a first-class metric, per model and per tenant. Without it, you cannot tell whether an optimization worked.
- Attribute cost to tenants. In a multi-tenant platform, one customer with 100k-token prompts can consume 60% of your fleet. You need to know that (Section XII.05).
- Reserved capacity vs on-demand. GPU spot/on-demand pricing varies 2-4x. Steady baseline on reserved, spikes on on-demand, batch work on spot.
- The build-vs-buy calculation is real. At low volume, an API provider is cheaper than running your own GPUs (they have better utilization than you will). The crossover is usually somewhere around continuous utilization of a few GPUs. Compute it honestly, including the engineer-months.
- Cost and quality trade against each other explicitly. A router that sends easy queries to an 8B model and hard ones to 70B (Section XII.09) can cut cost 3-5x with small quality loss. Make that tradeoff deliberately, with measurements.
9. Common mistakes#
Optimizing the kernel before the scheduler. A 20% faster attention kernel on a system running at batch 4 is worth far less than fixing why the batch is 4.
Measuring cost at peak. Your cost is set by the average utilization across the whole day, including 3 a.m. Report cost/M tokens over 24h, not during a load test.
Ignoring prefill in the cost model. In RAG workloads prefill can be 70% of GPU time. Pricing and capacity models that only count output tokens will be badly wrong.
Forgetting redundancy and headroom. Running at 100% of capacity is not a plan. Budget N+1 and spike headroom; that is 20-40% of your fleet.
Assuming bigger GPUs are always better value. Compare $/(GB/s) and $/TFLOP across SKUs for your regime. For memory-bound decode, an H200 or even an L40S cluster can beat H100s per dollar, depending on price.
Not measuring the cost of your own platform overhead. Gateways, sidecars, logging pipelines, and observability can quietly add 10-15%.
10. Hands-on exercise#
A. Build your cost model. In a spreadsheet or script, implement the model from section 5 for a specific model and hardware you have access to. Include prefill, utilization, redundancy, and waste. Produce cost per million input tokens and per million output tokens separately.
B. Sensitivity analysis. Vary each lever (batch, precision, utilization, model size) ±50% and rank them by impact on cost. Which three matter most for your configuration? Does the ranking match the table in section 7?
C. Compare to market. Look up current API prices for a model of similar size. Compute the utilization you would need to break even against buying the API. What does that tell you about build-vs-buy at your volume?
D. Find the waste. On a running system, measure: average running batch size, KV occupancy,
fraction of GPU time in prefill vs decode vs idle. Compute effective utilization as
(time busy) × (avg batch / max batch). Most teams are shocked by this number the first time.
11. Interview questions#
- Derive cost per token from hardware cost and throughput. What are the terms?
- Why do output tokens cost more than input tokens?
- Why does cost scale superlinearly with context length? Give three mechanisms.
- Rank the top five cost optimizations for an LLM service and justify the order.
- Your GPU shows 100% utilization but cost per token is 4x your estimate. What do you investigate?
- When does it make sense to self-host versus use an API? Show the calculation.
- How would you attribute cost to individual tenants in a shared inference cluster?
12. Further reading#
- [ESTABLISHED] Public inference pricing pages (compare input vs output pricing across providers and note the ratios)
- [ESTABLISHED] Kwon et al., PagedAttention (SOSP 2023) — the throughput gains are cost gains
- [ESTABLISHED] Yu et al., “Orca” (OSDI 2022) — quantifies static-batching waste
- [REFERENCE] MLPerf Inference results — for hardware-per-dollar comparisons
- Next: Section II — Computer Systems Foundations