The idea in one minute#
For AI infrastructure, cost is not an accounting afterthought; it is an engineering metric to graph next to latency. Two figures summarize a serving system: dollars per million tokens and tokens per joule. Both are simple divisions of counters you already have — GPU-hours and energy on top, tokens on the bottom.
The gap between what a GPU could produce and what it does produce is where the money goes: idle capacity, low batch occupancy, cache misses, and wasted work.
An analogy#
A delivery van. The lease and the driver cost the same whether the van is full or empty, so the cost per parcel is decided almost entirely by how full it runs. Fuel per parcel is a second, separate efficiency. A fleet manager tracks both, per route.
A picture#
flowchart TB
subgraph COST["What you pay for"]
GH["GPU-hours<br/>allocated, not used"]
EN["Energy<br/>joules from the GPU counter"]
end
subgraph WORK["What you got"]
TOK["Tokens served<br/>input, cached, output"]
GOOD["Good tokens<br/>from requests that met the SLO"]
end
GH --> D1["$ per million tokens"]
TOK --> D1
EN --> D2["tokens per joule"]
TOK --> D2
D1 --> LOSS["Where the gap goes:<br/>idle, low occupancy, cache misses,<br/>preempted and abandoned work"]
D2 --> LOSS
class GH,EN neutral
class TOK,GOOD memory
class D1,D2 compute
class LOSS warnHow it really works#
Dollars per million tokens#
cost per M tokens = (GPUs × $ per GPU-hour) ÷ (tokens per hour) × 1,000,000- The numerator is what you allocated, including idle time. An on-demand GPU and a reserved GPU sitting idle cost the same per hour.
- The denominator comes from the engine’s token counters (lesson 03) or the gateway’s.
- Split input and output: output tokens cost several times more to produce. A common practice is to apportion GPU time by measured phase time (prefill seconds vs decode seconds).
- Add the non-GPU share — CPU hosts, storage, network, the gateway — or state that you left it out.
As PromQL, with a recording rule holding your price per GPU-hour by node pool:
sum by (model) (gpu_allocated * on (node_pool) group_left gpu_hourly_price_dollars)
/
(sum by (model) (rate(vllm:generation_tokens_total[1h])) * 3600)
* 1e6The efficiency ladder#
Each step is a ratio you can measure; multiplied together they explain the cost.
| Ratio | Definition | Lost to |
|---|---|---|
| Allocation | GPUs allocated ÷ GPUs owned | Unscheduled capacity, fragmentation |
| Availability | Time ready to serve ÷ time allocated | Cold starts, loading, restarts, faults |
| Occupancy | Tokens served ÷ tokens the replica could serve at SLO | Under-filled batches, over-provisioning for peaks |
| Useful work | Tokens delivered ÷ tokens computed | Preempted work recomputed, cancelled streams, retries |
| Goodput | Tokens in SLO-meeting requests ÷ tokens delivered | Overload |
| Cache leverage | Cached prompt tokens ÷ prompt tokens | Poor routing, unstable prompts |
Occupancy is usually the largest loss. A fleet sized for peak traffic with a 3:1 peak-to-average ratio cannot exceed ~33% average occupancy without batch work to fill the troughs.
Power and energy#
DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION is a counter in millijoules (lesson 02), so:
# Average GPU power in watts
sum(rate(DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION[5m])) / 1000
# Output tokens per joule (GPU only)
sum(rate(vllm:generation_tokens_total[5m]))
/ (sum(rate(DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION[5m])) / 1000)GPU energy is not the whole bill. Scale it up to the wall:
facility energy = GPU energy × (server overhead: CPUs, fans, NICs, PSU loss ~1.2-1.5×)
× PUE (cooling and power distribution ~1.1-1.5×)PUE (power usage effectiveness) is total facility power ÷ IT power; modern liquid-cooled sites report values near 1.1, older air-cooled ones 1.4 or more.
Why power is now a first-class metric:
- It is the binding constraint. A site has a fixed number of megawatts. Tokens per watt decides how much revenue fits in the building; vendors now headline tokens per megawatt.
- Racks are dense. Current flagship racks draw well over 100 kW. Power capping and throttling are routine, and a capped GPU silently produces fewer tokens (lesson 02).
- It is an honest load signal. Power tracks real work far better than the utilization gauge.
- Reporting. Energy and carbon per request are increasingly asked for by customers and regulators.
An idle GPU still draws a meaningful fraction of its peak, so tokens per joule collapses at low occupancy — the same lever as cost.
Attribution#
Showback needs each unit of cost tied to an owner:
| Level | Method |
|---|---|
| Per model | GPU-hours of that model’s replicas |
| Per tenant | Their share of tokens (weighted input vs output), from gateway events |
| Per request | Estimated: prefill and decode seconds × the replica’s cost per second ÷ batch occupancy |
| Per agent task | Sum over the trace’s model spans (lesson 04) |
Hosted APIs make this easy — the bill is tokens × price — and the same fields
(gen_ai.usage.*) feed both. For self-hosted serving, publish an internal price per token per
model and revisit it when occupancy changes.
Unit economics dashboards#
A short list that belongs on one screen:
- $ per million output tokens, per model, against target.
- Occupancy and allocation ratio.
- Prefix-cache hit rate.
- Tokens per joule, and average power against the site’s cap.
- Spend by tenant, top ten.
- Idle GPU-hours this week.
The observability bill itself#
One GPU-hour costs more than storing a great deal of telemetry. Spending 1–2% of the infrastructure budget to recover a few points of occupancy pays for itself many times over. The usual mistake here is the opposite of web services: too little detail on expensive hardware.
Code#
From counters to unit cost and energy efficiency, and the effect of occupancy.
// unitcost.go — $ per million tokens and tokens per joule from counters; the occupancy lever.
package main
import "fmt"
type Sample struct {
OutputTokens float64 // counter
EnergyMilliJ float64 // counter, per GPU summed
}
func main() {
const (
gpus = 8.0
pricePerHour = 3.20 // $ per GPU-hour, illustrative
serverFactor = 1.3 // CPUs, fans, power supply losses
pue = 1.2
capacityTokPS = 2500.0 // output tokens/s per GPU at the SLO, from a load test
idleWatts = 120.0
peakWatts = 650.0
)
fmt.Println("occupancy tokens/s $ / M output tokens GPU W each tokens/J (GPU) tokens/J (facility)")
for _, occ := range []float64{0.1, 0.25, 0.5, 0.75, 0.95} {
// One hour of counters at this occupancy.
tps := gpus * capacityTokPS * occ
watts := idleWatts + (peakWatts-idleWatts)*occ
start := Sample{0, 0}
end := Sample{tps * 3600, gpus * watts * 3600 * 1000}
tokens := end.OutputTokens - start.OutputTokens
joules := (end.EnergyMilliJ - start.EnergyMilliJ) / 1000
cost := gpus * pricePerHour / tokens * 1e6
fmt.Printf("%8.0f%% %8.0f %19.2f %11.0f %15.2f %19.2f\n",
occ*100, tps, cost, watts, tokens/joules, tokens/(joules*serverFactor*pue))
}
fmt.Println("\nThe GPU-hour price never changed. Occupancy moved the unit cost almost 10x,")
fmt.Println("and the energy per token by about 2.6x because an idle GPU still draws power.")
}Remember this#
- $ per million tokens = allocated GPU cost ÷ tokens served. Tokens per joule = tokens ÷ energy.
- Both are divisions of counters you already export.
- The efficiency ladder — allocation, availability, occupancy, useful work, goodput, cache — locates the loss. Occupancy is usually the largest.
- Power is the binding constraint of modern sites and an honest load signal.
- Under-observing expensive hardware is the common mistake.
Try it#
- Run
unitcost.gowith a GPU price and capacity you can justify. At what occupancy do you match a hosted API’s price for a comparable model? - Write the PromQL for idle GPU-hours per week per team.
- Design a per-request cost estimate using only fields from a wide event.
Check yourself#
- Why is cost computed from allocated GPU-hours and not used ones?
- Name the six ratios of the efficiency ladder.
- Why does tokens per joule fall at low occupancy?