PidokuInfra

Cost, Power and Efficiency

Advanced 50 min Difficulty 3/5 Topic 06 of 07

Prerequisites 02, 03, 05; helpful: Inference Engineering XI.03, GPU Engineering VII.03

The idea in one minute#

For AI infrastructure, cost is not an accounting afterthought; it is an engineering metric to graph next to latency. Two figures summarize a serving system: dollars per million tokens and tokens per joule. Both are simple divisions of counters you already have — GPU-hours and energy on top, tokens on the bottom.

The gap between what a GPU could produce and what it does produce is where the money goes: idle capacity, low batch occupancy, cache misses, and wasted work.

An analogy#

A delivery van. The lease and the driver cost the same whether the van is full or empty, so the cost per parcel is decided almost entirely by how full it runs. Fuel per parcel is a second, separate efficiency. A fleet manager tracks both, per route.

A picture#

flowchart TB
  subgraph COST["What you pay for"]
    GH["GPU-hours<br/>allocated, not used"]
    EN["Energy<br/>joules from the GPU counter"]
  end
  subgraph WORK["What you got"]
    TOK["Tokens served<br/>input, cached, output"]
    GOOD["Good tokens<br/>from requests that met the SLO"]
  end
  GH --> D1["$ per million tokens"]
  TOK --> D1
  EN --> D2["tokens per joule"]
  TOK --> D2
  D1 --> LOSS["Where the gap goes:<br/>idle, low occupancy, cache misses,<br/>preempted and abandoned work"]
  D2 --> LOSS
  class GH,EN neutral
  class TOK,GOOD memory
  class D1,D2 compute
  class LOSS warn

How it really works#

Dollars per million tokens#

cost per M tokens = (GPUs × $ per GPU-hour) ÷ (tokens per hour) × 1,000,000
  • The numerator is what you allocated, including idle time. An on-demand GPU and a reserved GPU sitting idle cost the same per hour.
  • The denominator comes from the engine’s token counters (lesson 03) or the gateway’s.
  • Split input and output: output tokens cost several times more to produce. A common practice is to apportion GPU time by measured phase time (prefill seconds vs decode seconds).
  • Add the non-GPU share — CPU hosts, storage, network, the gateway — or state that you left it out.

As PromQL, with a recording rule holding your price per GPU-hour by node pool:

PromQL
  sum by (model) (gpu_allocated * on (node_pool) group_left gpu_hourly_price_dollars)
/
  (sum by (model) (rate(vllm:generation_tokens_total[1h])) * 3600)
* 1e6

The efficiency ladder#

Each step is a ratio you can measure; multiplied together they explain the cost.

RatioDefinitionLost to
AllocationGPUs allocated ÷ GPUs ownedUnscheduled capacity, fragmentation
AvailabilityTime ready to serve ÷ time allocatedCold starts, loading, restarts, faults
OccupancyTokens served ÷ tokens the replica could serve at SLOUnder-filled batches, over-provisioning for peaks
Useful workTokens delivered ÷ tokens computedPreempted work recomputed, cancelled streams, retries
GoodputTokens in SLO-meeting requests ÷ tokens deliveredOverload
Cache leverageCached prompt tokens ÷ prompt tokensPoor routing, unstable prompts

Occupancy is usually the largest loss. A fleet sized for peak traffic with a 3:1 peak-to-average ratio cannot exceed ~33% average occupancy without batch work to fill the troughs.

Power and energy#

DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION is a counter in millijoules (lesson 02), so:

PromQL
# Average GPU power in watts
sum(rate(DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION[5m])) / 1000

# Output tokens per joule (GPU only)
sum(rate(vllm:generation_tokens_total[5m]))
/ (sum(rate(DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION[5m])) / 1000)

GPU energy is not the whole bill. Scale it up to the wall:

facility energy = GPU energy × (server overhead: CPUs, fans, NICs, PSU loss  ~1.2-1.5×)
                             × PUE (cooling and power distribution           ~1.1-1.5×)

PUE (power usage effectiveness) is total facility power ÷ IT power; modern liquid-cooled sites report values near 1.1, older air-cooled ones 1.4 or more.

Why power is now a first-class metric:

  • It is the binding constraint. A site has a fixed number of megawatts. Tokens per watt decides how much revenue fits in the building; vendors now headline tokens per megawatt.
  • Racks are dense. Current flagship racks draw well over 100 kW. Power capping and throttling are routine, and a capped GPU silently produces fewer tokens (lesson 02).
  • It is an honest load signal. Power tracks real work far better than the utilization gauge.
  • Reporting. Energy and carbon per request are increasingly asked for by customers and regulators.

An idle GPU still draws a meaningful fraction of its peak, so tokens per joule collapses at low occupancy — the same lever as cost.

Attribution#

Showback needs each unit of cost tied to an owner:

LevelMethod
Per modelGPU-hours of that model’s replicas
Per tenantTheir share of tokens (weighted input vs output), from gateway events
Per requestEstimated: prefill and decode seconds × the replica’s cost per second ÷ batch occupancy
Per agent taskSum over the trace’s model spans (lesson 04)

Hosted APIs make this easy — the bill is tokens × price — and the same fields (gen_ai.usage.*) feed both. For self-hosted serving, publish an internal price per token per model and revisit it when occupancy changes.

Unit economics dashboards#

A short list that belongs on one screen:

  1. $ per million output tokens, per model, against target.
  2. Occupancy and allocation ratio.
  3. Prefix-cache hit rate.
  4. Tokens per joule, and average power against the site’s cap.
  5. Spend by tenant, top ten.
  6. Idle GPU-hours this week.

The observability bill itself#

One GPU-hour costs more than storing a great deal of telemetry. Spending 1–2% of the infrastructure budget to recover a few points of occupancy pays for itself many times over. The usual mistake here is the opposite of web services: too little detail on expensive hardware.

Code#

From counters to unit cost and energy efficiency, and the effect of occupancy.

Go
// unitcost.go — $ per million tokens and tokens per joule from counters; the occupancy lever.
package main

import "fmt"

type Sample struct {
	OutputTokens float64 // counter
	EnergyMilliJ float64 // counter, per GPU summed
}

func main() {
	const (
		gpus          = 8.0
		pricePerHour  = 3.20 // $ per GPU-hour, illustrative
		serverFactor  = 1.3  // CPUs, fans, power supply losses
		pue           = 1.2
		capacityTokPS = 2500.0 // output tokens/s per GPU at the SLO, from a load test
		idleWatts     = 120.0
		peakWatts     = 650.0
	)
	fmt.Println("occupancy  tokens/s   $ / M output tokens   GPU W each   tokens/J (GPU)   tokens/J (facility)")
	for _, occ := range []float64{0.1, 0.25, 0.5, 0.75, 0.95} {
		// One hour of counters at this occupancy.
		tps := gpus * capacityTokPS * occ
		watts := idleWatts + (peakWatts-idleWatts)*occ
		start := Sample{0, 0}
		end := Sample{tps * 3600, gpus * watts * 3600 * 1000}

		tokens := end.OutputTokens - start.OutputTokens
		joules := (end.EnergyMilliJ - start.EnergyMilliJ) / 1000
		cost := gpus * pricePerHour / tokens * 1e6
		fmt.Printf("%8.0f%%  %8.0f  %19.2f  %11.0f  %15.2f  %19.2f\n",
			occ*100, tps, cost, watts, tokens/joules, tokens/(joules*serverFactor*pue))
	}
	fmt.Println("\nThe GPU-hour price never changed. Occupancy moved the unit cost almost 10x,")
	fmt.Println("and the energy per token by about 2.6x because an idle GPU still draws power.")
}

Remember this#

  • $ per million tokens = allocated GPU cost ÷ tokens served. Tokens per joule = tokens ÷ energy.
  • Both are divisions of counters you already export.
  • The efficiency ladder — allocation, availability, occupancy, useful work, goodput, cache — locates the loss. Occupancy is usually the largest.
  • Power is the binding constraint of modern sites and an honest load signal.
  • Under-observing expensive hardware is the common mistake.

Try it#

  1. Run unitcost.go with a GPU price and capacity you can justify. At what occupancy do you match a hosted API’s price for a comparable model?
  2. Write the PromQL for idle GPU-hours per week per team.
  3. Design a per-request cost estimate using only fields from a wide event.

Check yourself#

  1. Why is cost computed from allocated GPU-hours and not used ones?
  2. Name the six ratios of the efficiency ladder.
  3. Why does tokens per joule fall at low occupancy?

↑↓ navigate↵ openesc close