The idea in one minute#
At scale, the GPU chip is not the hard part. Electricity, cooling and networking are. A modern AI rack draws as much power as a small neighbourhood, must be liquid-cooled, and is worthless without the network that ties racks together. All of that ends up in one number you can reason about: the cost of a GPU-hour, and from it, the cost per token.
A picture#
flowchart TB GRID["Power grid"] --> SUB["Substation and UPS"] SUB --> PDU["Rack power<br/>40 to 130 kW or more"] PDU --> RACK["Rack<br/>GPU servers, NVLink switches"] RACK -->|"heat"| CDU["Coolant loop<br/>cold plates on chips"] CDU --> CHILL["Chillers or cooling towers"] RACK <--> LEAF["Leaf switches"] <--> SPINE["Spine switches"] SPINE <--> OTHER["Other racks"] class GRID,SUB,PDU warn class RACK compute class CDU,CHILL memory class LEAF,SPINE,OTHER io
How it really works#
Power density#
| Rack type | Power |
|---|---|
| Traditional CPU rack | 5–15 kW |
| Air-cooled GPU rack (H100 servers) | 30–45 kW |
| Liquid-cooled Blackwell-generation rack (72 GPUs) | ~120–140 kW |
| Liquid-cooled Rubin-generation rack (72 GPUs) | roughly 200 kW or more (reported) |
| Announced next-generation racks | ~600 kW |
Air stops being practical around 40 kW per rack: you cannot push enough of it through the chassis. Above that, direct liquid cooling pipes coolant to cold plates sitting on the GPUs and CPUs. This is now the baseline for flagship systems, and it means AI capacity cannot simply be dropped into an existing data-center hall.
PUE (power usage effectiveness) is total facility power divided by the power reaching the computers. A PUE of 1.2 means 20% overhead for cooling and conversion. Liquid cooling lowers it.
The binding constraint for new AI capacity is frequently the grid connection: getting hundreds of megawatts delivered takes years.
The networks#
A GPU cluster has more than one network:
- Scale-up — NVLink within a server or rack: the GPUs act as one machine.
- Scale-out (back-end) — InfiniBand or Ethernet with RDMA between racks, for multi-node training and distributed inference. Built as a fat tree (leaf and spine switches) so any node can reach any other at full speed.
- Front-end — ordinary Ethernet for user traffic, storage and management.
The network is a large share of cluster cost — commonly 10–20% — and its optics and switches draw significant power.
What a GPU-hour costs#
Owning:
cost per GPU-hour = (capital ÷ useful life in hours) + power + facility + operations
─────────────────────────────────────────────────────────────────
utilizationThe division by utilization is the part people forget. A GPU that is busy 40% of the time costs 2.5x more per useful hour than its sticker rate. Idle GPUs are the most expensive thing in an AI company.
Renting moves the same costs into a price: on-demand (flexible, highest), reserved (commit for 1–3 years, substantially cheaper), spot (cheapest, can be taken away).
From GPU-hours to tokens#
cost per million tokens = (cost per GPU-hour ÷ tokens per hour) × 1,000,000Tokens per hour depends on everything in modules IV and V: batching, the regime, quantization, how full the KV cache is. This is why software efficiency translates so directly into money: a 3x throughput improvement is a 3x cut in cost per token, on hardware you already have.
Code#
// tco.go — what a GPU-hour and a million tokens really cost.
package main
import "fmt"
func main() {
const (
gpuCapex = 30000.0 // per GPU, including its share of server, network, storage
lifeYears = 4.0
gpuWatts = 700.0
serverOverhead = 1.45 // CPUs, fans, network per GPU
pue = 1.25
dollarsPerKWh = 0.08
facilityPerYear = 2500.0 // space, cooling plant, staff, per GPU
)
hours := lifeYears * 365 * 24
capexPerHour := gpuCapex / hours
powerPerHour := gpuWatts * serverOverhead * pue / 1000 * dollarsPerKWh
facilityPerHour := facilityPerYear / (365 * 24)
raw := capexPerHour + powerPerHour + facilityPerHour
fmt.Printf("capital $%.2f/h power $%.2f/h facility $%.2f/h total $%.2f/h at 100%% use\n\n",
capexPerHour, powerPerHour, facilityPerHour, raw)
fmt.Println("utilization $/useful GPU-hour $/M tokens at 2,500 tok/s at 7,500 tok/s")
for _, u := range []float64{1.0, 0.7, 0.4, 0.2} {
perHour := raw / u
fmt.Printf("%10.0f%% %17.2f %24.3f %14.3f\n",
u*100, perHour, perHour/(2500*3600)*1e6, perHour/(7500*3600)*1e6)
}
}Two levers dominate the output: utilization (down the rows) and throughput (across the columns). Electricity, despite the headlines, is the smallest term.
Remember this#
- Flagship GPU racks exceed 100 kW and require liquid cooling.
- Clusters have a scale-up network (NVLink), a scale-out network (RDMA), and a front-end network.
- Cost per useful GPU-hour = total cost ÷ utilization. Idle GPUs are the biggest waste.
- Cost per token = GPU-hour cost ÷ throughput. Software efficiency is money.
Try it#
- Run
tco.gowith a GPU price and electricity rate you can find today. - At what utilization does owning beat renting at $2.50 per GPU-hour?
- A team doubles throughput through batching but p99 latency rises 30%. Write the argument for and against shipping it.
Check yourself#
- Why does liquid cooling become necessary?
- What are the three networks in a GPU cluster?
- Why does utilization appear in the denominator of cost?