PidokuInfra

Racks, Power and Cost

Expert Advanced 50 min Difficulty 3/5 Topic 03 of 04

Prerequisites II.05, VI.01

The idea in one minute#

At scale, the GPU chip is not the hard part. Electricity, cooling and networking are. A modern AI rack draws as much power as a small neighbourhood, must be liquid-cooled, and is worthless without the network that ties racks together. All of that ends up in one number you can reason about: the cost of a GPU-hour, and from it, the cost per token.

A picture#

flowchart TB
  GRID["Power grid"] --> SUB["Substation and UPS"]
  SUB --> PDU["Rack power<br/>40 to 130 kW or more"]
  PDU --> RACK["Rack<br/>GPU servers, NVLink switches"]
  RACK -->|"heat"| CDU["Coolant loop<br/>cold plates on chips"]
  CDU --> CHILL["Chillers or cooling towers"]
  RACK <--> LEAF["Leaf switches"] <--> SPINE["Spine switches"]
  SPINE <--> OTHER["Other racks"]
  class GRID,SUB,PDU warn
  class RACK compute
  class CDU,CHILL memory
  class LEAF,SPINE,OTHER io

How it really works#

Power density#

Rack typePower
Traditional CPU rack5–15 kW
Air-cooled GPU rack (H100 servers)30–45 kW
Liquid-cooled Blackwell-generation rack (72 GPUs)~120–140 kW
Liquid-cooled Rubin-generation rack (72 GPUs)roughly 200 kW or more (reported)
Announced next-generation racks~600 kW

Air stops being practical around 40 kW per rack: you cannot push enough of it through the chassis. Above that, direct liquid cooling pipes coolant to cold plates sitting on the GPUs and CPUs. This is now the baseline for flagship systems, and it means AI capacity cannot simply be dropped into an existing data-center hall.

PUE (power usage effectiveness) is total facility power divided by the power reaching the computers. A PUE of 1.2 means 20% overhead for cooling and conversion. Liquid cooling lowers it.

The binding constraint for new AI capacity is frequently the grid connection: getting hundreds of megawatts delivered takes years.

The networks#

A GPU cluster has more than one network:

  • Scale-up — NVLink within a server or rack: the GPUs act as one machine.
  • Scale-out (back-end) — InfiniBand or Ethernet with RDMA between racks, for multi-node training and distributed inference. Built as a fat tree (leaf and spine switches) so any node can reach any other at full speed.
  • Front-end — ordinary Ethernet for user traffic, storage and management.

The network is a large share of cluster cost — commonly 10–20% — and its optics and switches draw significant power.

What a GPU-hour costs#

Owning:

cost per GPU-hour = (capital ÷ useful life in hours) + power + facility + operations
                    ─────────────────────────────────────────────────────────────────
                                          utilization

The division by utilization is the part people forget. A GPU that is busy 40% of the time costs 2.5x more per useful hour than its sticker rate. Idle GPUs are the most expensive thing in an AI company.

Renting moves the same costs into a price: on-demand (flexible, highest), reserved (commit for 1–3 years, substantially cheaper), spot (cheapest, can be taken away).

From GPU-hours to tokens#

cost per million tokens = (cost per GPU-hour ÷ tokens per hour) × 1,000,000

Tokens per hour depends on everything in modules IV and V: batching, the regime, quantization, how full the KV cache is. This is why software efficiency translates so directly into money: a 3x throughput improvement is a 3x cut in cost per token, on hardware you already have.

Code#

Go
// tco.go — what a GPU-hour and a million tokens really cost.
package main

import "fmt"

func main() {
	const (
		gpuCapex        = 30000.0 // per GPU, including its share of server, network, storage
		lifeYears       = 4.0
		gpuWatts        = 700.0
		serverOverhead  = 1.45 // CPUs, fans, network per GPU
		pue             = 1.25
		dollarsPerKWh   = 0.08
		facilityPerYear = 2500.0 // space, cooling plant, staff, per GPU
	)
	hours := lifeYears * 365 * 24
	capexPerHour := gpuCapex / hours
	powerPerHour := gpuWatts * serverOverhead * pue / 1000 * dollarsPerKWh
	facilityPerHour := facilityPerYear / (365 * 24)
	raw := capexPerHour + powerPerHour + facilityPerHour

	fmt.Printf("capital  $%.2f/h   power $%.2f/h   facility $%.2f/h   total $%.2f/h at 100%% use\n\n",
		capexPerHour, powerPerHour, facilityPerHour, raw)

	fmt.Println("utilization   $/useful GPU-hour   $/M tokens at 2,500 tok/s   at 7,500 tok/s")
	for _, u := range []float64{1.0, 0.7, 0.4, 0.2} {
		perHour := raw / u
		fmt.Printf("%10.0f%%   %17.2f   %24.3f   %14.3f\n",
			u*100, perHour, perHour/(2500*3600)*1e6, perHour/(7500*3600)*1e6)
	}
}

Two levers dominate the output: utilization (down the rows) and throughput (across the columns). Electricity, despite the headlines, is the smallest term.

Remember this#

  • Flagship GPU racks exceed 100 kW and require liquid cooling.
  • Clusters have a scale-up network (NVLink), a scale-out network (RDMA), and a front-end network.
  • Cost per useful GPU-hour = total cost ÷ utilization. Idle GPUs are the biggest waste.
  • Cost per token = GPU-hour cost ÷ throughput. Software efficiency is money.

Try it#

  1. Run tco.go with a GPU price and electricity rate you can find today.
  2. At what utilization does owning beat renting at $2.50 per GPU-hour?
  3. A team doubles throughput through batching but p99 latency rises 30%. Write the argument for and against shipping it.

Check yourself#

  1. Why does liquid cooling become necessary?
  2. What are the three networks in a GPU cluster?
  3. Why does utilization appear in the denominator of cost?

↑↓ navigate↵ openesc close