The idea in one minute#
Data-center GPUs use a special memory called HBM (high-bandwidth memory): memory chips stacked vertically and placed right next to the GPU, connected by thousands of microscopic wires. The goal is not capacity. It is bandwidth — how many bytes per second can flow between memory and the arithmetic units.
For large AI models, bandwidth is frequently the number that sets the speed. A GPU that can do a trillion multiplications per second is still stuck if the numbers to multiply arrive slowly.
An analogy#
A kitchen with fifty chefs and one narrow door to the pantry. Hiring more chefs does nothing; they queue at the door. Widening the door is the only thing that helps.
HBM is a very wide door.
A picture#
flowchart LR
subgraph PKG["One GPU package"]
direction LR
subgraph STACK["HBM stack (several around the chip)"]
direction TB
D1["DRAM layer"] --- D2["DRAM layer"] --- D3["DRAM layer"] --- D4["... 8 to 16 layers"]
end
INT["Silicon interposer<br/>thousands of short wires"]
DIE["GPU die"]
STACK --- INT --- DIE
end
class D1,D2,D3,D4 memory
class INT neutral
class DIE computeHow it really works#
What bandwidth means#
bandwidth = (number of data wires) × (bits per second on each wire)Ordinary memory modules sit centimetres from the CPU on a circuit board; only a few hundred wires fit, and long wires cannot be driven very fast. HBM attacks the first term: by stacking memory dies and connecting them with through-silicon vias (vertical wires through the chips), then mounting the stack on a silicon interposer beside the GPU, it gets over a thousand data wires per stack, each only millimetres long.
The generations#
| Memory | Used in | Bandwidth per GPU |
|---|---|---|
| GDDR6 / GDDR6X (not stacked) | T4, L4, RTX 4090 | 0.3–1.0 TB/s |
| HBM2e | A100 | ~2.0 TB/s |
| HBM3 | H100 | ~3.35 TB/s |
| HBM3e | H200, B200, B300 | ~4.8–8 TB/s |
| HBM4 | Rubin (shipping since August 2026); AMD MI455X | ~22 TB/s (vendor figures) |
Consumer cards use GDDR: fast, cheap, conventional chips around the GPU. It is why a gaming card with strong arithmetic still has a fraction of a data-center card’s bandwidth.
HBM is also the scarce part. It is difficult to manufacture, made by three companies, and for several years its supply — not the GPU chips themselves — has limited how many AI GPUs exist.
Why AI cares so much#
Generating one token with a language model means reading every weight once. For a 16 GB model:
time per token ≥ 16 GB ÷ bandwidth
H100: 16 ÷ 3350 GB/s = 4.8 ms → at most ~210 tokens/s for one sequence
L4: 16 ÷ 300 GB/s = 53 ms → at most ~19 tokens/sThe arithmetic for that token would take the H100 well under a millisecond. The GPU spends most of the step waiting for bytes. Work like this is called memory-bound.
The cure is to make each byte do more work — process many sequences per pass over the weights (batching), or store each weight in fewer bytes (quantization). Module V covers both.
Capacity vs bandwidth#
They are different numbers and they fail differently:
- Too little capacity → the program does not run (out of memory).
- Too little bandwidth → the program runs, slowly, with the GPU reporting “100% utilization”.
Code#
Measure your own machine’s memory bandwidth. The result will be far below a GPU’s, and the experiment shows what “memory-bound” feels like: adding goroutines stops helping long before you run out of cores.
// bandwidth.go — measure how fast this machine can stream bytes through memory.
package main
import (
"fmt"
"runtime"
"sync"
"time"
)
func main() {
const n = 1 << 28 // 256M float32 = 1 GiB per array
src, dst := make([]float32, n), make([]float32, n)
for i := range src {
src[i] = 1
}
for workers := 1; workers <= runtime.NumCPU(); workers *= 2 {
chunk := n / workers
t0 := time.Now()
var wg sync.WaitGroup
for w := 0; w < workers; w++ {
a, b := src[w*chunk:(w+1)*chunk], dst[w*chunk:(w+1)*chunk]
wg.Add(1)
go func() {
defer wg.Done()
copy(b, a) // read 4 bytes, write 4 bytes, no arithmetic at all
}()
}
wg.Wait()
dt := time.Since(t0).Seconds()
fmt.Printf("%2d workers: %6.1f GB/s\n", workers, 2*4*float64(n)/dt/1e9)
}
}Typical laptop: 15–30 GB/s with one worker, flattening at 30–100 GB/s no matter how many you add. That ceiling is your memory bandwidth. An H100’s is 3,350.
Remember this#
- HBM = stacked memory beside the GPU, built for bytes per second, not capacity.
- Bandwidth = wires × speed per wire. HBM wins by having far more, far shorter wires.
- Reading every weight once per token makes LLM generation memory-bound at small batch sizes.
- Capacity failures crash; bandwidth shortages just make things slow.
Try it#
- Run
bandwidth.go. At what worker count does the number stop improving? Why does it stop? - With your measured bandwidth, what is the upper limit on tokens/s for a 4 GB model on your CPU?
- A 70 GB model on a GPU with 8 TB/s: what is the maximum tokens/s for one sequence?
Check yourself#
- What physical trick gives HBM its bandwidth?
- Why does a gaming GPU with similar FLOP/s serve LLMs more slowly than a data-center GPU?
- What does “memory-bound” mean?