The idea in one minute#
A GPU chip is a grid of identical blocks called streaming multiprocessors (SMs). Each SM runs many threads, but not one at a time: it groups them into bundles of 32 called warps and gives all 32 the same instruction at the same moment, each on its own data.
This scheme is SIMT — single instruction, multiple threads. It is how one small control unit can drive a lot of arithmetic, and it is the source of the GPU’s biggest rule: threads in a warp should all want to do the same thing.
An analogy#
A rowing boat with 32 rowers and one cox calling the stroke. Everyone pulls together on the call, so the boat is fast and needs only one person giving orders.
If half the crew needs to do something different, the cox must call two strokes: first “left side, row” (right side waits), then “right side, row” (left side waits). The boat still works, but for those strokes half the crew is idle.
A picture#
flowchart TB
subgraph CHIP["GPU chip (H100: 132 SMs)"]
direction LR
subgraph SM["One SM"]
direction TB
SCHED["Warp schedulers<br/>pick a ready warp each cycle"]
subgraph W["A warp = 32 threads"]
direction LR
T0["t0"] --- T1["t1"] --- T2["t2"] --- T3["... t31"]
end
EX["Execution units<br/>FP32, INT, tensor cores"]
REG[("Registers")]
SHM[("Shared memory<br/>and L1 cache")]
SCHED --> W --> EX
EX --- REG
EX --- SHM
end
SM2["SM"]
SM3["SM"]
SM4["... x132"]
end
class SCHED neutral
class T0,T1,T2,T3,EX compute
class REG,SHM memory
class SM2,SM3,SM4 neutralHow it really works#
The hierarchy of names#
| Name | What it is | Size on an H100 |
|---|---|---|
| Thread | One instance of your kernel, working on one element | up to ~2,000 resident per SM |
| Warp | 32 threads that execute in lockstep | up to 64 resident per SM |
| SM | A block with its own schedulers, execution units, registers, shared memory | 132 per chip |
| Chip | All SMs plus a shared L2 cache and the memory controllers | 1 |
“Resident” means loaded and ready, not executing this instant. An SM holds many more warps than it can execute at once, and that is deliberate.
Why hold more warps than you can run?#
Reading GPU memory takes hundreds of clock cycles. A CPU hides such waits with caches and prediction. A GPU hides them with numbers: when a warp asks for memory and must wait, the scheduler instantly switches to another warp that is ready. Switching costs nothing, because each warp keeps its own registers — there is no state to save.
So an SM stays busy as long as some warp is ready. Keeping enough warps resident for this to work is called occupancy (module IV).
SIMT and the cost of branches#
All 32 threads of a warp share one program counter. When they reach an if and disagree:
if x[i] > 0 { // 20 threads say yes, 12 say no
A() // step 1: the 20 run A, the 12 are masked off (idle)
} else {
B() // step 2: the 12 run B, the 20 are masked off (idle)
}Both sides are executed, one after the other, with part of the warp idle each time. This is branch divergence. Correctness is unaffected — each thread gets the right answer — but the warp took the time of A plus B.
Divergence is only a problem within a warp. Two different warps taking different branches cost nothing extra.
SIMT vs SIMD#
CPUs have a similar idea, SIMD (single instruction, multiple data): one instruction operates on a vector of 8 or 16 numbers. The difference is who does the work:
| CPU SIMD | GPU SIMT | |
|---|---|---|
| You write | Vector code (or hope the compiler does) | Ordinary scalar code for one thread |
| Branches | You must convert to masks by hand | Hardware masks threads automatically |
| Width | 8–16 lanes | 32 lanes per warp, thousands of warps |
SIMT’s gift is that you write a normal-looking function for one element, and the hardware runs it in lockstep groups for you.
Code#
A warp simulator. Each “instruction” costs one step for the whole warp; a branch where threads disagree costs both sides.
// warp.go — SIMT in miniature: 32 threads, one program counter, masks on divergence.
package main
import (
"fmt"
"math/rand"
)
const warpSize = 32
// run executes "if data[i] > 0 { 4 steps } else { 6 steps }" for one warp
// and returns how many steps the warp as a whole needed.
func run(data [warpSize]int) (steps int) {
var takeThen, takeElse int
for _, v := range data {
if v > 0 {
takeThen++
} else {
takeElse++
}
}
steps++ // evaluate the condition: everyone together
if takeThen > 0 {
steps += 4 // "then" side runs; the others are masked
}
if takeElse > 0 {
steps += 6 // "else" side runs; the others are masked
}
return steps
}
func main() {
var sorted, mixed [warpSize]int
for i := range sorted {
sorted[i] = 1 // every thread takes the same branch
mixed[i] = rand.Intn(2)*2 - 1
}
fmt.Println("uniform warp: ", run(sorted), "steps") // 5
fmt.Println("divergent warp:", run(mixed), "steps") // 11
}The same data, grouped so that each warp agrees with itself, runs in less than half the steps. Real GPU code uses exactly this trick: sort or partition work so neighbours behave alike.
Remember this#
- Chip → SMs → warps → threads. A warp is 32 threads sharing one instruction stream.
- An SM hides slow memory by switching between many resident warps for free.
- Threads in a warp that disagree at a branch make the warp execute both sides.
- You write code for one thread; SIMT runs it in lockstep groups.
Try it#
- In
warp.go, make 1 thread out of 32 take theelsebranch. How many steps? What does this say about “rare” branches in GPU code? - Extend the simulator to 1,000 warps over random data, then sort the data first. Compare total steps.
- An H100 has 132 SMs with up to 64 warps of 32 threads each. How many threads can be resident at once?
Check yourself#
- What is a warp, and why is it 32 threads rather than 1?
- Why can a GPU switch warps at no cost when a CPU context switch is expensive?
- Why does branch divergence hurt performance but not correctness?