The idea in one minute#
An SM hides memory waits by switching between many resident warps. Occupancy is how many warps a kernel actually keeps resident, as a fraction of the SM’s maximum. Each thread needs registers and each block may need shared memory; a kernel that is greedy with either leaves room for fewer warps, and the SM runs out of ready work.
Divergence is the other way to waste an SM: warps whose threads disagree at branches execute both paths (II.01). This lesson shows how to compute the first and reduce the second.
A picture#
flowchart LR
subgraph LOW["Low occupancy"]
direction TB
A1["warp: waiting on memory"]
A2["warp: waiting on memory"]
A3["SM idle: nothing ready"]
end
subgraph HIGH["High occupancy"]
direction TB
B1["warp: waiting"]
B2["warp: waiting"]
B3["warp: waiting"]
B4["warp: ready, runs now"]
B5["warp: ready next"]
end
class A1,A2,B1,B2,B3 queue
class A3 warn
class B4,B5 computeHow it really works#
What limits resident warps#
An SM has fixed pools (H100 figures):
| Resource | Per SM | Consumed by |
|---|---|---|
| Warp slots | 64 | Every warp |
| Registers | 65,536 | registers per thread × threads |
| Shared memory | up to ~228 KB | shared memory per block × blocks |
| Block slots | 32 | Every block |
The number of blocks that fit is the minimum allowed by each resource. Occupancy is:
occupancy = resident warps ÷ 64Example: a kernel using 80 registers per thread, block size 256 (8 warps), no shared memory.
by registers: 65,536 ÷ (80 × 256) = 3 blocks → 24 warps
by warp slots: 64 ÷ 8 = 8 blocks → 64 warps
resident = 24 warps → occupancy 37.5%Registers are the limit. Trimming the kernel to 64 registers per thread gives 4 blocks, 32 warps, 50%.
Higher is not always better#
Occupancy exists to hide memory latency. A kernel that mostly waits on memory benefits from many warps. A kernel that is pure arithmetic on data already in registers is saturated by a few warps; raising occupancy does nothing. Past roughly 50%, gains are usually small.
Treat occupancy as a diagnostic: very low occupancy (under ~25%) in a memory-heavy kernel is a finding. Chasing 100% is not a goal.
Reducing divergence#
Three standard moves:
- Replace the branch with arithmetic.
max(x, 0)instead ofif x < 0 { x = 0 }. The compiler turns simple conditionals into a “select” instruction that every thread runs identically. - Group similar work. Sort or partition elements so each warp’s threads take the same path (the trick from II.01’s simulator).
- Move the decision outward. If a condition is the same for a whole batch, choose between two kernels on the host instead of branching inside one.
Loops with data-dependent trip counts are divergence too: a warp runs until its slowest thread finishes, with the finished ones idle.
Why neural networks are friendly#
Matrix multiplies, convolutions and activations do the same thing for every element: no data-dependent branches, no variable-length loops. That uniformity is a large part of why AI workloads fit GPUs so well. Divergence shows up in the parts around the model — sampling, tokenization, sparse or ragged data — which is why those often stay on the CPU.
Code#
An occupancy calculator. Real toolchains print the same numbers; computing them yourself makes the limits obvious.
// occupancy.go — how many warps can an SM keep resident for a given kernel?
package main
import "fmt"
type SM struct{ WarpSlots, Registers, SharedBytes, BlockSlots int }
type Kernel struct {
Name string
BlockThreads int
RegsPerThread int
SharedPerBlock int // bytes
}
func (sm SM) Occupancy(k Kernel) (pct float64, limiter string) {
warpsPerBlock := (k.BlockThreads + 31) / 32
blocks, limiter := sm.WarpSlots/warpsPerBlock, "warp slots"
if b := sm.Registers / (k.RegsPerThread * k.BlockThreads); b < blocks {
blocks, limiter = b, "registers"
}
if k.SharedPerBlock > 0 {
if b := sm.SharedBytes / k.SharedPerBlock; b < blocks {
blocks, limiter = b, "shared memory"
}
}
if sm.BlockSlots < blocks {
blocks, limiter = sm.BlockSlots, "block slots"
}
return 100 * float64(blocks*warpsPerBlock) / float64(sm.WarpSlots), limiter
}
func main() {
h100 := SM{WarpSlots: 64, Registers: 65536, SharedBytes: 228 * 1024, BlockSlots: 32}
for _, k := range []Kernel{
{"light elementwise", 256, 16, 0},
{"register-hungry", 256, 80, 0},
{"trimmed to 64 regs", 256, 64, 0},
{"big shared tile", 256, 32, 96 * 1024},
{"tiny blocks", 32, 16, 0},
} {
pct, lim := h100.Occupancy(k)
fmt.Printf("%-20s occupancy %5.1f%% limited by %s\n", k.Name, pct, lim)
}
}Remember this#
- Occupancy = resident warps ÷ maximum. Limited by registers, shared memory, or slot counts.
- It matters for kernels that wait on memory; it is irrelevant for pure-arithmetic kernels.
- Very low occupancy is a red flag; maximum occupancy is not a target.
- Cut divergence by using arithmetic instead of branches, grouping similar work, and deciding on the host.
Try it#
- Run
occupancy.go. Why is “tiny blocks” limited to 50%? What block size fixes it? - For the “big shared tile” kernel, what is the largest tile that still gives 50% occupancy?
- Rewrite
if x < 0 { y = 0 } else { y = x }andif a > b { m = a } else { m = b }without branches.
Check yourself#
- Name three resources that can limit occupancy.
- When does raising occupancy not help?
- Give two ways to reduce divergence.