The idea in one minute#
You give a GPU a kernel — a function — and ask for it to be run by a huge number of threads. Every thread runs the same code; the only difference between them is an index that says “which one am I?”. Each thread uses its index to pick the piece of data it is responsible for.
Threads are organized in two levels: threads are grouped into blocks, and blocks form a grid. A block’s threads run on the same SM and can cooperate through shared memory. Blocks are independent of each other.
An analogy#
A stadium of exam markers. Everyone gets the same marking instructions (the kernel). Each person’s seat — row and seat number — tells them which exam paper is theirs. A row (a block) shares one desk where they can pass notes (shared memory). Rows do not talk to each other.
A picture#
flowchart TB
subgraph GRID["Grid: one kernel launch"]
direction LR
subgraph B0["Block 0"]
direction LR
a0["thread 0"] --- a1["thread 1"] --- a2["..."] --- a3["thread 255"]
end
subgraph B1["Block 1"]
direction LR
b0["thread 0"] --- b1["thread 1"] --- b2["..."] --- b3["thread 255"]
end
BN["... more blocks"]
end
DATA[("Array of N elements")]
B0 -->|"elements 0..255"| DATA
B1 -->|"elements 256..511"| DATA
class a0,a1,a2,a3,b0,b1,b2,b3 compute
class DATA memory
class BN neutralHow it really works#
The index calculation#
Inside a kernel, each thread can read:
threadIdx— its position within its blockblockIdx— its block’s position within the gridblockDim— threads per block
and computes its global index:
i = blockIdx * blockDim + threadIdxThen it works on element i. That single line appears, in some form, in every CUDA kernel ever
written.
Mapping to hardware#
| Software | Hardware |
|---|---|
| Thread | One lane in a warp |
| 32 consecutive threads of a block | One warp |
| Block | Runs entirely on one SM; may share that SM with other blocks |
| Grid | Spread across all SMs by the scheduler |
A block is limited to 1,024 threads. The grid can have billions.
Choosing sizes#
You choose the block size; the number of blocks follows:
blocks = ceil(N / blockSize)Because the last block may run past the end of the data, every kernel starts with a bounds
check: if (i >= n) return;.
Typical block sizes are 128, 256 or 512 — multiples of the warp size (32), so no warp is partly empty. The exact choice affects occupancy (IV.03) but rarely by much; 256 is a sensible default.
Blocks and grids can be 2D or 3D#
For images or matrices it is convenient to index threads as (x, y). CUDA lets blocks and
grids have up to three dimensions. This is only a convenience for the index arithmetic; the
hardware does not care.
What threads may and may not do#
- Threads in the same block can share data through shared memory and wait for each other
with a barrier (
__syncthreads()). - Threads in different blocks cannot cooperate directly, and there is no guaranteed order in which blocks run.
- Two threads writing the same global memory location is a data race, exactly as in Go.
Go programmers have a head start here: a kernel launch is “start N goroutines over a slice, no channels, results written to disjoint indices”.
Code#
The CUDA execution model, rebuilt in Go. Launch plays the GPU: it runs the kernel once per
thread, passing each its coordinates.
// grid.go — the CUDA thread hierarchy, simulated with goroutines.
package main
import (
"fmt"
"sync"
)
// Ctx is what a thread knows about itself — the CUDA built-ins.
type Ctx struct {
ThreadIdx, BlockIdx, BlockDim int
}
// Launch runs kernel on a grid of `blocks` blocks of `blockDim` threads each.
// One goroutine per block stands in for one SM running that block.
func Launch(blocks, blockDim int, kernel func(Ctx)) {
var wg sync.WaitGroup
for b := 0; b < blocks; b++ {
wg.Add(1)
go func() {
defer wg.Done()
for t := 0; t < blockDim; t++ {
kernel(Ctx{ThreadIdx: t, BlockIdx: b, BlockDim: blockDim})
}
}()
}
wg.Wait() // like waiting for the GPU to finish
}
func main() {
const n = 1000
x, y, out := make([]float32, n), make([]float32, n), make([]float32, n)
for i := range x {
x[i], y[i] = float32(i), 2
}
// The kernel: out = a*x + y ("SAXPY"), written for ONE thread.
const a = 3
saxpy := func(c Ctx) {
i := c.BlockIdx*c.BlockDim + c.ThreadIdx // which element am I?
if i >= n { // the last block overshoots
return
}
out[i] = a*x[i] + y[i]
}
const blockDim = 256
blocks := (n + blockDim - 1) / blockDim // ceil(n / blockDim) = 4
Launch(blocks, blockDim, saxpy)
fmt.Println(blocks, "blocks;", out[0], out[1], out[999]) // 4 blocks; 2 5 2999
}No mutex is needed because each thread writes only out[i] for its own i. That is the
discipline of GPU programming in one sentence.
Remember this#
- A kernel is run by many threads; each finds its work from its index.
i = blockIdx * blockDim + threadIdx, followed by a bounds check.- Blocks ≤ 1,024 threads; threads in a block share an SM and shared memory.
- Blocks are independent and unordered.
Try it#
- Remove the
if i >= ncheck ingrid.go. What happens, and why would the same bug on a real GPU be worse than a panic? - Write a 2D version: a kernel with
(x, y)thread and block indices that inverts a 1920×1080 grayscale image. - Write a kernel where every thread adds its value to a single shared
sumvariable. Run it withgo run -race. How would you fix it? (GPU kernels fix it with a reduction.)
Check yourself#
- How does a thread know which element to process?
- Why does every kernel need a bounds check?
- Which threads can share memory directly, and which cannot?