PidokuInfra

Kernels, Threads, Blocks, Grids

Intermediate 1h Difficulty 3/5 Topic 01 of 05

Prerequisites II.01

The idea in one minute#

You give a GPU a kernel — a function — and ask for it to be run by a huge number of threads. Every thread runs the same code; the only difference between them is an index that says “which one am I?”. Each thread uses its index to pick the piece of data it is responsible for.

Threads are organized in two levels: threads are grouped into blocks, and blocks form a grid. A block’s threads run on the same SM and can cooperate through shared memory. Blocks are independent of each other.

An analogy#

A stadium of exam markers. Everyone gets the same marking instructions (the kernel). Each person’s seat — row and seat number — tells them which exam paper is theirs. A row (a block) shares one desk where they can pass notes (shared memory). Rows do not talk to each other.

A picture#

flowchart TB
  subgraph GRID["Grid: one kernel launch"]
    direction LR
    subgraph B0["Block 0"]
      direction LR
      a0["thread 0"] --- a1["thread 1"] --- a2["..."] --- a3["thread 255"]
    end
    subgraph B1["Block 1"]
      direction LR
      b0["thread 0"] --- b1["thread 1"] --- b2["..."] --- b3["thread 255"]
    end
    BN["... more blocks"]
  end
  DATA[("Array of N elements")]
  B0 -->|"elements 0..255"| DATA
  B1 -->|"elements 256..511"| DATA
  class a0,a1,a2,a3,b0,b1,b2,b3 compute
  class DATA memory
  class BN neutral

How it really works#

The index calculation#

Inside a kernel, each thread can read:

  • threadIdx — its position within its block
  • blockIdx — its block’s position within the grid
  • blockDim — threads per block

and computes its global index:

i = blockIdx * blockDim + threadIdx

Then it works on element i. That single line appears, in some form, in every CUDA kernel ever written.

Mapping to hardware#

SoftwareHardware
ThreadOne lane in a warp
32 consecutive threads of a blockOne warp
BlockRuns entirely on one SM; may share that SM with other blocks
GridSpread across all SMs by the scheduler

A block is limited to 1,024 threads. The grid can have billions.

Choosing sizes#

You choose the block size; the number of blocks follows:

blocks = ceil(N / blockSize)

Because the last block may run past the end of the data, every kernel starts with a bounds check: if (i >= n) return;.

Typical block sizes are 128, 256 or 512 — multiples of the warp size (32), so no warp is partly empty. The exact choice affects occupancy (IV.03) but rarely by much; 256 is a sensible default.

Blocks and grids can be 2D or 3D#

For images or matrices it is convenient to index threads as (x, y). CUDA lets blocks and grids have up to three dimensions. This is only a convenience for the index arithmetic; the hardware does not care.

What threads may and may not do#

  • Threads in the same block can share data through shared memory and wait for each other with a barrier (__syncthreads()).
  • Threads in different blocks cannot cooperate directly, and there is no guaranteed order in which blocks run.
  • Two threads writing the same global memory location is a data race, exactly as in Go.

Go programmers have a head start here: a kernel launch is “start N goroutines over a slice, no channels, results written to disjoint indices”.

Code#

The CUDA execution model, rebuilt in Go. Launch plays the GPU: it runs the kernel once per thread, passing each its coordinates.

Go
// grid.go — the CUDA thread hierarchy, simulated with goroutines.
package main

import (
	"fmt"
	"sync"
)

// Ctx is what a thread knows about itself — the CUDA built-ins.
type Ctx struct {
	ThreadIdx, BlockIdx, BlockDim int
}

// Launch runs kernel on a grid of `blocks` blocks of `blockDim` threads each.
// One goroutine per block stands in for one SM running that block.
func Launch(blocks, blockDim int, kernel func(Ctx)) {
	var wg sync.WaitGroup
	for b := 0; b < blocks; b++ {
		wg.Add(1)
		go func() {
			defer wg.Done()
			for t := 0; t < blockDim; t++ {
				kernel(Ctx{ThreadIdx: t, BlockIdx: b, BlockDim: blockDim})
			}
		}()
	}
	wg.Wait() // like waiting for the GPU to finish
}

func main() {
	const n = 1000
	x, y, out := make([]float32, n), make([]float32, n), make([]float32, n)
	for i := range x {
		x[i], y[i] = float32(i), 2
	}

	// The kernel: out = a*x + y ("SAXPY"), written for ONE thread.
	const a = 3
	saxpy := func(c Ctx) {
		i := c.BlockIdx*c.BlockDim + c.ThreadIdx // which element am I?
		if i >= n {                              // the last block overshoots
			return
		}
		out[i] = a*x[i] + y[i]
	}

	const blockDim = 256
	blocks := (n + blockDim - 1) / blockDim // ceil(n / blockDim) = 4
	Launch(blocks, blockDim, saxpy)

	fmt.Println(blocks, "blocks;", out[0], out[1], out[999]) // 4 blocks; 2 5 2999
}

No mutex is needed because each thread writes only out[i] for its own i. That is the discipline of GPU programming in one sentence.

Remember this#

  • A kernel is run by many threads; each finds its work from its index.
  • i = blockIdx * blockDim + threadIdx, followed by a bounds check.
  • Blocks ≤ 1,024 threads; threads in a block share an SM and shared memory.
  • Blocks are independent and unordered.

Try it#

  1. Remove the if i >= n check in grid.go. What happens, and why would the same bug on a real GPU be worse than a panic?
  2. Write a 2D version: a kernel with (x, y) thread and block indices that inverts a 1920×1080 grayscale image.
  3. Write a kernel where every thread adds its value to a single shared sum variable. Run it with go run -race. How would you fix it? (GPU kernels fix it with a reduction.)

Check yourself#

  1. How does a thread know which element to process?
  2. Why does every kernel need a bounds check?
  3. Which threads can share memory directly, and which cannot?

↑↓ navigate↵ openesc close