Short definitions of every term the course uses. The lesson in brackets is where it is explained.
| Term | Meaning |
|---|---|
| All-reduce | Collective operation after which every GPU holds the sum of all GPUs’ data. (VI.02) |
| Arithmetic intensity | FLOPs performed per byte moved through memory; a property of an operation. (IV.01) |
| Block | A group of up to 1,024 threads that run on one SM and can share memory. (III.01) |
| Coalescing | Merging a warp’s neighbouring memory requests into few transactions. (IV.02) |
| Compute-bound | Limited by arithmetic rate, not memory traffic. (IV.01) |
| Context | A process’s private state on a GPU. (VI.03) |
| CUDA | NVIDIA’s programming model and toolkit for general-purpose GPU computing. (I.03, III) |
| Device | The GPU and its memory, as opposed to the host. (I.01) |
| Device plugin | Kubernetes component that advertises GPUs to the kubelet. (VI.04) |
| Divergence | Threads of one warp taking different branches, forcing both paths to run. (II.01, IV.03) |
| ECC | Error-correcting code protecting memory against bit flips. (II.05) |
| FLOP | One floating-point operation. TFLOP/s = 10^12 per second. (I.04) |
| Fusion | Combining several operations into one kernel. (IV.04) |
| GEMM / GEMV | Matrix-matrix / matrix-vector multiplication. (V.01) |
| Grid | All the blocks of one kernel launch. (III.01) |
| HBM | High-bandwidth memory: stacked DRAM packaged beside the GPU. (II.03) |
| Host | The CPU and system RAM. (I.01) |
| Kernel | A function executed on the GPU by many threads. (I.01, III.01) |
| KV cache | Stored attention keys and values for previous tokens in an LLM. (V.03) |
| Memory-bound | Limited by memory bandwidth, not arithmetic. (II.03, IV.01) |
| MIG | Multi-Instance GPU: hardware partitioning into isolated smaller GPUs. (VI.03) |
| Mixed precision | Heavy arithmetic in low precision, sensitive steps in higher precision. (II.04) |
| MPS | Multi-Process Service: lets several processes run kernels concurrently on one GPU. (VI.03) |
| NCCL | NVIDIA’s library of multi-GPU collective operations. (VI.02) |
| NVLink | NVIDIA’s direct GPU-to-GPU interconnect. (VI.01) |
| NVML | NVIDIA Management Library: read GPU state. go-nvml is the Go binding. (III.05, IV.05) |
| Occupancy | Resident warps on an SM as a fraction of the maximum. (IV.03) |
| PCIe | The bus connecting a GPU to its host. (II.05) |
| Pinned memory | Page-locked host memory the GPU can read directly. (III.03) |
| PUE | Facility power ÷ IT power. (VII.03) |
| Quantization | Storing numbers in fewer bits. (V.02) |
| RDMA | Network transfers directly between memories, bypassing the CPUs. (VI.01) |
| Ridge point | Peak FLOP/s ÷ bandwidth; the intensity where a GPU switches regime. (IV.01) |
| Roofline | Model: attainable FLOP/s = min(peak, bandwidth × intensity). (IV.01) |
| Shared memory | Small fast per-SM scratchpad shared by a block’s threads. (II.02) |
| SIMT | Single instruction, multiple threads: how a warp executes. (II.01) |
| SM | Streaming multiprocessor: the repeating compute block of a GPU. (II.01) |
| Stream | An ordered queue of GPU operations. (III.04) |
| Tensor core | Circuit that multiplies small matrices in low precision. (II.04) |
| Tensor parallelism | Splitting every layer’s weights across GPUs. (VI.02) |
| Thermal throttling | The GPU lowering its clock to protect itself from heat. (II.05) |
| Unified memory | One address space the system migrates between host and device. (III.03) |
| VRAM | GPU memory. (I.01) |
| Warp | 32 threads that execute the same instruction together. (II.01) |
| XID | NVIDIA driver error event code. (II.05) |