PidokuInfra
Glossary

Glossary

Short definitions of every term the course uses. The lesson in brackets is where it is explained.

TermMeaning
All-reduceCollective operation after which every GPU holds the sum of all GPUs’ data. (VI.02)
Arithmetic intensityFLOPs performed per byte moved through memory; a property of an operation. (IV.01)
BlockA group of up to 1,024 threads that run on one SM and can share memory. (III.01)
CoalescingMerging a warp’s neighbouring memory requests into few transactions. (IV.02)
Compute-boundLimited by arithmetic rate, not memory traffic. (IV.01)
ContextA process’s private state on a GPU. (VI.03)
CUDANVIDIA’s programming model and toolkit for general-purpose GPU computing. (I.03, III)
DeviceThe GPU and its memory, as opposed to the host. (I.01)
Device pluginKubernetes component that advertises GPUs to the kubelet. (VI.04)
DivergenceThreads of one warp taking different branches, forcing both paths to run. (II.01, IV.03)
ECCError-correcting code protecting memory against bit flips. (II.05)
FLOPOne floating-point operation. TFLOP/s = 10^12 per second. (I.04)
FusionCombining several operations into one kernel. (IV.04)
GEMM / GEMVMatrix-matrix / matrix-vector multiplication. (V.01)
GridAll the blocks of one kernel launch. (III.01)
HBMHigh-bandwidth memory: stacked DRAM packaged beside the GPU. (II.03)
HostThe CPU and system RAM. (I.01)
KernelA function executed on the GPU by many threads. (I.01, III.01)
KV cacheStored attention keys and values for previous tokens in an LLM. (V.03)
Memory-boundLimited by memory bandwidth, not arithmetic. (II.03, IV.01)
MIGMulti-Instance GPU: hardware partitioning into isolated smaller GPUs. (VI.03)
Mixed precisionHeavy arithmetic in low precision, sensitive steps in higher precision. (II.04)
MPSMulti-Process Service: lets several processes run kernels concurrently on one GPU. (VI.03)
NCCLNVIDIA’s library of multi-GPU collective operations. (VI.02)
NVIDIA’s direct GPU-to-GPU interconnect. (VI.01)
NVMLNVIDIA Management Library: read GPU state. go-nvml is the Go binding. (III.05, IV.05)
OccupancyResident warps on an SM as a fraction of the maximum. (IV.03)
PCIeThe bus connecting a GPU to its host. (II.05)
Pinned memoryPage-locked host memory the GPU can read directly. (III.03)
PUEFacility power ÷ IT power. (VII.03)
QuantizationStoring numbers in fewer bits. (V.02)
RDMANetwork transfers directly between memories, bypassing the CPUs. (VI.01)
Ridge pointPeak FLOP/s ÷ bandwidth; the intensity where a GPU switches regime. (IV.01)
RooflineModel: attainable FLOP/s = min(peak, bandwidth × intensity). (IV.01)
Shared memorySmall fast per-SM scratchpad shared by a block’s threads. (II.02)
SIMTSingle instruction, multiple threads: how a warp executes. (II.01)
SMStreaming multiprocessor: the repeating compute block of a GPU. (II.01)
StreamAn ordered queue of GPU operations. (III.04)
Tensor coreCircuit that multiplies small matrices in low precision. (II.04)
Tensor parallelismSplitting every layer’s weights across GPUs. (VI.02)
Thermal throttlingThe GPU lowering its clock to protect itself from heat. (II.05)
Unified memoryOne address space the system migrates between host and device. (III.03)
VRAMGPU memory. (I.01)
Warp32 threads that execute the same instruction together. (II.01)
XIDNVIDIA driver error event code. (II.05)

↑↓ navigate↵ openesc close