PidokuInfra
06GPU Computing
On this page

GPU Computing

IntermediateModule 0612 topics~15h

Topics, in order

01 GPU ArchitectureA GPU is a chip organized around one idea: do the same operation to enormous amounts of data, and spend every available transistor on arithmetic and memory bandwidth rather than on making a … Beginner 1h 15m 02 Execution Model: Threads, Warps, Blocks, GridsThe hierarchy CUDA uses to organize parallel work. Intermediate 1h 15m 03 GPU Memory HierarchyFive levels of storage, each roughly 10x faster and 10x smaller than the one below. Intermediate 1h 15m 04 Bandwidth and the Roofline on GPUThe roofline model applied to GPUs: a single plot that tells you, for any kernel, whether it's limited by memory bandwidth or by arithmetic, and how far from the limit you are. Intermediate 1h 05 Kernel Launch and Host-Device InteractionEverything that happens at the CPU-GPU boundary: launching kernels, copying data, and synchronizing. This boundary costs microseconds, and at decode's kernel granularity, microseconds … Intermediate 1h 06 Streams and SynchronizationA stream is an ordered queue of GPU operations. Operations in the same stream execute in order; operations in different streams may overlap. Intermediate 1h 07 CUDA GraphsRecording a sequence of GPU operations once, then replaying the whole sequence with a single launch. Advanced 1h 08 CUDA Programming FundamentalsYou will rarely write production CUDA. You should write enough to read it fluently, to understand what your engine is doing, and to write a custom fused op when nothing else fits. Advanced 2h 09 Coalescing and Access PatternsCoalescing is the hardware combining a warp's 32 memory accesses into the minimum number of memory transactions. Advanced 1h 15m 10 Occupancy and Warp DivergenceTwo related properties of how well a kernel uses the SM. Advanced 1h 15m 11 Tensor CoresA hardware unit that performs a small matrix multiply-accumulate in one instruction, rather than a scalar one. Advanced 1h 15m 12 Profiling CUDAMeasuring what the GPU actually did, at two levels: Advanced 1h 30m

About this module

Goal: understand the machine well enough to predict its behavior, read a profiler, and write a kernel when you need to.

You do not need to become a CUDA expert to be an excellent inference engineer. You do need to understand the execution model, the memory hierarchy, and how to profile — because without them you cannot tell a hardware limit from a software bug.

Hardware numbers for every card you are likely to meet — VRAM, memory type, bandwidth, interconnects — live in GPU-HARDWARE.md. Keep it open while reading this section.

Files#

#FileLevelTime
01GPU architectureBeginner75 min
02Execution model: warps, blocks, gridsIntermediate75 min
03Memory hierarchyIntermediate75 min
04Bandwidth and the roofline on GPUIntermediate60 min
05Kernel launch and host-device interactionIntermediate60 min
06Streams and synchronizationIntermediate60 min
07CUDA graphsAdvanced60 min
08CUDA programming fundamentalsAdvanced120 min
09Coalescing and access patternsAdvanced75 min
10Occupancy and warp divergenceAdvanced75 min
11Tensor coresAdvanced75 min
12Profiling CUDAAdvanced90 min

The thread#

flowchart TD
  N0["A GPU is many simple cores plus enormous bandwidth<br/><b>01</b>"]
  N1["Organized as warps of 32 lanes executing in lockstep<br/><b>02</b>"]
  N2["Fed by a memory hierarchy: registers → shared → L2 → HBM<br/><b>03</b>"]
  N3["Whose bandwidth, not FLOPs, usually binds you<br/><b>04</b>"]
  N4["Driven by the CPU via launches (05) on streams<br/><b>06</b>"]
  N5["Which can be pre-recorded to eliminate overhead<br/><b>07</b>"]
  N6["And you can write your own<br/><b>08</b>"]
  N7["If you respect coalescing (09), occupancy, and divergence<br/><b>10</b>"]
  N8["And use tensor cores for anything matmul-shaped<br/><b>11</b>"]
  N9["And measure everything<br/><b>12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8,N9 memory

No GPU?#

Sections 01-07 and 09-11 are readable and the exercises have CPU-analogue or Colab alternatives. Google Colab’s free T4 is sufficient for every exercise in this section — it has tensor cores, Nsight works, and the concepts all transfer. Do not skip this section for lack of an H100.

↑↓ navigate↵ openesc close