PidokuInfra
04Neural Network Inference
On this page

Neural Network Inference

BasicModule 0412 topics~13h 30m

Topics, in order

01 Tracing One Request Through a ModelA step-by-step account of everything that happens between "user sends a prompt" and "user sees a token." Not the abstraction — the actual operations, in order, with sizes. Beginner 1h 15m 02 Computational Graphs and OperatorsA model is a directed acyclic graph: nodes are operators (matmul, add, softmax), edges are tensors. Inference is executing that graph in topological order. Intermediate 1h 03 GEMM and GEMV- GEMM — GEneral Matrix Multiply: C = α·A·B + β·C. Matrix times matrix. - GEMV — GEneral Matrix-Vector: y = α·A·x + β·y. Matrix times vector. Intermediate 1h 30m 04 ConvolutionsLLMs don't use convolutions. This file exists because (a) vision encoders in multimodal models do, (b) speech models do, (c) convolution is where many inference optimization techniques were … Intermediate 45 min 05 Normalization LayersAn operation that rescales activations so their magnitude stays in a predictable range. Intermediate 45 min 06 Attention Computation in PracticeHow attention is actually computed on a GPU, as opposed to how it's written in a paper. The difference is substantial: the textbook formula is unusable at production sequence lengths, and … Advanced 1h 30m 07 Tensor Layouts and MemoryLayout is how a logical tensor's elements map to physical memory addresses. The same (B, h, S, D) tensor can be stored in 24 different orders, and the choice changes kernel performance by … Advanced 1h 15m 08 Kernels and Kernel LaunchesA kernel is a function that runs on the GPU. A launch is the CPU telling the GPU to run one. Launches have a fixed cost — roughly 3-10 µs of CPU time and ~1-3 µs of GPU-side scheduling — … Intermediate 1h 09 Operator FusionCombining several operations into one kernel so intermediate results stay in registers or shared memory instead of round-tripping through HBM. Advanced 1h 15m 10 Graph Optimization and CompilersTaking the computational graph and rewriting it into a faster but mathematically equivalent graph, then generating code for it. Advanced 1h 15m 11 Static vs Dynamic ShapesStatic shapes are known at compile time. Dynamic shapes vary per request. LLM serving is inherently dynamic — prompt lengths, batch sizes, and generation lengths all vary — while nearly … Advanced 1h 12 Numerical Precision and Stability in PracticeThe practical discipline of keeping an inference system numerically correct while running it in the lowest precision you can get away with, with fused kernels, on multiple GPUs, at varying … Advanced 1h

About this module

Goal: turn “a model is math” into “a model is a program that a specific machine executes.”

Section III taught the mathematics. This section is about what actually runs: operators, kernels, layouts, graphs, and the transformations a compiler applies. It is the bridge between the model and the hardware.

Files#

#FileLevelTime
01Tracing one request through a modelBeginner75 min
02Computational graphs and operatorsIntermediate60 min
03GEMM and GEMVIntermediate90 min
04ConvolutionsIntermediate45 min
05Normalization layersIntermediate45 min
06Attention computation in practiceAdvanced90 min
07Tensor layouts and memoryAdvanced75 min
08Kernels and kernel launchesIntermediate60 min
09Operator fusionAdvanced75 min
10Graph optimization and compilersAdvanced75 min
11Static vs dynamic shapesAdvanced60 min
12Numerical precision and stability in practiceAdvanced60 min

The thread#

flowchart TD
  N0["A model is a graph of operators<br/><b>02</b>"]
  N1["Each operator becomes one or more kernels<br/><b>08</b>"]
  N2["Kernels read tensors whose LAYOUT determines their speed<br/><b>07</b>"]
  N3["Most kernels are GEMM (03) or attention (06) or memory-bound glue<br/><b>05</b>"]
  N4["The glue should be fused away<br/><b>09</b>"]
  N5["A compiler can do that automatically — if shapes cooperate<br/><b>10, 11</b>"]
  N6["And all of it must stay numerically sane<br/><b>12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6

  class N0,N1 neutral
  class N2 io
  class N3,N4 queue
  class N5 compute
  class N6 memory

↑↓ navigate↵ openesc close