Goal: turn “a model is math” into “a model is a program that a specific machine executes.”
Section III taught the mathematics. This section is about what actually runs: operators, kernels, layouts, graphs, and the transformations a compiler applies. It is the bridge between the model and the hardware.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | Tracing one request through a model | Beginner | 75 min |
| 02 | Computational graphs and operators | Intermediate | 60 min |
| 03 | GEMM and GEMV | Intermediate | 90 min |
| 04 | Convolutions | Intermediate | 45 min |
| 05 | Normalization layers | Intermediate | 45 min |
| 06 | Attention computation in practice | Advanced | 90 min |
| 07 | Tensor layouts and memory | Advanced | 75 min |
| 08 | Kernels and kernel launches | Intermediate | 60 min |
| 09 | Operator fusion | Advanced | 75 min |
| 10 | Graph optimization and compilers | Advanced | 75 min |
| 11 | Static vs dynamic shapes | Advanced | 60 min |
| 12 | Numerical precision and stability in practice | Advanced | 60 min |
The thread#
flowchart TD N0["A model is a graph of operators<br/><b>02</b>"] N1["Each operator becomes one or more kernels<br/><b>08</b>"] N2["Kernels read tensors whose LAYOUT determines their speed<br/><b>07</b>"] N3["Most kernels are GEMM (03) or attention (06) or memory-bound glue<br/><b>05</b>"] N4["The glue should be fused away<br/><b>09</b>"] N5["A compiler can do that automatically — if shapes cooperate<br/><b>10, 11</b>"] N6["And all of it must stay numerically sane<br/><b>12</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 class N0,N1 neutral class N2 io class N3,N4 queue class N5 compute class N6 memory