Goal: teach exactly the ML you need to reason about inference — no more, no less.
This is not an ML course. There is no training theory, no optimization, no generalization bounds. What is here: the linear algebra you will do by hand, the architecture you will trace through, and the numerics you will trade away for speed.
If you already know ML, do not skip files 07, 11, and 12 — softmax stability, number formats, and quantization are where ML people most often have gaps that matter for inference.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | Vectors and matrices | Beginner | 60 min |
| 02 | Matrix multiplication by hand | Beginner | 75 min |
| 03 | Tensors, shapes, and broadcasting | Beginner | 60 min |
| 04 | Neural networks from scratch | Beginner | 90 min |
| 05 | Activation functions | Beginner | 45 min |
| 06 | Embeddings | Beginner | 60 min |
| 07 | Softmax and numerical stability | Intermediate | 75 min |
| 08 | Attention from first principles | Intermediate | 120 min |
| 09 | The transformer | Intermediate | 120 min |
| 10 | Training vs inference: the math | Intermediate | 45 min |
| 11 | Number formats: FP32, FP16, BF16, FP8 | Intermediate | 90 min |
| 12 | Quantization fundamentals | Intermediate | 90 min |
The thread#
flowchart TD N0["Numbers in arrays<br/><b>01, 03</b>"] N1["Multiplied by learned arrays<br/><b>02</b>"] N2["Stacked with nonlinearity between them<br/><b>04, 05</b>"] N3["Text becomes vectors<br/><b>06</b>"] N4["Scores become probabilities<br/><b>07</b>"] N5["Tokens look at each other<br/><b>08</b>"] N6["Assembled into a transformer<br/><b>09</b>"] N7["Which we run forward only<br/><b>10</b>"] N8["In as few bits as we can get away with<br/><b>11, 12</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 class N0,N1 neutral class N2,N3 io class N4,N5 queue class N6,N7 compute class N8 memory
Checkpoint B#
- Multiply a 2×3 by a 3×2 matrix by hand, showing every dot product.
- Explain attention in three sentences with no equations.
- Why is softmax computed with a max subtraction? What breaks without it?
- What does BF16 trade away relative to FP16, and why does that trade matter?
- Write the parameter count formula for a decoder-only transformer and apply it to a real model.
- What is the difference between per-tensor and per-channel quantization, and why does the latter usually win?