Goal: turn every section of the curriculum into something that runs, and that you measured.
Fifteen builds. Each one reuses the code from the ones before it, so keep them in one repo
(inference-lab/ is a fine name) with one folder per project and a shared numbers.md.
Index#
| # | Project | After section | Level | Time | GPU needed? |
|---|---|---|---|---|---|
| 01 | From-scratch inference engine | III | Beginner | 6-8 h | No |
| 02 | CPU matmul benchmark | II, III | Beginner | 6-8 h | No |
| 03 | First GPU kernel | VI | Intermediate | 8-10 h | Yes (T4 is enough) |
| 04 | Tiny transformer engine | IV, V | Intermediate | 8-12 h | No |
| 05 | KV cache ★ | V | Intermediate | 6-8 h | No |
| 06 | LLM inference server | VIII | Intermediate | 8-12 h | No |
| 07 | Dynamic batching | VIII | Intermediate | 8-10 h | Helps |
| 08 | Continuous batching ★ | V, VIII | Advanced | 12-18 h | Helps |
| 09 | KV cache manager ★ | V, X | Advanced | 10-14 h | No |
| 10 | Quantized inference | VII | Advanced | 10-14 h | Helps |
| 11 | GPU benchmark suite | VI, X | Advanced | 8-12 h | Yes |
| 12 | Inference gateway | VIII, XI | Advanced | 12-16 h | No |
| 13 | Multi-GPU inference | IX | Expert | 12-16 h | 2 GPUs, or simulate |
| 14 | Distributed inference | IX | Expert | 14-20 h | Optional |
| 15 | Mini inference platform | XII | Expert | 20-30 h | Optional |
★ = do not skip. These three are the ones interviewers ask you to whiteboard.
The thread#
flowchart TD N0["A forward pass is just array math<br/><b>01</b>"] N1["Whose speed is set by FLOPs and bytes moved, which you can measure<br/><b>02</b>"] N2["On a GPU, where you launch kernels yourself<br/><b>03</b>"] N3["A transformer is the same thing with attention<br/><b>04</b>"] N4["Made affordable by caching K and V<br/><b>05</b>"] N5["Wrapped in an API that streams<br/><b>06</b>"] N6["That batches requests to amortize weight reads<br/><b>07</b>"] N7["At every step rather than once per batch<br/><b>08</b>"] N8["With KV memory managed in pages<br/><b>09</b>"] N9["And weights shrunk to move fewer bytes<br/><b>10</b>"] N10["On hardware you have characterized yourself<br/><b>11</b>"] N11["Behind a front door that protects it<br/><b>12</b>"] N12["Split across GPUs when one is not enough<br/><b>13</b>"] N13["And across machines when one box is not enough<br/><b>14</b>"] N14["All of it run as a platform<br/><b>15</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10 --> N11 --> N12 --> N13 --> N14 class N0,N1,N2 neutral class N3,N4,N5 io class N6,N7,N8 queue class N9,N10,N11 compute class N12,N13,N14 memory
Rules that apply to every project#
- Correctness before speed. Every project has a reference to match (PyTorch, Hugging Face, or your own earlier project). Do not benchmark code that produces different tokens.
- Write the prediction first. Before each measurement, write down the number you expect and why. The gap between prediction and measurement is the lesson.
- Record everything in
numbers.mdwith the machine, the date, and the command. - Warm up, then measure. Discard the first runs; report median and p95, never a single run.
- Synchronize before timing GPU code, or you are timing the launch, not the work (Section VI.06).
Doing the projects in Go#
The projects are designed to be built in Go, with one deliberate exception: the model weights you compare against come from the Python ecosystem, because that is where trained models are published.
| Projects | In Go | What still touches Python / CUDA |
|---|---|---|
| 01, 02, 04, 05 — engine, matmul, transformer, KV cache | Everything: tensors, operators, the forward pass, the cache. Standard library only. | A ten-line script, run once, to export reference weights and expected outputs from PyTorch to safetensors (I.03 shows how to read that format in Go). |
| 06, 07, 08, 09, 12 — server, batching, KV manager, gateway | Everything. These are the projects closest to professional Go work: net/http, channels, contexts, schedulers. | Nothing. The engine behind your server is your own Project 04/05 model, or any OpenAI-compatible server you put it in front of. |
| 10 — quantization | The quantizers and the accuracy measurements. | Reference perplexity numbers, if you want to compare. |
| 03 — first GPU kernel | The host program, through cgo (see the GPU course, III.02). | The kernel itself is CUDA C. That is true in every language. |
| 11, 13, 14 — GPU benchmarks, multi-GPU, distributed | Load generators, harnesses, the KV transfer protocol. | The on-GPU model runs in an existing engine (vLLM or PyTorch). The skeletons in these three projects are shown in Python for that reason. |
| 15 — platform | The gateway, router and autoscaler. | The operator skeleton is shown with a Python framework; in Go you would use controller-runtime. |
Every project asks for a Model you can swap. Define it once and reuse it everywhere:
// Model is the seam between your serving code and whatever does the arithmetic.
type Model interface {
Prefill(ids []int) (logits []float32, cache KVCache)
Decode(token int, cache KVCache) (logits []float32, next KVCache)
}Start with your own pure-Go implementation (slow, fully understood). Later, add a second implementation that forwards to a real engine over HTTP. Your server, batcher and gateway code does not change — which is the point of the interface.
Suggested repo layout#
inference-lab/
numbers.md
go.mod
common/ # timers, load generator, plotting — grows as you go
p01_engine/
p02_matmul_bench/
...
p15_platform/Model to use throughout#
GPT-2 small (124M) is the default from Project 04 onward: it runs on a laptop CPU, its weights are public, and every formula in Section V applies to it unchanged. Where a project benefits from a modern architecture (RoPE, GQA, RMSNorm), swap in a ≤1B model such as Qwen2.5-0.5B or SmolLM2-360M.