PidokuInfra

Projects

Module15 topics~148h

Topics, in order

01 A From-Scratch Inference EngineRun a trained network with nothing but Go's standard library, and match PyTorch to five decimal places. Beginner 6h 02 CPU Matmul BenchmarkMake the same matrix multiply 1,000× faster without changing the math — and explain every factor. Beginner 6h 03 First GPU KernelWrite CUDA by hand, time it correctly, and find out why your first version loses to the CPU. Intermediate 8h 04 Tiny Transformer EngineLoad real GPT-2 weights into your own code and generate text — deliberately the slow way. Intermediate 8h 05 KV CacheAdd the single most important optimization in LLM inference to your own engine, and prove the outputs did not change. Intermediate 6h 06 LLM Inference ServerWrap your engine in an OpenAI-compatible streaming API that behaves well when things go wrong. Intermediate 8h 07 Dynamic BatchingGroup requests that arrive close together, and map the throughput/latency trade you just bought. Intermediate 8h 08 Continuous BatchingBuild the scheduler loop at the heart of vLLM, SGLang, and TGI: decide the batch at every step. Advanced 12h 09 KV Cache ManagerReplace "one big slot per request" with paged blocks, and get prefix sharing almost for free. Advanced 10h 10 Quantized InferenceShrink the weights, measure what you lost and what you gained, and learn why the two are not the same question. Advanced 10h 11 GPU Benchmark SuiteCharacterize a GPU in ten minutes: produce the handful of numbers from which you can predict any model's performance on it. Advanced 8h 12 Inference GatewayBuild the front door: the component that decides who gets in, where they go, and what happens when a backend dies. Advanced 12h 13 Multi-GPU InferenceSplit one model across two devices by hand — tensor parallelism — and find out when it is worse than two independent copies. Expert 12h 14 Distributed InferenceCross the machine boundary: split prefill from decode, ship the KV cache over a network, and learn when that is cheaper than recomputing it. Expert 14h 15 Mini Inference PlatformThe capstone. Put a control plane around everything you have built: declare a model, and the platform deploys it, routes to it, scales it, rolls it, and bills for it. Expert 20h

About this module

Goal: turn every section of the curriculum into something that runs, and that you measured.

Fifteen builds. Each one reuses the code from the ones before it, so keep them in one repo (inference-lab/ is a fine name) with one folder per project and a shared numbers.md.

Index#

#ProjectAfter sectionLevelTimeGPU needed?
01From-scratch inference engineIIIBeginner6-8 hNo
02CPU matmul benchmarkII, IIIBeginner6-8 hNo
03First GPU kernelVIIntermediate8-10 hYes (T4 is enough)
04Tiny transformer engineIV, VIntermediate8-12 hNo
05KV cache ★VIntermediate6-8 hNo
06LLM inference serverVIIIIntermediate8-12 hNo
07Dynamic batchingVIIIIntermediate8-10 hHelps
08Continuous batching ★V, VIIIAdvanced12-18 hHelps
09KV cache manager ★V, XAdvanced10-14 hNo
10Quantized inferenceVIIAdvanced10-14 hHelps
11GPU benchmark suiteVI, XAdvanced8-12 hYes
12Inference gatewayVIII, XIAdvanced12-16 hNo
13Multi-GPU inferenceIXExpert12-16 h2 GPUs, or simulate
14Distributed inferenceIXExpert14-20 hOptional
15Mini inference platformXIIExpert20-30 hOptional

★ = do not skip. These three are the ones interviewers ask you to whiteboard.

The thread#

flowchart TD
  N0["A forward pass is just array math<br/><b>01</b>"]
  N1["Whose speed is set by FLOPs and bytes moved, which you can measure<br/><b>02</b>"]
  N2["On a GPU, where you launch kernels yourself<br/><b>03</b>"]
  N3["A transformer is the same thing with attention<br/><b>04</b>"]
  N4["Made affordable by caching K and V<br/><b>05</b>"]
  N5["Wrapped in an API that streams<br/><b>06</b>"]
  N6["That batches requests to amortize weight reads<br/><b>07</b>"]
  N7["At every step rather than once per batch<br/><b>08</b>"]
  N8["With KV memory managed in pages<br/><b>09</b>"]
  N9["And weights shrunk to move fewer bytes<br/><b>10</b>"]
  N10["On hardware you have characterized yourself<br/><b>11</b>"]
  N11["Behind a front door that protects it<br/><b>12</b>"]
  N12["Split across GPUs when one is not enough<br/><b>13</b>"]
  N13["And across machines when one box is not enough<br/><b>14</b>"]
  N14["All of it run as a platform<br/><b>15</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10 --> N11 --> N12 --> N13 --> N14

  class N0,N1,N2 neutral
  class N3,N4,N5 io
  class N6,N7,N8 queue
  class N9,N10,N11 compute
  class N12,N13,N14 memory

Rules that apply to every project#

  1. Correctness before speed. Every project has a reference to match (PyTorch, Hugging Face, or your own earlier project). Do not benchmark code that produces different tokens.
  2. Write the prediction first. Before each measurement, write down the number you expect and why. The gap between prediction and measurement is the lesson.
  3. Record everything in numbers.md with the machine, the date, and the command.
  4. Warm up, then measure. Discard the first runs; report median and p95, never a single run.
  5. Synchronize before timing GPU code, or you are timing the launch, not the work (Section VI.06).

Doing the projects in Go#

The projects are designed to be built in Go, with one deliberate exception: the model weights you compare against come from the Python ecosystem, because that is where trained models are published.

ProjectsIn GoWhat still touches Python / CUDA
01, 02, 04, 05 — engine, matmul, transformer, KV cacheEverything: tensors, operators, the forward pass, the cache. Standard library only.A ten-line script, run once, to export reference weights and expected outputs from PyTorch to safetensors (I.03 shows how to read that format in Go).
06, 07, 08, 09, 12 — server, batching, KV manager, gatewayEverything. These are the projects closest to professional Go work: net/http, channels, contexts, schedulers.Nothing. The engine behind your server is your own Project 04/05 model, or any OpenAI-compatible server you put it in front of.
10 — quantizationThe quantizers and the accuracy measurements.Reference perplexity numbers, if you want to compare.
03 — first GPU kernelThe host program, through cgo (see the GPU course, III.02).The kernel itself is CUDA C. That is true in every language.
11, 13, 14 — GPU benchmarks, multi-GPU, distributedLoad generators, harnesses, the KV transfer protocol.The on-GPU model runs in an existing engine (vLLM or PyTorch). The skeletons in these three projects are shown in Python for that reason.
15 — platformThe gateway, router and autoscaler.The operator skeleton is shown with a Python framework; in Go you would use controller-runtime.

Every project asks for a Model you can swap. Define it once and reuse it everywhere:

Go
// Model is the seam between your serving code and whatever does the arithmetic.
type Model interface {
	Prefill(ids []int) (logits []float32, cache KVCache)
	Decode(token int, cache KVCache) (logits []float32, next KVCache)
}

Start with your own pure-Go implementation (slow, fully understood). Later, add a second implementation that forwards to a real engine over HTTP. Your server, batcher and gateway code does not change — which is the point of the interface.

Suggested repo layout#

inference-lab/
  numbers.md
  go.mod
  common/            # timers, load generator, plotting — grows as you go
  p01_engine/
  p02_matmul_bench/
  ...
  p15_platform/

Model to use throughout#

GPT-2 small (124M) is the default from Project 04 onward: it runs on a laptop CPU, its weights are public, and every formula in Section V applies to it unchanged. Where a project benefits from a modern architecture (RoPE, GQA, RMSNorm), swap in a ≤1B model such as Qwen2.5-0.5B or SmolLM2-360M.

↑↓ navigate↵ openesc close