This is the core section of the curriculum. Everything before it was preparation; everything after it is elaboration.
If you learn one section properly, make it this one. The concepts here — prefill vs decode, the KV cache, continuous batching, paged attention — are what separate people who can operate an LLM serving system from people who can only start one.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | Transformer inference overview | Intermediate | 60 min |
| 02 | Tokenization and its consequences | Beginner | 60 min |
| 03 | Prefill vs decode ★ | Intermediate | 120 min |
| 04 | Autoregressive generation | Intermediate | 60 min |
| 05 | The KV cache ★ | Intermediate | 120 min |
| 06 | KV cache math ★ | Intermediate | 90 min |
| 07 | Context length and how it scales | Advanced | 75 min |
| 08 | Batching: static and dynamic | Intermediate | 75 min |
| 09 | Continuous batching ★ | Advanced | 120 min |
| 10 | PagedAttention ★ | Advanced | 120 min |
| 11 | Prefix and prompt caching | Advanced | 90 min |
| 12 | Speculative decoding | Advanced | 90 min |
| 13 | Sampling and decoding strategies | Intermediate | 75 min |
| 14 | Streaming inference | Intermediate | 60 min |
| 15 | Capacity math: worked examples ★ | Advanced | 120 min |
★ = do not skip.
The thread#
flowchart TD N0["Text becomes tokens<br/><b>02</b>"] N1["The prompt is processed in one parallel pass — PREFILL<br/><b>03</b>"] N2["Then tokens are generated one at a time — DECODE<br/><b>03, 04</b>"] N3["Which is only affordable because we cache K and V<br/><b>05</b>"] N4["And that cache costs memory that we can compute exactly<br/><b>06</b>"] N5["Memory that grows with context length<br/><b>07</b>"] N6["So we batch to amortize weight reads<br/><b>08</b>"] N7["But static batching wastes 60-80% of slots<br/><b>08</b>"] N8["So we schedule at every step — CONTINUOUS BATCHING<br/><b>09</b>"] N9["And manage KV memory in pages to avoid fragmentation<br/><b>10</b>"] N10["And reuse KV across requests when prefixes match<br/><b>11</b>"] N11["And generate several tokens per weight read when we can<br/><b>12</b>"] N12["Choosing each token from a distribution<br/><b>13</b>"] N13["Streaming them out as they appear<br/><b>14</b>"] N14["All of which we can size on paper before we build it<br/><b>15</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10 --> N11 --> N12 --> N13 --> N14 class N0,N1,N2 neutral class N3,N4,N5 io class N6,N7,N8 queue class N9,N10,N11 compute class N12,N13,N14 memory
Checkpoint C — the big one#
Before moving on, you should be able to answer without notes:
- Draw the timeline of a single request through prefill and decode.
- Compute the KV cache size for Llama-3-8B at 8k context, batch 32. Show every term.
- Explain continuous batching to a backend engineer in 60 seconds.
- Why does PagedAttention exist? What exactly was wasteful before it?
- At what batch size does a decode step stop being memory-bound on your hardware?
- A user reports 4-second TTFT. List the five things you check, in order.
- Given 8×H100 and Llama-3-70B, how many concurrent users at 8k context? Show the arithmetic.