PidokuInfra
05LLM Inference
On this page

LLM Inference

BasicModule 0515 topics~22h 15m

Topics, in order

01 Transformer Inference OverviewThe map of everything in this section, and how the pieces connect. Read this file, then read it again after finishing the section — it will mean something different the second time. Intermediate 1h 02 Tokenization and Its ConsequencesConverting text to integers the model can process, and back again. Beginner 1h 03 Prefill vs DecodeThe single most important distinction in LLM serving. Everything else in this section follows from it. Intermediate 2h 04 Autoregressive GenerationGenerating a sequence one element at a time, where each new element is conditioned on all previous ones — including the ones the model itself just generated. Intermediate 1h 05 The KV CacheThe most important optimization in LLM inference — and the source of most of its problems. Intermediate 2h 06 KV Cache MathPure arithmetic. Work every example with a calculator. These calculations are what you'll be asked to do in an interview and what you'll do weekly on the job. Intermediate 1h 30m 07 Context Length and How It ScalesThe maximum number of tokens (prompt + generation) a model can attend over. Modern models advertise 128k, 200k, or 1M. Serving those numbers is a fundamentally different engineering problem … Advanced 1h 15m 08 Batching: Static and DynamicThis file covers the first three and shows exactly why they're insufficient for autoregressive generation, which motivates file 09. Intermediate 1h 15m 09 Continuous BatchingThe single largest throughput win in modern LLM serving — typically 3-10x. Also called iteration-level scheduling (Orca) or in-flight batching (TensorRT-LLM). Advanced 2h 10 PagedAttentionIf you read Section II.06 (virtual memory), this file will feel inevitable rather than clever. That's the point — PagedAttention is operating systems applied to KV cache. Advanced 2h 11 Prefix and Prompt CachingReusing the KV cache computed for a shared prompt prefix across requests, instead of recomputing it. Advanced 1h 30m 12 Speculative DecodingUse a cheap model to guess the next few tokens, then use the expensive model to verify all of them in one forward pass. Accept the guesses that match what the big model would have produced; … Advanced 1h 30m 13 Sampling and Decoding StrategiesThe model produces a probability distribution over the vocabulary. Sampling is choosing one token from it. Intermediate 1h 15m 14 Streaming InferenceSending tokens to the client as they're generated, rather than waiting for the complete response. Intermediate 1h 15 Capacity Math: Worked ExamplesThe capstone of Section V. Work every example with a calculator. This is the skill that gets you hired and that you'll use weekly. Advanced 2h

About this module

This is the core section of the curriculum. Everything before it was preparation; everything after it is elaboration.

If you learn one section properly, make it this one. The concepts here — prefill vs decode, the KV cache, continuous batching, paged attention — are what separate people who can operate an LLM serving system from people who can only start one.

Files#

#FileLevelTime
01Transformer inference overviewIntermediate60 min
02Tokenization and its consequencesBeginner60 min
03Prefill vs decode ★Intermediate120 min
04Autoregressive generationIntermediate60 min
05The KV cache ★Intermediate120 min
06KV cache math ★Intermediate90 min
07Context length and how it scalesAdvanced75 min
08Batching: static and dynamicIntermediate75 min
09Continuous batching ★Advanced120 min
10PagedAttention ★Advanced120 min
11Prefix and prompt cachingAdvanced90 min
12Speculative decodingAdvanced90 min
13Sampling and decoding strategiesIntermediate75 min
14Streaming inferenceIntermediate60 min
15Capacity math: worked examples ★Advanced120 min

★ = do not skip.

The thread#

flowchart TD
  N0["Text becomes tokens<br/><b>02</b>"]
  N1["The prompt is processed in one parallel pass — PREFILL<br/><b>03</b>"]
  N2["Then tokens are generated one at a time — DECODE<br/><b>03, 04</b>"]
  N3["Which is only affordable because we cache K and V<br/><b>05</b>"]
  N4["And that cache costs memory that we can compute exactly<br/><b>06</b>"]
  N5["Memory that grows with context length<br/><b>07</b>"]
  N6["So we batch to amortize weight reads<br/><b>08</b>"]
  N7["But static batching wastes 60-80% of slots<br/><b>08</b>"]
  N8["So we schedule at every step — CONTINUOUS BATCHING<br/><b>09</b>"]
  N9["And manage KV memory in pages to avoid fragmentation<br/><b>10</b>"]
  N10["And reuse KV across requests when prefixes match<br/><b>11</b>"]
  N11["And generate several tokens per weight read when we can<br/><b>12</b>"]
  N12["Choosing each token from a distribution<br/><b>13</b>"]
  N13["Streaming them out as they appear<br/><b>14</b>"]
  N14["All of which we can size on paper before we build it<br/><b>15</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10 --> N11 --> N12 --> N13 --> N14

  class N0,N1,N2 neutral
  class N3,N4,N5 io
  class N6,N7,N8 queue
  class N9,N10,N11 compute
  class N12,N13,N14 memory

Checkpoint C — the big one#

Before moving on, you should be able to answer without notes:

  1. Draw the timeline of a single request through prefill and decode.
  2. Compute the KV cache size for Llama-3-8B at 8k context, batch 32. Show every term.
  3. Explain continuous batching to a backend engineer in 60 seconds.
  4. Why does PagedAttention exist? What exactly was wasteful before it?
  5. At what batch size does a decode step stop being memory-bound on your hardware?
  6. A user reports 4-second TTFT. List the five things you check, in order.
  7. Given 8×H100 and Llama-3-70B, how many concurrent users at 8k context? Show the arithmetic.

↑↓ navigate↵ openesc close