Goal: answer “why is my inference system slow?” systematically, every time, without guessing.
This section is the toolbox. Sections I-IX taught you what the system does; this one teaches you how to find out what it’s actually doing.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | The performance debugging methodology ★ | Advanced | 90 min |
| 02 | The roofline model in practice | Advanced | 60 min |
| 03 | Memory-bound vs compute-bound | Advanced | 60 min |
| 04 | GPU utilization is a lie ★ | Intermediate | 60 min |
| 05 | Memory fragmentation and allocators | Advanced | 75 min |
| 06 | OOM: causes and cures | Advanced | 60 min |
| 07 | Benchmarking inference correctly ★ | Advanced | 90 min |
| 08 | GPU profiling with Nsight | Advanced | 90 min |
| 09 | CPU and Python profiling | Intermediate | 60 min |
| 10 | Metrics, Prometheus, Grafana | Intermediate | 75 min |
| 11 | Case studies ★ | Advanced | 90 min |
The method, in one diagram#
flowchart TB S["It's slow"] --> A["1. WHICH METRIC?<br/>TTFT, ITL, throughput, cost"] A --> B["2. WHICH PHASE?<br/>queue, tokenize, prefill, decode, detokenize, network<br/><b>file 10</b>"] B --> C["3. GPU, CPU, OR WAITING?<br/>timeline analysis<br/><b>files 08, 09</b>"] C --> D["4. WHICH REGIME?<br/>memory, compute, latency, launch<br/><b>files 02, 03</b>"] D --> E["5. WHICH KERNEL?<br/>per-kernel analysis<br/><b>file 08</b>"] E --> F["6. FIX ONE THING, RE-MEASURE"] F -.->|"still slow"| A class S warn class A,B neutral class C io class D queue class E compute class F memory
Never skip a level. Starting at step 5 optimizes kernels that don’t matter.
Checkpoint E#
- Given a slow inference service, list your first five measurements in order.
- Draw a roofline and place decode, prefill, and RMSNorm on it.
nvidia-smisays 100% GPU utilization. Why might the GPU still be mostly idle?- What makes an inference benchmark valid? Name five requirements.
- Your p99 ITL is 5x your p50. Give four hypotheses and how you’d distinguish them.