PidokuInfra
10Memory and Performance Engineering
On this page

Memory and Performance Engineering

AdvancedModule 1011 topics~13h 30m

Topics, in order

01 The Performance Debugging MethodologyIf you internalize one thing from Section X, make it this procedure. Advanced 1h 30m 02 The Roofline Model in PracticeSections I.08 and VI.04 introduced the roofline. This file is about using it on a real system: how to place your kernels on it, and what each position tells you to do. Advanced 1h 03 Memory-Bound vs Compute-BoundEvery optimization helps exactly one of these and does nothing for the other two. Applying the wrong one is the single largest source of wasted performance work. Advanced 1h 04 GPU Utilization Is a LieThis one misconception has probably wasted more GPU-hours than any other in the field. Intermediate 1h 05 Memory Fragmentation and Allocators"Reserved but unallocated" is fragmentation. PyTorch holds it, your tensors don't use it, and it can't satisfy your 2 GB request because it's in pieces. Advanced 1h 15m 06 OOM: Causes and CuresNot all OOMs are the same. The first job is to classify. Advanced 1h 07 Benchmarking Inference CorrectlyA wrong benchmark is worse than no benchmark, because it produces confident wrong decisions. Advanced 1h 30m 08 GPU Profiling with NsightSection VI.12 introduced the tools. This file is a working recipe book for inference-specific profiling. Advanced 1h 30m 09 CPU and Python ProfilingBecause a substantial fraction of inference performance problems are on the CPU: Intermediate 1h 10 Metrics, Prometheus, and DashboardsThe complete metric set for an LLM inference service, organized by what each answers. Intermediate 1h 15m 11 Case StudiesEight worked investigations. Read each symptom, stop, and write down your hypotheses before reading the diagnosis. That exercise is the point of this file. Advanced 1h 30m

About this module

Goal: answer “why is my inference system slow?” systematically, every time, without guessing.

This section is the toolbox. Sections I-IX taught you what the system does; this one teaches you how to find out what it’s actually doing.

Files#

#FileLevelTime
01The performance debugging methodology ★Advanced90 min
02The roofline model in practiceAdvanced60 min
03Memory-bound vs compute-boundAdvanced60 min
04GPU utilization is a lie ★Intermediate60 min
05Memory fragmentation and allocatorsAdvanced75 min
06OOM: causes and curesAdvanced60 min
07Benchmarking inference correctly ★Advanced90 min
08GPU profiling with NsightAdvanced90 min
09CPU and Python profilingIntermediate60 min
10Metrics, Prometheus, GrafanaIntermediate75 min
11Case studies ★Advanced90 min

The method, in one diagram#

flowchart TB
  S["It's slow"] --> A["1. WHICH METRIC?<br/>TTFT, ITL, throughput, cost"]
  A --> B["2. WHICH PHASE?<br/>queue, tokenize, prefill, decode, detokenize, network<br/><b>file 10</b>"]
  B --> C["3. GPU, CPU, OR WAITING?<br/>timeline analysis<br/><b>files 08, 09</b>"]
  C --> D["4. WHICH REGIME?<br/>memory, compute, latency, launch<br/><b>files 02, 03</b>"]
  D --> E["5. WHICH KERNEL?<br/>per-kernel analysis<br/><b>file 08</b>"]
  E --> F["6. FIX ONE THING, RE-MEASURE"]
  F -.->|"still slow"| A

  class S warn
  class A,B neutral
  class C io
  class D queue
  class E compute
  class F memory

Never skip a level. Starting at step 5 optimizes kernels that don’t matter.

Checkpoint E#

  1. Given a slow inference service, list your first five measurements in order.
  2. Draw a roofline and place decode, prefill, and RMSNorm on it.
  3. nvidia-smi says 100% GPU utilization. Why might the GPU still be mostly idle?
  4. What makes an inference benchmark valid? Name five requirements.
  5. Your p99 ITL is 5x your p50. Give four hypotheses and how you’d distinguish them.

↑↓ navigate↵ openesc close