PidokuInfra
07Inference Optimization
On this page

Inference Optimization

IntermediateModule 0714 topics~17h 30m

Topics, in order

01 Optimization MethodologyA repeatable procedure for making an inference system faster, that doesn't waste months on the wrong thing. Intermediate 1h 02 Quantization Overview and Decision GuideShipping by weight. Your freight cost is per kilogram, and your cargo is mostly packaging. Quantization is repacking into smaller boxes: same goods, less weight, lower cost — up to the point … Intermediate 1h 30m 03 Weight-Only Quantization: GPTQ and AWQBut the layer's output error is X · (W - Ŵ). A weight column that multiplies large activations contributes proportionally more error. Two insights follow: Advanced 1h 30m 04 Activation Quantization and SmoothQuantDettmers et al. (LLM.int8()) discovered this empirically: above roughly 6.7B parameters, transformers develop emergent outlier features — specific hidden dimensions where activation … Advanced 1h 15m 05 FP8 and INT8 InferenceThis is the production default for modern inference on Hopper and Blackwell. It is [ESTABLISHED], not experimental. Advanced 1h 15m 06 INT4 and Low-Bit InferenceNeural network weights are heavily over-parameterized and remarkably robust to precision reduction. Empirically: Advanced 1h 15m 07 Kernel and Operator FusionSection IV.09 covered fusion mechanically. This file covers it as an optimization decision: what to fuse in an LLM, what it's worth, and when it isn't. Advanced 1h 08 CUDA Graphs in ServingSection VI.07 covered CUDA graphs mechanically. This file covers the serving-specific decisions: capture lists, memory budgets, and feature interactions. Advanced 45 min 09 TensorRT and TensorRT-LLMTensorRT is NVIDIA's ahead-of-time inference compiler: it takes a model graph and builds an optimized, hardware-specific binary "engine." Advanced 1h 15m 10 FlashAttentionThe most important algorithmic optimization in modern transformer inference. Not because it's the fastest — because without it, long context is impossible. Advanced 2h 11 KV Cache OptimizationThis file is the catalogue and the decision guide. Individual techniques get their own files (12 for quantization, XIII for MLA/offload/disaggregation). Advanced 1h 30m 12 KV Cache QuantizationThey attack different terms and their benefits are situational in opposite ways: Advanced 1h 13 Speculative Decoding in PracticeSection V.12 covered the algorithm and the theory. This file covers deployment: which variant, how to tune it, how it interacts with everything else, and how to tell whether it's helping. Advanced 1h 15m 14 Pruning, Distillation, and Architecture OptimizationThree ways to make the model itself smaller, rather than making the same model run faster. Advanced 1h

About this module

Goal: a catalogue of the techniques that make inference faster, each analyzed the same way:

Problem  →  Why it happens  →  Optimization  →  How it works
         →  Trade-offs  →  When to use  →  When NOT to use

Every file follows that structure inside the standard 12-part template. The “when NOT to use” sections are the ones worth rereading.

Files#

#FileLevelTime
01Optimization methodologyIntermediate60 min
02Quantization overview and decision guideIntermediate90 min
03Weight-only quantization: GPTQ and AWQAdvanced90 min
04Activation quantization and SmoothQuantAdvanced75 min
05FP8 and INT8 inferenceAdvanced75 min
06INT4 and low-bit inferenceAdvanced75 min
07Kernel and operator fusionAdvanced60 min
08CUDA graphs in servingAdvanced45 min
09TensorRT and TensorRT-LLMAdvanced75 min
10FlashAttention ★Advanced120 min
11KV cache optimizationAdvanced90 min
12KV cache quantizationAdvanced60 min
13Speculative decoding in practiceAdvanced75 min
14Pruning, distillation, architectureAdvanced60 min

The decision tree#

flowchart TB
  S["System is slow"] --> M["Measure first - 01<br/>which phase, which regime?"]
  M --> R{"What binds you?"}
  R -->|"scheduling: low batch, idle slots"| SC["Fix this FIRST<br/>continuous batching - V.09, VIII"]
  R -->|"decode: memory bandwidth"| DEC["Move fewer bytes per token"]
  R -->|"prefill: compute"| PRE["Do less or cheaper math"]
  R -->|"launch / CPU"| CPU["Cut host overhead"]
  DEC --> D1["Quantize weights - 02 to 06"]
  DEC --> D2["Shrink or quantize KV - 11, 12"]
  DEC --> D3["Speculative decoding - 13"]
  PRE --> P1["Prefix caching - V.11"]
  PRE --> P2["FlashAttention - 10"]
  PRE --> P3["FP8 compute - 05"]
  PRE --> P4["TensorRT, fusion - 09, 07"]
  CPU --> C1["CUDA graphs - 08"]
  CPU --> C2["Fusion - 07"]

  class S,M neutral
  class R,SC queue
  class DEC,D1,D2,D3 memory
  class PRE,P1,P2,P3,P4 compute
  class CPU,C1,C2 io

The same tree with every option listed:

Is your system slow?
├─ Measure first (01). Which phase? Which regime?
│
├─ DECODE-BOUND (memory bandwidth)
│   ├─ Reduce weight bytes ....... quantization (02-06)
│   ├─ Reduce KV bytes ........... GQA/MLA (XIII), KV quantization (12)
│   ├─ More tokens per read ...... speculative decoding (13)
│   ├─ More sequences per read ... batching (V.09)
│   └─ More aggregate bandwidth .. tensor parallelism (IX)
│
├─ PREFILL-BOUND (compute)
│   ├─ Skip redundant work ....... prefix caching (V.11)
│   ├─ Faster attention .......... FlashAttention (10)
│   ├─ Lower-precision compute ... FP8 (05), not INT4 weight-only
│   └─ Better kernels ............ TensorRT (09), fusion (07)
│
├─ LAUNCH/CPU-BOUND
│   ├─ CUDA graphs ............... (08)
│   ├─ Fusion .................... (07)
│   └─ Move logic off the hot path
│
└─ SCHEDULING-BOUND (low batch, idle slots)
    └─ This is Section V.09 and VIII, not this section.
       Fix it FIRST — it's a bigger lever than anything here.

Order matters. A 20% kernel improvement on a system running at batch 4 when it could run at batch 64 is worth almost nothing. Fix scheduling, then precision, then kernels.

↑↓ navigate↵ openesc close