PidokuInfra

Optimization Methodology

Intermediate 1h Difficulty 3/5 Topic 01 of 14

Prerequisites I.07, VI.12


1. What is it?#

A repeatable procedure for making an inference system faster, that doesn’t waste months on the wrong thing.

1. DEFINE  what "faster" means (which metric, what target)
2. MEASURE where the time goes (phase, then kernel)
3. CLASSIFY the bottleneck (memory / compute / launch / scheduling / queueing)
4. SELECT   an optimization that attacks THAT bottleneck
5. ESTIMATE the expected gain before doing the work
6. APPLY    one change
7. VERIFY   the gain, and the quality
8. REPEAT

Steps 3 and 5 are the ones people skip, and skipping them is why optimization projects fail.

Diagram — The optimization loop#

flowchart LR
  M["Measure<br/>baseline + SLO"] --> F["Find the bottleneck<br/>profile, do not guess"]
  F --> C["Classify<br/>memory / compute / host / queue"]
  C --> X["Apply ONE change"]
  X --> V["Verify<br/>speed AND output quality"]
  V -->|"still short of goal"| F
  V -->|"goal met"| S["Stop"]

  class M,F neutral
  class C queue
  class X compute
  class V memory
  class S io

2. Why does it exist?#

Because inference performance is counterintuitive, and the intuitions people bring from other domains actively mislead:

"We reduced FLOPs by 40%"          → zero gain if memory-bound
"We got 100% GPU utilization"      → means nothing (Section X.04)
"The kernel is 2x faster"          → 5% end-to-end if it was 10% of time
"Quantization made it 4x faster"   → for decode; prefill may be unchanged
"More GPUs will help"              → only with the right parallelism, and not always

Every one of those is a real statement someone has made in a real project, and every one wasted real time.


3. Simple analogy#

Amdahl’s law, restated as a rule of thumb: if a component is 10% of your time, making it infinitely fast gives you 11% overall. Optimizing anything below 20% of total time is rarely worth it until the big pieces are done.

speedup_overall = 1 / ((1 - p) + p/s)

p = fraction of time in the optimized part
s = speedup of that part

p=0.10, s=∞  → 1.11x
p=0.50, s=2  → 1.33x
p=0.80, s=2  → 1.67x
p=0.80, s=10 → 3.57x

Always compute this before starting work. It takes 30 seconds and frequently kills a bad idea.


4. Tiny example — the estimation habit#

PROPOSAL: "Let's write a fused RMSNorm kernel."

ESTIMATE FIRST:
  Profile says: rmsnorm kernels = 8% of decode step time.
  Fused version: ~3x faster on that op (6 passes → 2).
  Amdahl: 1/((1-0.08) + 0.08/3) = 1/(0.92+0.027) = 1.056x = 5.6% gain.
  Effort: 2 days.
  
  Compare: enabling CUDA graphs = 25% gain, 1 hour.
           Enabling FP8 = 80% gain, 3 days + validation.

DECISION: do CUDA graphs first, then FP8, then maybe the kernel.

That five-minute calculation reorders a quarter’s worth of work.


5. The methodology in detail#

Step 1 — Define the target#

BAD:  "make it faster"
GOOD: "reduce p95 ITL from 45 ms to 30 ms at batch 32, without increasing
       p95 TTFT above 500 ms or degrading MMLU by more than 0.5 points"

Without a target you can’t tell when you’re done, and you can’t trade off correctly. LLM optimization always involves a tradeoff; naming the constraint names the tradeoff.

Step 2 — Measure, at three levels#

LEVEL 1 — Request phases (application metrics)
  queue_wait, tokenize, prefill, decode, detokenize, network
  → tells you WHICH PHASE

LEVEL 2 — Timeline (Nsight Systems)
  kernel time vs gaps, sync points, transfers
  → tells you GPU vs CPU vs waiting

LEVEL 3 — Kernel (Nsight Compute)
  sm%, dram%, occupancy, sectors/request, stall reasons
  → tells you WHY a kernel is slow

Always go 1 → 2 → 3. Starting at level 3 optimizes kernels that don’t matter.

Step 3 — Classify#

Symptom                                    Bottleneck        Section
────────────────────────────────────────────────────────────────────
Low batch despite available memory         scheduling        V.09, VIII.03
High queue_wait                            capacity          XI.02
Gaps between kernels, CPU busy             launch/dispatch   VI.05, 08
dram% high, sm% low                        memory bandwidth  02-06, 11, 12
sm% high, dram% low                        compute           05, 09, 10
Both low                                   latency/occupancy VI.10
Long prefill blocking decode               scheduling        XIII.05
Client TTFT >> server TTFT                 network/proxy     II.08

Step 4-5 — Select and estimate#

Use the decision tree in this section’s README. Then estimate:

expected_gain = (fraction of time in the bottleneck)
              × (improvement factor for that bottleneck)
              adjusted by Amdahl

If the estimate is under 10%, ask whether there’s a bigger lever available first.

Step 6-7 — Apply and verify#

One change at a time. Two simultaneous changes with a 15% net gain could be +40% and -25%, and you’d never know.

Verify both performance and quality. For any precision change, run the Level 1-4 validation from Section IV.12. A 2x speedup with a 5% quality regression may be a bad trade — but you can only decide that if you measured both.

The optimization ranking (typical, not universal)#

RankOptimizationTypical gainEffortRisk
1Use an engine with continuous batching3-10xhourslow
2Fix scheduling (batch size, admission)1.5-4xdayslow
3Prefix caching (if prefixes are shared)1.3-5xhourslow
4Abort on client disconnect1.1-1.4xhourslow
5FP8/INT8 quantization1.5-2xdaysmedium (quality)
6CUDA graphs1.1-1.4xhourslow
7Right-size the model2-10xweekshigh (quality)
8Chunked prefill1.1-1.5x ITLhourslow
9INT4 weight-only (decode-heavy only)1.5-2.5xdaysmedium
10Speculative decoding1.3-2.5xweeksmedium
11Tensor parallelism (for latency)1.5-3x latencydayslow
12Custom kernels1.05-1.3xweeksmedium

Items 1-6 are config and hours of work. Most teams jump to 10-12. Do them in order.


6. Under the hood: the estimation formulas#

Keep these to hand:

Decode floor:      T = bytes_read / bandwidth
Prefill floor:     T = flops / peak_flops
Quantization gain: bytes_ratio (for memory-bound), flops_ratio (for compute-bound)
Batching gain:     min(target_batch/current_batch, ridge_point/current_intensity)
Fusion gain:       (passes_before / passes_after) on that op's time
Graph gain:        launch_overhead_total / step_time
Amdahl:            1/((1-p) + p/s)

With these you can estimate any optimization’s payoff in under five minutes.


7-9. Performance, production, mistakes#

Performance: the methodology itself has a performance benefit — it prevents you from spending three weeks on a 4% gain when a config flag gives 40%.

Production:

  • Keep a performance journal: baseline, change, measured gain, quality delta.
  • Automate the benchmark so anyone can run it.
  • Track achieved-vs-roofline as a regression signal.
  • Re-measure after every dependency upgrade; library heuristics change.

Mistakes:

  • Optimizing without measuring. The cardinal sin.
  • Multiple simultaneous changes. You learn nothing.
  • Not measuring quality. Half of inference optimizations trade accuracy.
  • Benchmarking with unrealistic workloads. Fixed lengths, no concurrency.
  • Ignoring Amdahl. Optimizing 5% of the time.
  • Optimizing the median when the SLO is on the tail.
  • Not re-measuring after the change. “It should be faster” is not data.

10. Hands-on exercise#

A. Build the benchmark. Write a reproducible benchmark for your system: realistic length distribution, realistic arrival pattern, reporting TTFT/ITL percentiles, throughput, and cost per million tokens. Everything else in this section depends on having this.

B. Full diagnosis. Take a running system and produce a complete diagnosis: phase breakdown, timeline analysis, top-5 kernel analysis, and a classification. Write it up in one page.

C. Estimate before doing. Pick three optimizations. For each, estimate the gain using the formulas in section 6. Then implement one and compare the actual gain to your estimate. How close were you? Adjust your estimation model.

D. Amdahl in practice. For your system, list every component with its share of time. For each, compute the maximum possible end-to-end gain from making it infinitely fast. Which are worth working on?

E. The journal. Start a performance-journal.md: date, baseline numbers, change made, measured result, quality delta, decision. Keep it for the rest of this section.


11. Interview questions#

  1. Walk me through your methodology for optimizing an inference system.
  2. A colleague reduced FLOPs by 40% with no speedup. What do you tell them?
  3. What is Amdahl’s law and how do you apply it before starting work?
  4. Rank the top five optimizations for a typical LLM service and justify the order.
  5. Why do you change one thing at a time?
  6. How would you decide between quantization and speculative decoding?
  7. What would you measure to know whether an optimization is safe to ship?

12. Further reading#

  • [FUNDAMENTAL] Amdahl, “Validity of the single processor approach” (1967)
  • [FUNDAMENTAL] Brendan Gregg, the USE method
  • [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” — analysis before implementation, done well
  • Next: 02 — Quantization overview

↑↓ navigate↵ openesc close