PidokuInfra

Case Studies

Advanced 1h 30m Difficulty 4/5 Topic 11 of 11

Prerequisites all of Section X

Eight worked investigations. Read each symptom, stop, and write down your hypotheses before reading the diagnosis. That exercise is the point of this file.


Case 1 — “TTFT is 4 seconds”#

SYMPTOM   p95 TTFT 4,100 ms (target 800 ms). p95 ITL 24 ms (fine).
          Throughput 3,900 tok/s. No errors.

STOP. What are your hypotheses?

MEASUREMENTS
  queue_wait p95:        3,650 ms      ← 89% of TTFT
  prefill p95:             380 ms
  tokenize p95:              3 ms
  running batch avg:        62 (max_num_seqs = 64)
  KV usage:                 94%
  arrival rate:             18 req/s
  Little's Law check: L = λW = 18 × 6.2 s = 112 in system, capacity 64
                      → 48 always waiting

DIAGNOSIS
  Genuinely at capacity. Not a tuning problem — an arithmetic one.
  The system can serve ~10 req/s at the current request shape; 18 are arriving.

WRONG FIXES
  ✗ optimize kernels (they're fine)
  ✗ increase max_num_seqs (KV is at 94%; you'd cause preemption)
  ✗ increase TP degree (doesn't help throughput; Section IX.12)

RIGHT FIXES, in order
  1. Check prefix caching. Chat workload, 1,100-token system prompt,
     hit rate was 0% (round-robin routing).
     → enable prefix-aware routing: prefill drops 68%, capacity +40%
  2. Enable FP8: capacity +80%
  3. Add replicas for the remainder.

RESULT  after (1) and (2): 18 req/s served at p95 TTFT 640 ms,
        with no new hardware.

Lesson: “at capacity” is a real diagnosis, but check whether you’re at capacity efficiently before buying hardware.


Case 2 — “ITL spikes to 300 ms randomly”#

SYMPTOM   p50 ITL 22 ms, p99 ITL 310 ms. Spikes are irregular.
          TTFT and throughput fine.

STOP. Hypotheses?

MEASUREMENTS
  Plotted ITL over time: spikes are ~450 ms and occur 3-8 times/minute.
  Correlated with: arrival of requests with prompt > 12,000 tokens.
  chunked prefill: DISABLED
  max_num_batched_tokens: 32,768

DIAGNOSIS
  Head-of-line blocking. A 12k-token prefill runs as one unit, occupying
  the GPU for ~450 ms, during which every decoding sequence stalls.

FIX
  --enable-chunked-prefill --max-num-batched-tokens 2048
  → the prefill is split into 6 chunks, each riding along with a decode step
  → each step is ~28 ms instead of 22 ms, but there are no 450 ms stalls

RESULT  p50 ITL 26 ms (slightly worse), p99 ITL 51 ms (6x better).
        p95 TTFT for long prompts: 480 → 620 ms (slightly worse).
        Overall throughput: +6%.

Lesson: a p99 that’s 14x the p50 is almost always a blocking event, not general slowness. Correlate the spikes with something.


Case 3 — “Throughput is a third of what we calculated”#

SYMPTOM   Predicted 3,000 tok/s from the roofline; measured 980 tok/s.
          Batch 32, 8B model, one H100.

STOP. Hypotheses?

MEASUREMENTS
  nsys: GPU busy 38% of wall clock
  API summary: 4,200 cudaLaunchKernel per step, 1 cudaStreamSynchronize
  py-spy: 41% of time in a custom logits processor (Python)
          which calls .cpu() on the logits every step
  CUDA graphs: disabled (--enforce-eager was in the config)

DIAGNOSIS
  Two compounding problems:
  1. --enforce-eager left in the config from a debugging session
  2. a per-step D2H copy + sync in a custom stopping-criterion callback

FIX
  1. Remove --enforce-eager                     → 980 → 1,510 tok/s
  2. Move the stopping check to the GPU          → 1,510 → 2,740 tok/s

RESULT  2,740 tok/s (91% of the roofline prediction).
        Total effort: 3 hours.

Lesson: a 3x gap from the roofline is almost never kernel efficiency. It’s CPU, launches, or synchronization. Check GPU-busy percentage first.


Case 4 — “OOM three times a day”#

SYMPTOM   Server OOMs 3-4 times daily, always near peak. Restarts cleanly.

STOP. Hypotheses?

MEASUREMENTS
  memory_summary() at failure: reserved 78.1 GB, allocated 71.4 GB
                               → 6.7 GB fragmentation, but that's not the
                                 whole story
  Correlation: every OOM within 2 s of a request with prompt > 28,000 tokens
  gpu_memory_utilization: 0.94
  chunked prefill: disabled
  Reproduced: 30k prompt + 40 decoding sequences → OOM

DIAGNOSIS
  Type 2 (Section X.06): activation spike during unchunked long prefill.
  The KV pool was sized using a profiling pass with max_num_batched_tokens
  tokens, but without chunking the actual prefill processed all 30k at once,
  producing an activation tensor 5x larger than profiled.

FIX
  1. --enable-chunked-prefill (bounds activation to max_num_batched_tokens)
  2. --gpu-memory-utilization 0.90 (headroom)
  3. Cap max input at 32,000 at the gateway with a clear 400

RESULT  Zero OOMs in 30 days. Bonus: p99 ITL improved 40% (same root cause
        as Case 2).

Lesson: OOMs correlate with something. Find the correlation before tuning memory settings.


Case 5 — “Adding GPUs made it slower”#

SYMPTOM   Went from TP=4 to TP=8 to increase throughput. Throughput per GPU
          dropped 35%; total throughput up only 30%.

STOP. Hypotheses?

MEASUREMENTS
  nvidia-smi topo -m:
     GPU0-3: NV18 to each other
     GPU4-7: NV18 to each other
     GPU0-4: SYS                       ← across the CPU interconnect
  nsys with NCCL trace: AllReduce time 4.2 ms of a 9.1 ms step (46%)
  At TP=4 (within one NVLink island): AllReduce 0.6 ms of 12.8 ms (5%)

DIAGNOSIS
  The node has two NVLink islands, not one NVSwitch fabric.
  TP=8 forces AllReduce across the CPU interconnect.

FIX
  Revert to TP=4, run TWO instances (DP=2 × TP=4), one per NVLink island.
  Pin each with numactl to its NUMA node.

RESULT  Total throughput +95% vs the original TP=4 single instance,
        vs +30% for TP=8. Same 8 GPUs.

Lesson: always check nvidia-smi topo -m before choosing a TP degree. This exact mistake is common on non-DGX hardware.


Case 6 — “The GPUs are at 100%, we need more”#

SYMPTOM   Capacity request for 64 additional GPUs. Justification:
          "all 64 existing GPUs at 100% utilization."

STOP. Hypotheses?

MEASUREMENTS
  DCGM_FI_PROF_DRAM_ACTIVE:  0.24
  DCGM_FI_PROF_SM_ACTIVE:    0.97
  avg_running_batch_size:    7  (max_num_seqs = 96)
  kv_cache_usage_ratio:      0.09
  queue depth:               0
  replica count:             16, each receiving ~4 req/s

DIAGNOSIS
  Not capacity-limited at all. Traffic is spread too thinly across too many
  replicas; each reads the full model per step to serve 7 sequences
  (Section IX.12, way 5).

FIX
  1. Consolidate 16 → 5 replicas
  2. Prefix-aware routing across them
  3. Keep 1 spare for headroom

RESULT  avg batch 7 → 26. DRAM_ACTIVE 0.24 → 0.71. Throughput per GPU 3.4x.
        44 of 64 GPUs freed. No purchase.

Lesson: nvidia-smi utilization is not a capacity signal (Section X.04). Require DRAM_ACTIVE and avg_running_batch_size in every capacity request.


Case 7 — “Quality dropped after a deploy, but all tests passed”#

SYMPTOM   User complaints about "worse answers" starting Tuesday.
          Error rate, latency, throughput all normal. All CI tests green.

STOP. Hypotheses?

MEASUREMENTS
  Deploy log: Tuesday's release changed the serving precision from BF16 to
              INT4 (W4A16, GPTQ, group 128) for cost reasons.
  Offline eval at deploy time: MMLU -0.4%, perplexity +0.08. Accepted.
  Output length distribution: mean 218 → 341 tokens (+56%)   ← the signal
  Regeneration rate: 4.1% → 11.3%
  Structured output validity: 99.2% → 91.7%

DIAGNOSIS
  INT4 quantization degraded instruction-following and long-form coherence
  in a way that MMLU and perplexity did not capture (Section VII.06).
  The model rambles and violates JSON schemas more often.

FIX
  1. Immediate: roll back to BF16 (old checkpoint still cached — good)
  2. Re-evaluate with the Level 3 tests (Section IV.12): long generations,
     instruction following, structured output validity
  3. Deploy FP8 instead: -0.1% MMLU, output length unchanged,
     structured validity 99.1%, and 1.9x throughput.

RESULT  Cost target met with FP8. Added output-length distribution and
        structured-validity to the canary guards.

Lesson: output length distribution is a cheap, sensitive early quality signal. Perplexity and MMLU are not sufficient validation for a precision change.


Case 8 — “Client sees 8 seconds, server says 200 ms”#

SYMPTOM   Users report the response appears all at once after ~8 seconds.
          Server-side TTFT metric: p95 210 ms. Server-side E2E: p95 7.9 s.

STOP. Hypotheses?

MEASUREMENTS
  curl -N -w '%{time_starttransfer}' through the production ingress: 7.82 s
  curl -N directly to the pod IP: 0.19 s                    ← the answer
  nginx ingress config: proxy_buffering not set → defaults to ON

DIAGNOSIS
  The ingress buffers the entire SSE response before forwarding.
  Server metrics are correct; the user experience is not.

FIX
  proxy_buffering off;
  proxy_cache off;
  proxy_read_timeout 3600s;
  add_header X-Accel-Buffering no;
  (and verify the CDN and WAF layers too)

RESULT  Client TTFT p95: 7.82 s → 0.26 s.
        Added a synthetic client-side TTFT check through the production
        ingress to monitoring.

Lesson: always measure from a real client through the real path. Server-side metrics cannot see this class of problem (Section II.08).


The patterns across all eight#

1. DECOMPOSE FIRST. Six of eight were resolved by phase decomposition
   or a single correlation, before any deep profiling.

2. CHECK THE CONFIGURATION. Cases 3, 4, 8 were configuration errors.
   Cases 2 and 5 were configuration choices made without measurement.

3. MEASURE FROM THE OUTSIDE. Case 8 is invisible from the inside.

4. THE OBVIOUS METRIC LIES. Case 6: GPU utilization. Case 7: MMLU.

5. CORRELATE ANOMALIES WITH SOMETHING. Cases 2 and 4 were both
   "correlate the spikes with the request shape."

6. THE FIX IS OFTEN A FLAG. Cases 2, 3, 4, 8. Total engineering time
   for those four: about a day. Combined impact: large.

7. AMDAHL. In every case, the right fix addressed the dominant term.
   In Case 3, someone could have spent a month on kernels for 5%.

Hands-on exercise#

A. Predict before reading. Re-read each case’s symptom and measurements, covering the diagnosis. Write your hypothesis. How often were you right? Which cases fooled you and why?

B. Reproduce one. Pick Case 2, 3, or 4 and reproduce it deliberately on a test system. Confirm the symptom, apply the fix, measure the improvement.

C. Write your own. Take a real performance problem you’ve encountered (in any system) and write it up in this format: symptom, measurements, diagnosis, wrong fixes, right fix, result, lesson. This is the format for your team’s incident postmortems.

D. Build the checklist. From the eight cases, derive a “first 15 minutes” checklist for an inference performance incident. What do you measure, in what order?

E. The capacity request template. Write the template for a GPU capacity request at your organization, requiring the evidence that Cases 1 and 6 show is necessary.


Interview questions#

  1. p95 TTFT is 4 s, p95 ITL is fine. Walk me through your investigation.
  2. p99 ITL is 14x p50, irregularly. What’s your first hypothesis?
  3. Measured throughput is a third of the roofline prediction. What do you check first?
  4. A team requests 64 more GPUs citing 100% utilization. What evidence do you require?
  5. Quality dropped after a deploy but all tests passed. How do you find it, and how do you prevent it next time?
  6. Client-side latency is 40x server-side latency. Diagnose.
  7. Adding GPUs reduced per-GPU throughput 35%. What happened?

Further reading#

↑↓ navigate↵ openesc close