PidokuInfra

Capacity Planning

Advanced 1h 30m Difficulty 4/5 Topic 02 of 11

Prerequisites V.15, 01


1. The question#

How many GPUs do I need? And its harder cousin: how many will I need in six months, and when do I have to order them?

Section V.15 did the per-node arithmetic. This file adds the production factors: traffic modeling, headroom, redundancy, growth, and lead times.


2. The formula#

GPUs = ceil(
    peak_demand
    ÷ per_node_capacity_at_SLO
    ÷ target_utilization
    × redundancy_factor
    × growth_headroom
) × GPUs_per_node

where:
  peak_demand           tokens/sec at your busiest 5-minute window
  per_node_capacity     measured, at your SLO (not peak throughput!)
  target_utilization    0.60-0.75 for latency-sensitive (the queueing wall)
  redundancy_factor     1 + (1/nodes) for N+1, or 1.5 for N+50%
  growth_headroom       1.2-1.5 depending on your planning horizon

Two terms people omit and shouldn’t: target_utilization (Section I.05 — latency explodes near 100%) and growth_headroom (GPU lead times are months).


3. Step 1 — model the traffic#

This is the input everything else depends on, and it’s usually the weakest part of a capacity plan.

FROM PRODUCT / BUSINESS
  daily active users
  interactions per user per day
  peak-to-average ratio
  growth rate

FROM MEASUREMENT (structured logs, Section X.10)
  prompt length distribution (p50, p90, p95, p99)
  output length distribution
  request rate by hour of day and day of week
  the shape of the peak (sharp spike or broad plateau?)

DERIVED
  peak_requests_per_second
  peak_prefill_tokens_per_second
  peak_decode_tokens_per_second
WORKED:
  50,000 DAU × 8 sessions/day × 6 turns = 2.4M requests/day
  Peak hour = 14% of daily volume (measured, typical for consumer)
     = 336,000 requests in the peak hour = 93 req/s
  Peak 5-minute burst = 1.4× the peak hour rate = 130 req/s
  
  p50 prompt 480 tokens, p95 3,200; mean 890 (the tail matters!)
  p50 output 190 tokens, p95 1,100; mean 320
  
  peak prefill: 130 × 890 = 115,700 tok/s
  peak decode:  130 × 320 =  41,600 tok/s

Use the MEAN for capacity, not the median. The mean is what determines total work; the tail pulls it well above the median. Using p50 will underprovision by 40-80% for a heavy-tailed distribution.


4. Step 2 — measure per-node capacity at the SLO#

Not peak throughput. Capacity at the SLO.

Run the benchmark from Section X.07, open-loop, at increasing arrival rates.
Find the rate at which goodput starts falling — that's your capacity.

  arrival rate   p95 TTFT   p95 ITL   goodput
      4 req/s      280 ms     22 ms     100%
      8            340 ms     26 ms     100%
     12            480 ms     31 ms     99.4%
     14            690 ms     38 ms     98.1%
     16          1,240 ms     44 ms     91.3%    ← knee
     18          3,100 ms     52 ms     71.0%
     
  → capacity at SLO (goodput ≥ 97%) = 14 req/s per node
  → peak throughput would say 18+, and would be wrong

The gap between “peak throughput” and “capacity at SLO” is typically 25-40%. Planning with the former guarantees SLO violations at peak.


5. Step 3 — assemble the plan#

peak demand:               130 req/s
capacity per node at SLO:   14 req/s
                          ──────────
nodes for peak:            9.3 → 10

target utilization 0.70:   10 / 0.70 = 14.3 → 15 nodes
                           (running 10 nodes' worth of work on 15 nodes'
                            capacity keeps you off the queueing wall)

N+1 redundancy:            16 nodes
growth headroom 1.3
  (6-month horizon,
   30% expected growth):   16 × 1.3 = 20.8 → 21 nodes

ANSWER: 21 nodes × 8 GPUs = 168 GPUs

Sanity-check it against the token arithmetic:

21 nodes × 14 req/s × 320 output tokens = 94,080 output tok/s of capacity
Peak demand: 41,600 output tok/s
Ratio: 2.26x

Is 2.26x too much? Decompose:
  1/0.70 (utilization) × 1.07 (N+1) × 1.3 (growth) = 1.99
  plus the rounding: 2.26
  
→ Consistent. The headroom is deliberate, not accidental.

Always do this cross-check. If the ratio is much larger than the product of your explicit factors, you’ve double-counted something.


6. The factors, justified#

Target utilization (the biggest and most-argued factor)#

From queueing theory (Section I.05): W_queue ≈ W_service × ρ/(1-ρ)

  ρ = 0.60 → queue adds 1.5× service time
  ρ = 0.70 → 2.3×
  ρ = 0.80 → 4×
  ρ = 0.90 → 9×
  ρ = 0.95 → 19×

For a latency SLO with 300 ms of queueing budget and 200 ms service:
  queue budget / service = 1.5  →  ρ ≈ 0.60

For batch workloads with no latency SLO: ρ = 0.95 is fine.

This factor is where the money is, and it’s where the argument happens. “Why are we running at 70%?” is answered by the queueing formula and your latency budget, not by preference.

Redundancy#

N+1:     survive one node failure. redundancy = (N+1)/N
N+2:     survive two, or one failure during a deployment
2N:      survive an AZ loss (Section XI.06)

For GPU fleets: N+1 within an AZ, and enough capacity in a second AZ to
serve degraded traffic if the first is lost.

Growth headroom and lead time#

GPU procurement lead times: weeks to quarters, depending on the part and
your relationship with the supplier.

Plan: headroom = expected_growth_over(lead_time + safety_margin)

  lead time 12 weeks, growth 8%/month → 12 weeks = 2.8 months
  → 1.08^2.8 = 1.24
  → plus safety: 1.3

Order before you need it. Capacity planning that produces an order date after the need date is a plan to have an outage.


7. What changes the plan#

Re-run capacity planning when any of these change materially:

INPUT                          EFFECT
prompt length distribution     linear-to-quadratic on prefill
output length distribution     linear on decode
peak-to-average ratio          linear on peak sizing
model size or architecture     large, nonlinear
precision (quantization)       1.5-3x
prefix cache hit rate          large for prefill-heavy workloads
SLO thresholds                 large (via target utilization)
context length limits          large (via concurrency)
new product features           unpredictable — model them explicitly

Prompt length distribution drift is the most common surprise. A product change that adds retrieved context to every request can double your prefill load overnight without any change in request rate.


8. Producing the artifact#

A capacity plan is a document, not a number. It should contain:

1. TRAFFIC MODEL          assumptions, sources, measured distributions
2. UNIT CAPACITY          measured, with the benchmark methodology
3. THE CALCULATION        every factor, with justification
4. THE ANSWER             GPU count, instance types, timeline
5. COST                   monthly, and cost per million tokens
6. SENSITIVITY            what if traffic is 2x? what if prompts double?
7. TRIGGERS               what measurements would invalidate this plan
8. REVIEW DATE

Section 6 (sensitivity) is what makes it useful. A plan with a single number is fragile; a plan with “if peak exceeds 160 req/s we need 26 nodes, order by March” is actionable.


9. Production implications#

  • Base the traffic model on measurement. Structured per-request logs (Section X.10) are the source. Product forecasts are an input, not the input.
  • Measure unit capacity at the SLO, with a realistic benchmark.
  • Justify the utilization target from the queueing formula and your latency budget.
  • Track actual vs planned monthly. A plan you don’t check is a guess.
  • Set triggers: “if p95 prompt length exceeds 4,000 tokens, re-plan.”
  • Account for lead time. Order date = need date − lead time − safety.
  • Include the efficiency levers in the plan: “this assumes FP8 and 60% prefix cache hit rate; without them, add 8 nodes.”

10. Common mistakes#

Planning with peak throughput instead of capacity at SLO. 25-40% underprovisioned.

Using median instead of mean lengths. 40-80% underprovisioned for heavy-tailed traffic.

Planning for 90%+ utilization with a latency SLO. The queueing wall.

Forgetting growth and lead time. The plan is correct and arrives late.

No sensitivity analysis. One assumption changes and the plan is void.

Not accounting for the efficiency levers. Planning at FP16 when you’ll deploy FP8 means 2x overprovisioning.

Planning once. Traffic shape drifts continuously.


11. Hands-on exercise#

A. Build the traffic model. From real (or realistic synthetic) request logs, extract: the length distributions, the hourly rate profile, the peak-to-average ratio, and the peak 5-minute burst. Plot each.

B. Measure unit capacity at the SLO. Using your Section X.07 harness, run open-loop at increasing rates. Find the goodput knee. Compare to peak throughput — how big is the gap?

C. Produce the plan. Write the full capacity plan document from section 8 for a service you know. Include the sensitivity analysis.

D. Sensitivity. Vary each input ±50% and compute the effect on the GPU count. Rank the inputs by impact. Which most deserves careful measurement?

E. The efficiency lever. Compute the plan with and without: FP8, prefix caching, chunked prefill. How many GPUs does each lever save? Express each as annual dollars.

F. Cross-check. For your plan, verify the token arithmetic matches the request arithmetic (section 5’s sanity check).


12. Interview questions#

  1. Walk me through capacity planning for an LLM service.
  2. Why plan with capacity-at-SLO rather than peak throughput?
  3. Why target 70% utilization? Justify it quantitatively.
  4. Why use the mean rather than the median request length?
  5. What would invalidate a capacity plan, and how would you detect it?
  6. How does prefix caching change your capacity requirement?
  7. Your plan says 21 nodes. How do you sanity-check that?

13. Further reading#

  • [FUNDAMENTAL] Google SRE Book, capacity planning and demand forecasting chapters
  • [FUNDAMENTAL] Little’s Law and basic queueing theory
  • [ESTABLISHED] Section V.15 of this curriculum — the per-node arithmetic
  • Next: 03 — Cost per token

↑↓ navigate↵ openesc close