★ The capstone. This is the system design interview question for inference engineering roles, and it is a real design exercise. Work through it yourself before reading the solution.
1. The brief#
REQUIREMENTS
Model: Llama-3-70B-Instruct (or equivalent)
Users: 10,000 concurrent active conversations at peak
Workload: chat. Average 1,200 input tokens (with history),
350 output tokens. p95 input 6,000, p95 output 1,200.
SLO: p95 TTFT < 500 ms for inputs ≤ 4,000 tokens
p95 ITL < 40 ms
99.9% availability
Budget: fixed at 64× H100 GPUs (8 nodes of 8)
Constraints: single region initially; NVSwitch nodes; NDR InfiniBand
QUESTION: design the serving architecture. Does it meet the SLO?
If not, what do you change?Stop here. Sketch your answer before reading on.
2. Step 1 — is it even possible?#
Start with the arithmetic, before any architecture.
TRAFFIC DEMAND
10,000 concurrent conversations.
A "concurrent user" in chat is not continuously generating — they read,
think, and type. Estimate: a user is actively generating ~8% of the time
(a 350-token response at 25 tok/s = 14 s of generation, then ~3 minutes
of reading and typing).
→ concurrently GENERATING: 10,000 × 0.08 = 800 sequences
→ request rate: 10,000 users / 180 s per turn = 55.6 req/s
peak output tokens/sec: 55.6 × 350 = 19,460 tok/s
peak input tokens/sec: 55.6 × 1,200 = 66,720 tok/sThat “8% duty cycle” assumption is the most important number in the problem, and it’s the one candidates most often get wrong by assuming 10,000 simultaneous generations. State it explicitly and justify it.
CAPACITY AVAILABLE
64 H100s. At FP8, weights are 70.6 GB.
Option A: TP=8, 8 instances (one per node)
per GPU: 8.8 GB weights, 80 - 8.8 - 5 = 66 GB for KV
KV per token per GPU (TP=8, GQA-8, FP8 KV): 320 KiB / 8 / 2 = 20 KiB
Sequences per node at 4,000 avg context:
66 GB / (20 KiB × 4000) = 66e9 / 8.19e7 = 806 sequences per node
× 8 nodes = 6,448 concurrent sequences of KV capacity
We need 800 concurrently generating. Memory is NOT the constraint. ✓Memory is not the binding constraint here — a useful realization that redirects the design.
THROUGHPUT CHECK
Decode, per node, at batch 100 (800 / 8 nodes):
bytes per step per GPU = 8.8 GB (weights) + 100 × 4000 × 20 KiB
= 8.8 + 8.0 = 16.8 GB
T_theoretical = 16.8 / 3350 = 5.0 ms
+ TP AllReduce (~20%) + 70% efficiency → ~9.0 ms
output tok/s per node = 100 / 0.009 = 11,111
× 8 nodes = 88,888 tok/s of decode capacity
We need 19,460. Decode capacity is 4.6x demand. ✓
ITL = 9.0 ms ≪ 40 ms SLO ✓✓PREFILL CHECK ← this is where it gets interesting
Prefill throughput per node (FP8, 8×H100, TP=8):
FLOPs per prefill token ≈ 2 × 70e9 = 140 GFLOP
8 GPUs × 990 TFLOP/s × 0.5 achieved = 3,960 TFLOP/s
→ 3,960e12 / 140e9 = 28,286 prefill tokens/sec per node
× 8 nodes = 226,000 tok/s
We need 66,720. Prefill capacity is 3.4x demand. ✓
BUT: prefill and decode SHARE the GPUs.
prefill time fraction: 66,720 / 226,000 = 29.5%
decode time fraction: 19,460 / 88,888 = 21.9%
total: 51.4% utilizationVerdict: the fleet is adequate with roughly 2x headroom. Now design it properly.
3. Step 2 — the architecture#
┌──────────────────┐
clients ──────────► │ API GATEWAY │ auth, rate limit (tokens +
│ (CPU nodes, 6) │ concurrency), validation,
└────────┬─────────┘ tokenization, SSE streaming
│
┌────────▼─────────┐
│ ROUTER │ prefix-aware + least-KV-loaded
│ (CPU nodes, 3) │ session affinity by conv_id
└────────┬─────────┘
│
┌────────┬────────┬──────┼──────┬────────┬────────┬────────┐
▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼
┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐
│node0│ │node1│ │node2│ │node3│ │node4│ │node5│ │node6│ │node7│
│TP=8 │ │TP=8 │ │TP=8 │ │TP=8 │ │TP=8 │ │TP=8 │ │TP=8 │ │TP=8 │
└─────┘ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘
general pool (6 nodes) long-ctx pool spare/canary
(1 node) (1 node)Key decisions and their justifications:
DECISION JUSTIFICATION
TP=8, DP=8 (not TP=16 or 64) TP=8 fills one NVSwitch domain (Section IX.03).
Higher TP crosses nodes → unusable (IX.09).
DP for throughput (IX.02).
FP8 quantization 2x throughput and memory; well-validated;
Hopper native (VII.05). Without it, weights are
141 GB → 17.6 GB/GPU, halving KV capacity and
doubling decode time. FP8 is what makes the
budget work.
FP8 KV cache 2x concurrency (VII.12). Not strictly needed
here (memory isn't binding) but reduces KV
bandwidth, improving ITL at high batch.
Prefix caching ON Chat with conversation history: turn N re-prefills
turns 1..N-1. Expected hit rate 70-85%.
→ cuts prefill demand from 66,720 to ~15,000 tok/s.
THIS IS THE LARGEST SINGLE WIN.
Prefix-aware routing Without it, prefix cache hit rate is ~1/8.
Session affinity by conversation_id (VIII.06).
Chunked prefill ON p95 input is 6,000 tokens; unchunked, a 6k prefill
is ~210 ms of GPU, spiking ITL for everyone (V.03).
max_num_batched_tokens = 2048.
Separate long-context pool A 32k+ request consumes 8x the KV and blocks
prefill. Segregating protects the general pool's
p99 (VIII.06).
One spare node N+1 for the 99.9% SLO, and it serves as the
canary target for rollouts (XI.07).
CUDA graphs ON 20-30% at these batch sizes (VII.08).4. Step 3 — the configuration#
vllm serve meta-llama/Meta-Llama-3-70B-Instruct \
--quantization fp8 \
--kv-cache-dtype fp8 \
--tensor-parallel-size 8 \
--max-model-len 8192 \
--max-num-seqs 160 \
--max-num-batched-tokens 2048 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-chunked-prefill \
--scheduling-policy priority \
--port 8000JUSTIFICATION FOR EACH TUNED VALUE
max-model-len 8192
p95 input 6,000 + p95 output 1,200 = 7,200, +14% headroom.
NOT 128k — that would inflate block tables and reduce usable blocks
for no benefit (VIII.04). Long-context requests go to the other pool.
max-num-seqs 160
Memory allows ~806 at 4k avg context. But:
- we only need 800/6 = 133 per general-pool node
- 160 gives 20% headroom above expected
- and keeps ITL comfortable: at 160, step bytes = 8.8 + 12.8 = 21.6 GB
→ 6.4 ms theoretical, ~11.5 ms real. Still ≪ 40 ms SLO.
Deliberately well below the memory maximum (V.09, preemption avoidance).
max-num-batched-tokens 2048
From the ITL formula (VIII.04): prefill rate ~28,000 tok/s per node
→ 2048 tokens = 73 ms of prefill work per step... too much.
Recompute: we want the prefill chunk to add < 15 ms to a step.
28,000 tok/s → 15 ms = 420 tokens. But that's very small chunks.
Compromise: 2048 tokens adds ~73 ms to ONE step every few steps,
but chunked prefill spreads it. Measured: p95 ITL 24 ms. Acceptable.
→ TUNE THIS EMPIRICALLY. Start at 2048, measure p95 ITL, adjust.
gpu-memory-utilization 0.90
Not 0.95: the p99 input is much larger than p95, and we need
activation headroom for chunked prefill's largest chunk (XI/X.06).5. Step 4 — does it meet the SLO?#
TTFT BUDGET (target: p95 < 500 ms for inputs ≤ 4,000 tokens)
network RTT (regional) 25 ms
gateway (auth, rate limit, route) 12 ms
tokenization (4,000 tokens, Rust) 4 ms
queue wait (at 51% utilization) ~90 ms ← from queueing theory
prefill:
with 78% prefix cache hit rate,
effective new tokens ≈ 880
880 / 28,000 tok/s 31 ms
but chunked, interleaved with decode: ~55 ms
first sample 2 ms
SSE first write 6 ms
─────────────────────────────────────────────────
TOTAL 194 ms
p95 (with variance) ~340 ms ✓ under 500 ms
ITL (target: p95 < 40 ms)
decode step at batch 160: 11.5 ms
+ chunked prefill interleaving: +8 ms average
+ scheduler and sampling overhead: +3 ms
─────────────────────────────────────────────────
p50 ~19 ms
p95 (steps containing a prefill chunk) ~28 ms ✓ under 40 ms
AVAILABILITY (target 99.9%)
8 nodes, N+1 → survives one node failure with degradation
→ 7 nodes serving 51% × 8/7 = 59% utilization. Still fine.
Deployment: roll one node at a time; 7 nodes serve during each roll.
✓ achievable, given fast failure detection and restart (XI.05)Verdict: the SLO is met with headroom. Now stress it.
6. Step 5 — what breaks it#
FAILURE MODE IMPACT MITIGATION
Prefix cache hit rate drops to 20% prefill demand 3.5x → still fits
(e.g. a product change adds a → utilization 78% (barely).
unique timestamp to prompts) Monitor hit rate;
alert on drops.
Traffic doubles utilization 102% → degradation ladder;
need 8 more nodes.
Order lead time!
Input lengths double (product adds prefill demand 2x → utilization 81%.
retrieval) Tighter but OK.
A burst of 32k-context requests long-ctx pool → route them there;
saturates general pool
unaffected. ✓
(this is why we
segregated)
One node fails 7 nodes, 59% util → fine ✓
Two nodes fail 6 nodes, 68% util → fine, tighter
FP8 quality regression discovered must revert to BF16 → weights 141 GB
→ 17.6 GB/GPU
→ decode 2x slower
→ utilization 95%
→ SLO AT RISK
← THE REAL RISKThe last row is the one to plan for. The design depends on FP8. If FP8 proves unacceptable for quality, the budget does not support BF16 at this SLO. Mitigations: validate FP8 thoroughly before committing (Section IV.12), and have INT8 W8A8 as a fallback (similar speedup, different quantization method, so a quality issue with one may not affect the other).
State this dependency explicitly in the design. A design whose critical assumption is unstated is a design that will fail surprisingly.
7. Step 6 — cost#
64 H100s at $2.50/GPU-hour reserved = $160/hour = $1,401,600/year
Tokens per year:
19,460 output tok/s × 0.51 duty (peak vs average over 24h) × 3.15e7 s
= 3.13e11 output tokens/year
Cost per million output tokens:
$1,401,600 / 313,000 = $4.48
Hmm — that's high. Why?
Because we're sized for PEAK and the average is much lower.
Average utilization across 24 hours: ~30% (peak is 51%, trough is ~10%)
Improvements:
- run batch/offline workloads in the trough → +$0 marginal, better
amortization
- reduce the fleet and accept degradation at peak
- use spot instances for 2 of the 8 nodesThis is the honest answer: the fleet is sized for peak, and peak is a fraction of the day. Filling the trough with batch work is the single largest cost lever remaining.
8. The complete answer, summarized#
ARCHITECTURE
8 nodes × 8 H100, TP=8 within nodes, DP=8 across
6 general pool, 1 long-context pool, 1 spare/canary
CPU gateway tier (6 nodes) + router tier (3 nodes)
CONFIGURATION
FP8 weights and KV cache, prefix caching, chunked prefill,
CUDA graphs, max_model_len 8192, max_num_seqs 160,
max_num_batched_tokens 2048 (tuned empirically)
ROUTING
prefix-aware with session affinity by conversation_id
long-context requests (> 16k) segregated to their own pool
least-KV-loaded within the affinity group
RESULT
p95 TTFT ~340 ms (target 500) ✓
p95 ITL ~28 ms (target 40) ✓
utilization 51% at peak → headroom for 1.9x growth
availability 99.9% with N+1 ✓
cost ~$4.48/M output tokens (improvable by filling the trough)
CRITICAL DEPENDENCIES (state these!)
1. FP8 quality must be acceptable — validate before committing
2. Prefix cache hit rate ≥ 60% — monitor and alert
3. Prompt structure must remain cache-friendly — a product change
that adds per-request content to the prompt prefix would break it
4. Long-context traffic must stay < 5% of requests
WHAT I'D MEASURE FIRST IN PRODUCTION
prefix cache hit rate, average running batch, p95 ITL,
queue wait p95, cost per million tokens9. Variations to practice#
Work each of these; they’re the follow-up questions.
A. "Budget is 32 GPUs, not 64."
→ utilization 102%. Options: accept degradation, INT4 weights
(decode-heavy so it helps), reduce max context, tier the SLO,
or route overflow to a smaller model. Show the arithmetic for each.
B. "The SLO is p95 TTFT < 200 ms."
→ the network and gateway alone are 37 ms; prefill for 4k tokens
is unavoidable. Need: higher prefix hit rate, smaller model for
the first token, or accept that only cached-prefix requests
meet it. Consider a speculative/draft model for TTFT.
C. "Add 100 fine-tuned variants."
→ multi-LoRA (VIII.09). One base model, adapters swapped per request.
5-15% overhead, ~4 GB for 100 adapters. Adapter-aware routing.
D. "Users are global."
→ multi-region (XI.06). Latency is the driver. Size each region for
its traffic + degraded failover. Sessions stay regional.
E. "Peak is 10x average, spiky."
→ autoscaling can't help (cold start). Need headroom for the spike,
or a degradation ladder, or overflow to an external provider.
F. "What if we used a 405B model?"
→ FP8: 405 GB. TP=8 gives 50.6 GB/GPU — fits, but only 25 GB for KV.
Decode is 5.7x slower per token. At the same fleet: utilization
would far exceed 100%. Would need ~4x the GPUs or a smaller model.
G. "Can we use A100s instead?"
→ 2.04 TB/s vs 3.35: decode 1.64x slower. No FP8 (INT8 instead).
Need ~1.7x the nodes. But A100s may be cheaper per hour —
compute $/M tokens for both.10. How to answer this in an interview#
1. CLARIFY THE WORKLOAD FIRST.
"10,000 concurrent users" is ambiguous. Ask about the duty cycle,
the length distributions, and what "concurrent" means.
← this alone distinguishes strong candidates
2. DO THE ARITHMETIC BEFORE THE ARCHITECTURE.
Is it even possible? Which resource binds?
3. STATE ASSUMPTIONS EXPLICITLY.
Duty cycle, prefix hit rate, achieved efficiency fraction.
4. JUSTIFY EVERY CHOICE FROM A CONSTRAINT.
"TP=8 because that's the NVSwitch domain and TP=16 crosses nodes."
Not "TP=8 because that's typical."
5. IDENTIFY THE CRITICAL DEPENDENCY.
Here: FP8. Say what breaks if it fails.
6. SAY WHAT YOU'D MEASURE.
A design without a validation plan is a guess.
7. BE HONEST ABOUT WHAT YOU DON'T KNOW.
"I'd measure the prefix hit rate before committing to this;
if it's below 50% the numbers change materially."11. Hands-on exercise#
A. Do it yourself. Before re-reading the solution, work the whole problem from the brief. Compare your answer. Where did you differ, and was your reasoning sound?
B. Work the variations. Do all seven variations in section 9 with full arithmetic.
C. Build the calculator. Extend your Section V.15 and IX.09 calculators into a single tool that takes the brief and produces the design. Validate it against this solution.
D. Stress it. For the design in section 8, compute the utilization under each of the failure modes in section 6. Which is the binding one?
E. Present it. Write the design as a one-page document with the architecture diagram, configuration, SLO analysis, dependencies, and validation plan. This is the actual deliverable in a real job.
12. Interview questions#
- Design a serving architecture for a 70B model, 10,000 concurrent users, 64 H100s.
- What’s the first question you’d ask about “10,000 concurrent users”?
- Which resource binds in this design — memory, compute, or bandwidth?
- Why TP=8 and DP=8 rather than TP=64?
- What’s the critical dependency in your design, and what happens if it fails?
- Your cost is $4.48/M tokens. Why so high, and how would you reduce it?
- The budget is halved. What do you change?
13. Further reading#
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” — the analytical approach, applied at scale
- All of Sections V, VIII, IX, and XI of this curriculum
- Next: Section XII — Inference Platform Engineering