PidokuInfra

Scenario: 70B Model, 10,000 Concurrent Users, Fixed Budget

Advanced Expert 2h Difficulty 5/5 Topic 11 of 11

Prerequisites all of Sections V, VIII, IX, XI

★ The capstone. This is the system design interview question for inference engineering roles, and it is a real design exercise. Work through it yourself before reading the solution.


1. The brief#

REQUIREMENTS
  Model:        Llama-3-70B-Instruct (or equivalent)
  Users:        10,000 concurrent active conversations at peak
  Workload:     chat. Average 1,200 input tokens (with history),
                350 output tokens. p95 input 6,000, p95 output 1,200.
  SLO:          p95 TTFT < 500 ms for inputs ≤ 4,000 tokens
                p95 ITL  < 40 ms
                99.9% availability
  Budget:       fixed at 64× H100 GPUs (8 nodes of 8)
  Constraints:  single region initially; NVSwitch nodes; NDR InfiniBand

QUESTION: design the serving architecture. Does it meet the SLO?
          If not, what do you change?

Stop here. Sketch your answer before reading on.


2. Step 1 — is it even possible?#

Start with the arithmetic, before any architecture.

TRAFFIC DEMAND
  10,000 concurrent conversations.
  A "concurrent user" in chat is not continuously generating — they read,
  think, and type. Estimate: a user is actively generating ~8% of the time
  (a 350-token response at 25 tok/s = 14 s of generation, then ~3 minutes
  of reading and typing).

  → concurrently GENERATING: 10,000 × 0.08 = 800 sequences
  → request rate: 10,000 users / 180 s per turn = 55.6 req/s

  peak output tokens/sec: 55.6 × 350 = 19,460 tok/s
  peak input tokens/sec:  55.6 × 1,200 = 66,720 tok/s

That “8% duty cycle” assumption is the most important number in the problem, and it’s the one candidates most often get wrong by assuming 10,000 simultaneous generations. State it explicitly and justify it.

CAPACITY AVAILABLE
  64 H100s. At FP8, weights are 70.6 GB.
  
  Option A: TP=8, 8 instances (one per node)
     per GPU: 8.8 GB weights, 80 - 8.8 - 5 = 66 GB for KV
     KV per token per GPU (TP=8, GQA-8, FP8 KV): 320 KiB / 8 / 2 = 20 KiB
     
  Sequences per node at 4,000 avg context:
     66 GB / (20 KiB × 4000) = 66e9 / 8.19e7 = 806 sequences per node
     × 8 nodes = 6,448 concurrent sequences of KV capacity
     
  We need 800 concurrently generating. Memory is NOT the constraint. ✓

Memory is not the binding constraint here — a useful realization that redirects the design.

THROUGHPUT CHECK
  Decode, per node, at batch 100 (800 / 8 nodes):
     bytes per step per GPU = 8.8 GB (weights) + 100 × 4000 × 20 KiB
                            = 8.8 + 8.0 = 16.8 GB
     T_theoretical = 16.8 / 3350 = 5.0 ms
     + TP AllReduce (~20%) + 70% efficiency → ~9.0 ms
     
     output tok/s per node = 100 / 0.009 = 11,111
     × 8 nodes = 88,888 tok/s of decode capacity
     
  We need 19,460. Decode capacity is 4.6x demand. ✓

  ITL = 9.0 ms ≪ 40 ms SLO ✓✓
PREFILL CHECK  ← this is where it gets interesting
  Prefill throughput per node (FP8, 8×H100, TP=8):
     FLOPs per prefill token ≈ 2 × 70e9 = 140 GFLOP
     8 GPUs × 990 TFLOP/s × 0.5 achieved = 3,960 TFLOP/s
     → 3,960e12 / 140e9 = 28,286 prefill tokens/sec per node
     × 8 nodes = 226,000 tok/s
     
  We need 66,720. Prefill capacity is 3.4x demand. ✓
  
  BUT: prefill and decode SHARE the GPUs.
     prefill time fraction: 66,720 / 226,000 = 29.5%
     decode time fraction:  19,460 / 88,888 = 21.9%
     total: 51.4% utilization

Verdict: the fleet is adequate with roughly 2x headroom. Now design it properly.


3. Step 2 — the architecture#

                        ┌──────────────────┐
    clients ──────────► │  API GATEWAY     │  auth, rate limit (tokens +
                        │  (CPU nodes, 6)  │  concurrency), validation,
                        └────────┬─────────┘  tokenization, SSE streaming
                                 │
                        ┌────────▼─────────┐
                        │  ROUTER          │  prefix-aware + least-KV-loaded
                        │  (CPU nodes, 3)  │  session affinity by conv_id
                        └────────┬─────────┘
                                 │
        ┌────────┬────────┬──────┼──────┬────────┬────────┬────────┐
        ▼        ▼        ▼      ▼      ▼        ▼        ▼        ▼
     ┌─────┐  ┌─────┐  ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐
     │node0│  │node1│  │node2│ │node3│ │node4│ │node5│ │node6│ │node7│
     │TP=8 │  │TP=8 │  │TP=8 │ │TP=8 │ │TP=8 │ │TP=8 │ │TP=8 │ │TP=8 │
     └─────┘  └─────┘  └─────┘ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘
      general pool (6 nodes)            long-ctx pool    spare/canary
                                        (1 node)         (1 node)

Key decisions and their justifications:

DECISION                      JUSTIFICATION
TP=8, DP=8 (not TP=16 or 64)  TP=8 fills one NVSwitch domain (Section IX.03).
                              Higher TP crosses nodes → unusable (IX.09).
                              DP for throughput (IX.02).

FP8 quantization              2x throughput and memory; well-validated;
                              Hopper native (VII.05). Without it, weights are
                              141 GB → 17.6 GB/GPU, halving KV capacity and
                              doubling decode time. FP8 is what makes the
                              budget work.

FP8 KV cache                  2x concurrency (VII.12). Not strictly needed
                              here (memory isn't binding) but reduces KV
                              bandwidth, improving ITL at high batch.

Prefix caching ON             Chat with conversation history: turn N re-prefills
                              turns 1..N-1. Expected hit rate 70-85%.
                              → cuts prefill demand from 66,720 to ~15,000 tok/s.
                              THIS IS THE LARGEST SINGLE WIN.

Prefix-aware routing          Without it, prefix cache hit rate is ~1/8.
                              Session affinity by conversation_id (VIII.06).

Chunked prefill ON            p95 input is 6,000 tokens; unchunked, a 6k prefill
                              is ~210 ms of GPU, spiking ITL for everyone (V.03).
                              max_num_batched_tokens = 2048.

Separate long-context pool    A 32k+ request consumes 8x the KV and blocks
                              prefill. Segregating protects the general pool's
                              p99 (VIII.06).

One spare node                N+1 for the 99.9% SLO, and it serves as the
                              canary target for rollouts (XI.07).

CUDA graphs ON                20-30% at these batch sizes (VII.08).

4. Step 3 — the configuration#

Shell
vllm serve meta-llama/Meta-Llama-3-70B-Instruct \
  --quantization fp8 \
  --kv-cache-dtype fp8 \
  --tensor-parallel-size 8 \
  --max-model-len 8192 \
  --max-num-seqs 160 \
  --max-num-batched-tokens 2048 \
  --gpu-memory-utilization 0.90 \
  --enable-prefix-caching \
  --enable-chunked-prefill \
  --scheduling-policy priority \
  --port 8000
JUSTIFICATION FOR EACH TUNED VALUE

max-model-len 8192
  p95 input 6,000 + p95 output 1,200 = 7,200, +14% headroom.
  NOT 128k — that would inflate block tables and reduce usable blocks
  for no benefit (VIII.04). Long-context requests go to the other pool.

max-num-seqs 160
  Memory allows ~806 at 4k avg context. But:
    - we only need 800/6 = 133 per general-pool node
    - 160 gives 20% headroom above expected
    - and keeps ITL comfortable: at 160, step bytes = 8.8 + 12.8 = 21.6 GB
      → 6.4 ms theoretical, ~11.5 ms real. Still ≪ 40 ms SLO.
  Deliberately well below the memory maximum (V.09, preemption avoidance).

max-num-batched-tokens 2048
  From the ITL formula (VIII.04): prefill rate ~28,000 tok/s per node
  → 2048 tokens = 73 ms of prefill work per step... too much.
  Recompute: we want the prefill chunk to add < 15 ms to a step.
  28,000 tok/s → 15 ms = 420 tokens. But that's very small chunks.
  Compromise: 2048 tokens adds ~73 ms to ONE step every few steps,
  but chunked prefill spreads it. Measured: p95 ITL 24 ms. Acceptable.
  → TUNE THIS EMPIRICALLY. Start at 2048, measure p95 ITL, adjust.

gpu-memory-utilization 0.90
  Not 0.95: the p99 input is much larger than p95, and we need
  activation headroom for chunked prefill's largest chunk (XI/X.06).

5. Step 4 — does it meet the SLO?#

TTFT BUDGET (target: p95 < 500 ms for inputs ≤ 4,000 tokens)

  network RTT (regional)                    25 ms
  gateway (auth, rate limit, route)         12 ms
  tokenization (4,000 tokens, Rust)          4 ms
  queue wait (at 51% utilization)          ~90 ms   ← from queueing theory
  prefill:
     with 78% prefix cache hit rate,
     effective new tokens ≈ 880
     880 / 28,000 tok/s                     31 ms
     but chunked, interleaved with decode:  ~55 ms
  first sample                               2 ms
  SSE first write                            6 ms
  ─────────────────────────────────────────────────
  TOTAL                                    194 ms
  
  p95 (with variance)                     ~340 ms   ✓ under 500 ms

ITL (target: p95 < 40 ms)
  decode step at batch 160:                11.5 ms
  + chunked prefill interleaving:          +8 ms average
  + scheduler and sampling overhead:       +3 ms
  ─────────────────────────────────────────────────
  p50                                       ~19 ms
  p95 (steps containing a prefill chunk)    ~28 ms  ✓ under 40 ms

AVAILABILITY (target 99.9%)
  8 nodes, N+1 → survives one node failure with degradation
  → 7 nodes serving 51% × 8/7 = 59% utilization. Still fine.
  Deployment: roll one node at a time; 7 nodes serve during each roll.
  ✓ achievable, given fast failure detection and restart (XI.05)

Verdict: the SLO is met with headroom. Now stress it.


6. Step 5 — what breaks it#

FAILURE MODE                          IMPACT              MITIGATION
Prefix cache hit rate drops to 20%    prefill demand 3.5x  → still fits
  (e.g. a product change adds a          → utilization 78%    (barely).
   unique timestamp to prompts)                              Monitor hit rate;
                                                             alert on drops.

Traffic doubles                       utilization 102%     → degradation ladder;
                                                             need 8 more nodes.
                                                             Order lead time!

Input lengths double (product adds    prefill demand 2x    → utilization 81%.
  retrieval)                                                 Tighter but OK.

A burst of 32k-context requests       long-ctx pool         → route them there;
                                      saturates              general pool
                                                             unaffected. ✓
                                                             (this is why we
                                                              segregated)

One node fails                        7 nodes, 59% util    → fine ✓

Two nodes fail                        6 nodes, 68% util    → fine, tighter

FP8 quality regression discovered     must revert to BF16   → weights 141 GB
                                                             → 17.6 GB/GPU
                                                             → decode 2x slower
                                                             → utilization 95%
                                                             → SLO AT RISK
                                                             ← THE REAL RISK

The last row is the one to plan for. The design depends on FP8. If FP8 proves unacceptable for quality, the budget does not support BF16 at this SLO. Mitigations: validate FP8 thoroughly before committing (Section IV.12), and have INT8 W8A8 as a fallback (similar speedup, different quantization method, so a quality issue with one may not affect the other).

State this dependency explicitly in the design. A design whose critical assumption is unstated is a design that will fail surprisingly.


7. Step 6 — cost#

64 H100s at $2.50/GPU-hour reserved = $160/hour = $1,401,600/year

Tokens per year:
  19,460 output tok/s × 0.51 duty (peak vs average over 24h) × 3.15e7 s
  = 3.13e11 output tokens/year
  
Cost per million output tokens:
  $1,401,600 / 313,000 = $4.48

Hmm — that's high. Why?
  Because we're sized for PEAK and the average is much lower.
  Average utilization across 24 hours: ~30% (peak is 51%, trough is ~10%)
  
  Improvements:
    - run batch/offline workloads in the trough → +$0 marginal, better
      amortization
    - reduce the fleet and accept degradation at peak
    - use spot instances for 2 of the 8 nodes

This is the honest answer: the fleet is sized for peak, and peak is a fraction of the day. Filling the trough with batch work is the single largest cost lever remaining.


8. The complete answer, summarized#

ARCHITECTURE
  8 nodes × 8 H100, TP=8 within nodes, DP=8 across
  6 general pool, 1 long-context pool, 1 spare/canary
  CPU gateway tier (6 nodes) + router tier (3 nodes)

CONFIGURATION
  FP8 weights and KV cache, prefix caching, chunked prefill,
  CUDA graphs, max_model_len 8192, max_num_seqs 160,
  max_num_batched_tokens 2048 (tuned empirically)

ROUTING
  prefix-aware with session affinity by conversation_id
  long-context requests (> 16k) segregated to their own pool
  least-KV-loaded within the affinity group

RESULT
  p95 TTFT ~340 ms (target 500) ✓
  p95 ITL   ~28 ms (target 40)  ✓
  utilization 51% at peak → headroom for 1.9x growth
  availability 99.9% with N+1 ✓
  cost ~$4.48/M output tokens (improvable by filling the trough)

CRITICAL DEPENDENCIES (state these!)
  1. FP8 quality must be acceptable — validate before committing
  2. Prefix cache hit rate ≥ 60% — monitor and alert
  3. Prompt structure must remain cache-friendly — a product change
     that adds per-request content to the prompt prefix would break it
  4. Long-context traffic must stay < 5% of requests

WHAT I'D MEASURE FIRST IN PRODUCTION
  prefix cache hit rate, average running batch, p95 ITL,
  queue wait p95, cost per million tokens

9. Variations to practice#

Work each of these; they’re the follow-up questions.

A. "Budget is 32 GPUs, not 64."
   → utilization 102%. Options: accept degradation, INT4 weights
     (decode-heavy so it helps), reduce max context, tier the SLO,
     or route overflow to a smaller model. Show the arithmetic for each.

B. "The SLO is p95 TTFT < 200 ms."
   → the network and gateway alone are 37 ms; prefill for 4k tokens
     is unavoidable. Need: higher prefix hit rate, smaller model for
     the first token, or accept that only cached-prefix requests
     meet it. Consider a speculative/draft model for TTFT.

C. "Add 100 fine-tuned variants."
   → multi-LoRA (VIII.09). One base model, adapters swapped per request.
     5-15% overhead, ~4 GB for 100 adapters. Adapter-aware routing.

D. "Users are global."
   → multi-region (XI.06). Latency is the driver. Size each region for
     its traffic + degraded failover. Sessions stay regional.

E. "Peak is 10x average, spiky."
   → autoscaling can't help (cold start). Need headroom for the spike,
     or a degradation ladder, or overflow to an external provider.

F. "What if we used a 405B model?"
   → FP8: 405 GB. TP=8 gives 50.6 GB/GPU — fits, but only 25 GB for KV.
     Decode is 5.7x slower per token. At the same fleet: utilization
     would far exceed 100%. Would need ~4x the GPUs or a smaller model.

G. "Can we use A100s instead?"
   → 2.04 TB/s vs 3.35: decode 1.64x slower. No FP8 (INT8 instead).
     Need ~1.7x the nodes. But A100s may be cheaper per hour —
     compute $/M tokens for both.

10. How to answer this in an interview#

1. CLARIFY THE WORKLOAD FIRST.
   "10,000 concurrent users" is ambiguous. Ask about the duty cycle,
   the length distributions, and what "concurrent" means.
   ← this alone distinguishes strong candidates

2. DO THE ARITHMETIC BEFORE THE ARCHITECTURE.
   Is it even possible? Which resource binds?

3. STATE ASSUMPTIONS EXPLICITLY.
   Duty cycle, prefix hit rate, achieved efficiency fraction.

4. JUSTIFY EVERY CHOICE FROM A CONSTRAINT.
   "TP=8 because that's the NVSwitch domain and TP=16 crosses nodes."
   Not "TP=8 because that's typical."

5. IDENTIFY THE CRITICAL DEPENDENCY.
   Here: FP8. Say what breaks if it fails.

6. SAY WHAT YOU'D MEASURE.
   A design without a validation plan is a guess.

7. BE HONEST ABOUT WHAT YOU DON'T KNOW.
   "I'd measure the prefix hit rate before committing to this;
    if it's below 50% the numbers change materially."

11. Hands-on exercise#

A. Do it yourself. Before re-reading the solution, work the whole problem from the brief. Compare your answer. Where did you differ, and was your reasoning sound?

B. Work the variations. Do all seven variations in section 9 with full arithmetic.

C. Build the calculator. Extend your Section V.15 and IX.09 calculators into a single tool that takes the brief and produces the design. Validate it against this solution.

D. Stress it. For the design in section 8, compute the utilization under each of the failure modes in section 6. Which is the binding one?

E. Present it. Write the design as a one-page document with the architecture diagram, configuration, SLO analysis, dependencies, and validation plan. This is the actual deliverable in a real job.


12. Interview questions#

  1. Design a serving architecture for a 70B model, 10,000 concurrent users, 64 H100s.
  2. What’s the first question you’d ask about “10,000 concurrent users”?
  3. Which resource binds in this design — memory, compute, or bandwidth?
  4. Why TP=8 and DP=8 rather than TP=64?
  5. What’s the critical dependency in your design, and what happens if it fails?
  6. Your cost is $4.48/M tokens. Why so high, and how would you reduce it?
  7. The budget is halved. What do you change?

13. Further reading#

↑↓ navigate↵ openesc close