★ The counterintuitive file. Read it before your next capacity request.
1. The claim#
Adding GPUs to a model instance can reduce throughput, waste money, and increase tail latency. Not always — but often enough that “we’ll add more GPUs” should never be a reflex.
2. The six ways it goes wrong#
Way 1 — TP past the efficiency knee#
Llama-3-70B, measured:
TP=2: 1,230 tok/s (615 per GPU)
TP=4: 2,290 tok/s (573 per GPU) 93% efficiency
TP=8: 3,760 tok/s (470 per GPU) 82%
TP=16: 5,160 tok/s (323 per GPU) 56%
Going from TP=8 to TP=16:
throughput: +37%
GPUs: +100%
cost per token: +46% WORSEYou bought 8 more GPUs and made every token more expensive. If you needed throughput, DP=2 of the TP=8 configuration would have given +100% throughput at the same efficiency.
Way 2 — TP over a slow interconnect#
4 GPUs, PCIe Gen4 only (no NVLink):
1 GPU alone (if the model fits): 470 tok/s
TP=4 over PCIe: 760 tok/s
4 GPUs producing 1.6x the throughput of one.
Efficiency: 40%.
4 independent replicas (DP=4): 1,880 tok/s. Efficiency: 100%.TP over PCIe wasted 60% of the hardware. And this configuration is common — many cloud instance types have PCIe-only multi-GPU.
Way 3 — Pipeline bubbles at low concurrency#
PP=4, 8 concurrent sequences (2 microbatches of 4):
bubble fraction = (4-1)/(2+4-1) = 60%
4 GPUs delivering 40% of their potential.
1 GPU (if it fit) would deliver 100% of one GPU.
→ PP=4 gives 1.6 GPUs' worth of work from 4 GPUs.PP requires high concurrency to be worth anything. At low traffic it’s actively harmful.
Way 4 — MoE load imbalance#
EP=8, batch 32, skewed routing:
hottest expert gets 2.2x its fair share
→ all GPUs wait at the All-to-All barrier
→ effective utilization 46%
At batch 512, the routing distribution smooths:
hottest expert gets 1.15x
→ effective utilization 87%More GPUs with EP at low batch amplifies the imbalance, because each GPU has fewer experts and less averaging.
Way 5 — Splitting the batch#
This one is subtle and important.
Scenario: 100 requests/sec arriving. Two options:
A) 1 replica of TP=8:
all 100 req/s go to one scheduler
average running batch: 64
step bytes: 17.6 GB (weights) + KV for 64
B) 8 replicas of TP=1 (if the model fits):
12.5 req/s per replica
average running batch: 8
step bytes: 141 GB (weights!) + KV for 8
Option B reads 141 GB per step per replica to serve 8 sequences.
Option A reads 17.6 GB per GPU per step to serve 64.
Per-token weight-read cost:
A: 17.6 GB / 64 tokens = 275 MB per token per GPU
B: 141 GB / 8 tokens = 17.6 GB per token per replica = 17.6 GB per GPU
→ B reads 64x more per token!Splitting traffic across too many small replicas destroys batching efficiency. This is the argument for consolidation, and it’s why multi-tenant platforms that pool traffic beat per-team deployments (Section XII).
The general principle: batching efficiency requires traffic concentration. Every replica you add divides the traffic and lowers the average batch size.
Way 6 — Increased failure surface#
1 GPU: MTBF = M
TP=8 group: any of 8 GPUs failing kills the group → MTBF = M/8
2-node TP=8×PP=2: 16 GPUs + network → MTBF ≈ M/20
Plus: longer cold start (more to load), harder deployments (atomic groups),
more complex failure modes.Availability decreases with parallelism degree. At some point you’re spending more on redundancy than you gained in performance.
3. The diagnostic questions#
Before adding GPUs, answer these:
1. WHAT AM I TRYING TO IMPROVE?
Throughput → add REPLICAS (DP)
Latency → add TP, but check the efficiency curve
Capacity → add TP or PP, or quantize
2. WHAT IS MY CURRENT EFFICIENCY?
throughput_per_GPU at the current config vs at TP=1 (or the smallest
config that fits). If it's already below 70%, adding more will be worse.
3. WHAT IS MY AVERAGE RUNNING BATCH SIZE?
If it's much less than what memory allows, you are TRAFFIC-limited,
not capacity-limited. More GPUs will make it worse, not better.
← THIS IS THE MOST COMMONLY MISSED CHECK
4. WHAT IS MY INTERCONNECT?
No NVLink → don't add TP.
5. HAVE I QUANTIZED?
FP8 is a free 2x. Do it before buying hardware.
6. HAVE I FIXED SCHEDULING?
If continuous batching, prefix caching, and chunked prefill aren't
enabled, you have a 3-10x available for free.Question 3 catches the most common error. A system running at batch 6 when memory allows 60 does not need more GPUs; it needs more traffic per replica, which means fewer replicas with better routing.
4. The consolidation argument#
Because way 5 is counterintuitive, here it is as a positive recommendation:
GIVEN: 5 teams each running their own replica of the same model,
each getting 20 req/s.
CURRENT: 5 replicas, average batch 12 each.
Per-replica: 141 GB read per step for 12 tokens.
CONSOLIDATED: 2 replicas (with headroom), average batch 30 each.
Per-replica: 141 GB read per step for 30 tokens.
Throughput per GPU: 2.5x better.
GPUs used: 5 → 2 (plus a shared platform layer).Consolidating traffic onto fewer, better-utilized replicas is one of the largest available wins in a multi-team organization, and it’s an organizational problem more than a technical one. It’s the core value proposition of an internal inference platform (Section XII).
5. Worked example — the wrong decision and the right one#
SITUATION
Llama-3-70B on 8×H100 (TP=8). Traffic has doubled.
p95 TTFT has gone from 600 ms to 3,200 ms.
Team proposes: "add 8 more GPUs, go to TP=16."
DIAGNOSIS
1. Where is the time? → queue_wait is 2,700 ms of the 3,200 ms.
→ this is a CAPACITY problem, not a latency problem.
2. Average running batch: 58 of a possible 175.
→ not memory-limited. The scheduler is fine.
3. Efficiency at TP=8: 82%. At TP=16 it would be 56%.
WHY TP=16 IS WRONG
It would improve ITL (which is fine at 22 ms) and give only +37%
throughput for +100% GPUs. Cost per token gets 46% worse.
And it doesn't address the queue.
THE RIGHT ANSWER
Add a second TP=8 replica (DP=2 × TP=8).
→ +100% throughput for +100% GPUs (100% marginal efficiency)
→ queue_wait drops to near zero
→ ITL unchanged at 22 ms
→ cost per token unchanged
Even better, check first:
- Is prefix caching on? (chat workload → potentially 40% saving)
- Is FP8 viable? (2x, might mean no new GPUs at all)
- Is chunked prefill on? (long prompts blocking → TTFT spikes)In this example the right answer might have been “enable two config flags” rather than “buy eight H100s.” That’s a $200k/year difference from a ten-minute diagnosis.
6. The efficiency table to keep#
Strategy Marginal efficiency per added GPU
DP (add a replica) ~100% (if routing is good and traffic supports it)
TP 1→2 ~98%
TP 2→4 ~89%
TP 4→8 ~73%
TP 8→16 ~51%
TP 16→32 ~29%
PP (high concur.) ~85%
PP (low concur.) ~40%
EP (high batch) ~85%
EP (low batch) ~45%
TP over PCIe ~40%Anything below ~70% marginal efficiency should require justification.
7. Production implications#
- Make “average running batch size” a primary dashboard metric. It tells you immediately whether you’re traffic-limited or capacity-limited.
- Track throughput per GPU as a KPI, not just total throughput. It’s the number that degrades when you over-parallelize.
- Require an efficiency calculation in capacity requests. “We need 8 more GPUs” should come with “at 82% marginal efficiency, because…”
- Consolidate before scaling. Multiple under-utilized replicas of the same model are pure waste.
- Exhaust the free wins first: scheduling config, prefix caching, quantization. They’re typically 2-10x and cost engineering time, not hardware.
- Prefer DP for throughput. Always.
8. Common mistakes#
Increasing TP to fix a throughput problem. Use DP.
Increasing TP to fix a queueing problem. Add replicas.
Adding GPUs when the average batch is already low. Makes it worse.
Not measuring marginal efficiency.
TP over PCIe.
PP at low concurrency.
Running many small under-utilized replicas instead of consolidating.
Buying hardware before enabling FP8 and prefix caching.
9. Hands-on exercise#
A. Measure your efficiency curve. For a model you can run, measure throughput per GPU at TP=1, 2, 4, 8 (as far as hardware allows). Plot marginal efficiency. Where’s your knee?
B. Demonstrate way 5. With a fixed arrival rate, compare 1 replica handling all traffic versus 4 replicas each handling a quarter. Measure per-GPU throughput for both. Confirm the batching penalty.
C. The diagnostic. For a system you have access to, answer all six diagnostic questions from section 3. What’s the correct next action?
D. TP over PCIe. If you have PCIe-only GPUs (or can simulate with
NCCL_P2P_DISABLE=1), measure TP=4 efficiency and compare to DP=4.
E. The consolidation calculation. For a hypothetical organization with 5 teams each running a replica at 15 req/s, compute the GPU saving from consolidating to 2 shared replicas. Include the headroom you’d need.
10. Interview questions#
- Give three ways that adding GPUs can make an inference system worse.
- Why does TP efficiency degrade with degree? What’s the functional form?
- A system has p95 TTFT of 3 seconds and average running batch of 6 out of a possible 60. What do you do?
- Why does splitting traffic across more replicas hurt per-GPU throughput?
- What’s the marginal efficiency of TP=8→16, and why is DP better?
- What is the consolidation argument for a shared inference platform?
- What would you require in a capacity request before approving it?
11. Further reading#
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022)
- [FUNDAMENTAL] Amdahl’s law and Gustafson’s law — the general form of this argument
- [ESTABLISHED] Yu et al., “Orca” — batching efficiency and why concentration matters
- Next: Section X — Memory & Performance Engineering