1. What is it?#
Batching means processing several requests in one pass through the model, so the weights are read from memory once and used many times.
NO BATCHING BATCHING
read all weights → serve req 1 read all weights → serve reqs 1..32
read all weights → serve req 2
read all weights → serve req 3
...
read all weights → serve req 32
32 weight reads 1 weight readThat’s the whole idea. The weight read is the expensive part; batching amortizes it.
2. Why does it exist?#
Because at batch size 1, a modern GPU is almost entirely idle.
Concrete: generating one token with a 7B FP16 model requires
- 14 GFLOPs of arithmetic (2 × 7B), and
- 14 GB of memory reads (every weight, once).
An H100 delivers roughly 990 TFLOP/s (FP16, with sparsity off: ~990) and 3,350 GB/s of HBM bandwidth. Time for each part:
compute: 14e9 FLOP / 990e12 FLOP/s = 0.0141 ms
memory: 14e9 B / 3.35e12 B/s = 4.18 msThe memory read takes 296x longer than the arithmetic. The GPU spends 99.7% of the step waiting for weights to arrive. Its arithmetic units — the thing you paid for — are doing nothing.
Now batch 32 sequences. Memory reads: still 14 GB (same weights!). Compute: 32 × 14 GFLOPs = 448 GFLOPs → 0.45 ms. Still memory bound, but now you produced 32 tokens in ~4.2 ms instead of 1 token. 32x the throughput for ~0% extra memory cost.
That is why every production LLM system batches, and why batching is not an optimization but the architecture.
3. Simple analogy#
A bus versus a taxi.
A taxi (batch 1) picks you up immediately and takes you directly — best latency, terrible efficiency, one engine burning fuel for one person.
A bus (batch 32) makes you wait at the stop and takes a longer route — worse latency per passenger, dramatically better fuel per passenger.
The subtlety that makes LLM serving interesting: the bus’s fuel consumption barely changes with the number of passengers. Driving an empty bus costs almost as much as a full one. So running empty buses is pure waste, and the operator’s whole job is filling seats without making anyone wait too long.
Continuous batching (Section V.09) is the upgrade: a bus that lets passengers board and alight while moving, at every stop, instead of waiting for everyone to complete the full route.
4. Tiny example#
One matrix multiply, x · W, with W of 4096×4096 in FP16 on an A100-class GPU. The whole
behaviour falls out of two hardware numbers — how fast the GPU moves bytes and how fast it does
arithmetic — so a few lines of Go can predict it:
// batching.go — predicts the batching curve for one matmul from two hardware numbers.
package main
import "fmt"
func main() {
const (
d = 4096.0
bytesPer = 2.0 // FP16
bandwidth = 950e9 // bytes/s the GPU can actually stream (A100-class)
peak = 58e12 // FLOP/s the GPU can actually sustain on this op
)
for _, B := range []float64{1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024} {
flops := 2 * B * d * d // work grows with B
bytes := d*d*bytesPer + 2*B*d*bytesPer // W is read once, whatever B is
t := max(bytes/bandwidth, flops/peak) // the slower resource sets the time
fmt.Printf("B=%5.0f %7.1f us %6.2f TFLOP/s %6.1f GB/s intensity=%7.1f\n",
B, t*1e6, flops/t/1e12, bytes/t/1e9, flops/bytes)
}
}Output (a real A100 measures within about 10% of this, with a rounder knee):
B= 1 35.3 us 0.95 TFLOP/s 950.0 GB/s intensity= 1.0
B= 2 35.4 us 1.90 TFLOP/s 950.0 GB/s intensity= 2.0
B= 4 35.4 us 3.79 TFLOP/s 950.0 GB/s intensity= 4.0
B= 8 35.5 us 7.57 TFLOP/s 950.0 GB/s intensity= 8.0
B= 16 35.6 us 15.08 TFLOP/s 950.0 GB/s intensity= 15.9
B= 32 35.9 us 29.93 TFLOP/s 950.0 GB/s intensity= 31.5
B= 64 37.0 us 58.00 TFLOP/s 934.6 GB/s intensity= 62.1
B= 128 74.1 us 58.00 TFLOP/s 481.4 GB/s intensity= 120.5
B= 256 148.1 us 58.00 TFLOP/s 254.9 GB/s intensity= 227.6
B= 512 296.2 us 58.00 TFLOP/s 141.6 GB/s intensity= 409.6
B= 1024 592.4 us 58.00 TFLOP/s 85.0 GB/s intensity= 682.7Read the shape of this table carefully. It contains the entire economics of LLM serving.
- B = 1 to 16: time is flat. You are 16x-ing your work for free. The GPU was waiting on memory the whole time; you filled the wait with useful arithmetic.
- Around B = 64: the knee. Bandwidth utilization starts falling as compute rises. This is the transition from memory-bound to compute-bound.
- B ≥ 128: time is linear in B. Now you are compute-bound; every extra sequence costs real time. TFLOP/s has plateaued at the hardware’s practical ceiling.
The knee is the most important number in your system. Below it, batching is free throughput.
Above it, batching trades latency for throughput. Find it on your hardware for your model and
write it in numbers.md.
5. Technical explanation#
The arithmetic#
For a linear layer (B, d) @ (d, d):
FLOPs = 2 · B · d²
Bytes = 2·d² (weights, FP16)
+ 2·B·d (input)
+ 2·B·d (output)
Arithmetic intensity = 2·B·d² / (2d² + 4Bd)
≈ B when B << dArithmetic intensity ≈ batch size. That is the cleanest statement of why batching works, and you should be able to derive it in an interview.
A GPU’s ridge point — the intensity at which it stops being memory-bound — is:
ridge = peak_FLOP/s ÷ peak_bytes/s
A100: 312e12 / 2.0e12 = 156 FLOP/byte
H100: 990e12 / 3.35e12 = 296 FLOP/byte
H200: 990e12 / 4.8e12 = 206 FLOP/byteSo on an H100, you need a batch size of roughly 296 before dense FP16 decode becomes compute-bound. In practice you hit KV-cache memory limits, latency SLOs, and attention costs long before that. Most production LLM decode is memory-bound, always. Internalize this.
The three regimes#
throughput
^
| ___________________ compute-bound plateau
| /
| /
| / ← the knee (intensity ≈ ridge point)
| /
| /
| /
| / memory-bound: linear gain, ~free
| /
|/
+--------------------------------------------> batch size| Regime | Symptom | What to do |
|---|---|---|
| Memory-bound (small B) | throughput ∝ B, latency flat | increase B — it is free |
| Knee | latency starts rising | this is usually your operating point |
| Compute-bound (large B) | throughput flat, latency ∝ B | stop; add GPUs instead |
Why prefill and decode batch differently#
Prefill of a 2,000-token prompt already has 2,000 “rows” in its matmuls — it is already compute-bound with a batch of one request. Decode has one row per sequence — it needs 100+ sequences to reach the same regime.
Prefill: effective matmul rows = Σ prompt_lengths → large, compute-bound
Decode: effective matmul rows = number of sequences → small, memory-boundThis asymmetry is the root of chunked prefill, disaggregated serving, and most scheduler complexity. Keep it in mind through Section V.
Static batching and why it isn’t enough#
The naive scheme:
1. Collect requests until batch is full OR timeout elapses
2. Run the whole batch to completion
3. Return all results
4. Go to 1Problems for LLMs:
Batch of 4, generating different lengths:
seq A: ████████████████████████ 400 tokens
seq B: ████ 30 tokens [then idle for 370 steps]
seq C: ██████ 60 tokens [then idle for 340 steps]
seq D: ███ 20 tokens [then idle for 380 steps]
└──────────── all slots held until A finishes ────────────┘Sequences B, C, D finished early but their batch slots and KV cache are held hostage until A finishes. Measured waste in real workloads: 60-80% of slots idle. Also, new requests cannot join mid-flight — they wait for the entire batch.
That is what continuous batching fixes (Section V.09) by making scheduling decisions at every decode step rather than every request.
6. Under the hood#
At the kernel level, batching changes GEMV into GEMM:
B=1: (1, 4096) @ (4096, 4096) → GEMV. Each weight element used ONCE.
B=32: (32,4096) @ (4096, 4096) → GEMM. Each weight element used 32 times
after being loaded into shared memory.The GPU loads a tile of W into shared memory (fast SRAM) once, then reuses it across all 32
rows of the input. That reuse is the physical mechanism of the amortization. Tensor cores can
also finally be used efficiently: they operate on matrix tiles (e.g. 16×16), and a batch of 1
wastes 15 of the 16 rows.
At the framework level, batching requires padding or packing, because sequences have different lengths:
Padded: Packed (varlen):
[t t t t P P P P] [t t t t | t t | t t t t t t]
[t t P P P P P P] cu_seqlens = [0, 4, 6, 12]
[t t t t t t P P]
← wasted compute → ← no waste; needs varlen kernelsModern engines use packed/varlen layouts with FlashAttention’s cu_seqlens interface. Padding
can waste 30-50% of prefill compute on skewed length distributions.
7. Performance implications#
Throughput rises roughly linearly with batch size until the knee, then plateaus.
Latency per request is roughly flat until the knee, then rises linearly.
Cost per token falls as 1/B until the knee, then flattens. This means the economic
incentive to batch is enormous and saturating: going from B=1 to B=32 might cut cost 30x; going
from 32 to 64 might cut it 20%.
Memory is the constraint that stops you: each sequence in the batch needs its own KV cache.
max_batch = (GPU_memory − weights − overhead) / (KV bytes per sequence)For 70B FP16 on 8×H100 (640 GB): weights 140 GB, overhead ~20 GB, leaves 480 GB. KV for 70B (GQA, 8 KV heads, 80 layers, head_dim 128, FP16) is 320 KB/token. At 4k context that’s 1.28 GB per sequence → ~375 concurrent sequences. At 32k context → ~47. Context length is a concurrency multiplier in the denominator, which is why long-context serving is so expensive.
8. Production implications#
- Batch size is not a static config value. It is an outcome of a scheduling policy under a
memory constraint. Systems that expose
max_batch_sizeas a tuning knob are asking you to guess; systems with continuous batching decide it per step. - The batch you get depends on traffic. At 3 a.m. with 2 requests/sec you get batch 2 and terrible economics. This is why multi-tenant platforms and model consolidation matter (Section XII): pooling traffic raises everyone’s batch size.
- Latency SLOs cap batch size. If p95 ITL must be < 50 ms and each extra 32 sequences adds 10 ms, you have a hard ceiling regardless of memory.
- Offline/batch workloads should use enormous batches. Nightly embedding or summarization jobs have no latency SLO — run them at the compute-bound plateau and get 10x better cost.
9. Common mistakes#
Setting batch size to 1 “for latency.” Below the knee, larger batches cost you almost nothing in latency and multiply throughput. Measure your knee before deciding.
Setting batch size to the maximum that fits in memory. Above the knee you pay latency for diminishing throughput, and you have zero headroom for a long-context request — which will OOM you at 2 a.m.
Using static batching for autoregressive models. Wastes 60-80% of capacity. Use an engine with continuous batching.
Padding everything to max length. On a distribution with mean 300 and max 8,000, padding to 8,000 wastes 96% of your prefill compute. Use packed sequences.
Assuming the knee is the same for prefill and decode. It is not, by orders of magnitude. Schedule them differently.
Forgetting that batching changes numerics. Different batch sizes select different GEMM kernels with different reduction orders, so outputs can differ in the last bits — and after sampling, that can change the generated text. If you need bit-exact reproducibility, you have a hard problem (Section X.07).
10. Hands-on exercise#
A. Find your knee. Run the benchmark from section 4 on your hardware. Record the batch size at which time-per-call starts rising superlinearly. Compute your GPU’s ridge point from its spec sheet and compare. Do they agree? Explain any discrepancy.
B. Model the tradeoff. Using your measured numbers, build a small spreadsheet/script: for each batch size, compute per-request latency, system throughput, and cost per million tokens (assume $2/GPU-hour). Find the batch size that minimizes cost subject to per-request latency < 2x the batch-1 latency.
C. Simulate static-batching waste. Sample 1,000 output lengths from a lognormal distribution (µ=4.5, σ=1.0 — roughly realistic). Simulate static batching with batch size 32 and compute the fraction of slot-steps that were idle. Then simulate a scheduler that refills a slot as soon as it frees. Compare total time and utilization.
D. Padding waste. With the same length distribution, compute the fraction of compute wasted by padding to the batch max vs. packing.
11. Interview questions#
- Why does batching improve LLM throughput so dramatically? Give the mechanism in terms of memory traffic.
- Derive arithmetic intensity for a batched linear layer and explain why it ≈ B.
- What is the ridge point of an H100 in FP16, and what does it imply for decode batch size?
- Why do prefill and decode reach the compute-bound regime at such different batch sizes?
- What limits batch size in practice — and rank the limits by how often they bind.
- Explain the waste in static batching for autoregressive generation, quantitatively.
- Your throughput stopped improving past batch 64 but latency keeps rising. What is happening and what do you do?
12. Further reading#
- [ESTABLISHED] Yu et al., “Orca” (OSDI 2022) — the paper that formalized the waste in static batching
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022) §3
- [REFERENCE] NVIDIA “Matrix Multiplication Background User’s Guide” — tile quantization effects
- Next: 07 — Compute vs memory