1. Problem → Why → Optimization#
PROBLEM A long prefill occupies the GPU for hundreds of milliseconds,
during which every decoding sequence stalls. p99 ITL explodes.
WHY Prefill and decode compete for the same GPU. A 32k-token prefill
is ~500 ms of compute; a decode step is ~10 ms. Running the
prefill as one unit means a 50-step stall for everyone.
OPTIMIZE Split the prefill into chunks and process one chunk per
iteration, alongside the decode work.
TRADE-OFFS
✓ p99 ITL improves dramatically (5-20x in bad cases)
✓ throughput improves slightly (better GPU utilization —
prefill is compute-bound, decode is memory-bound; mixing them
uses both resources)
✗ TTFT for the long prompt is slightly worse (it's spread out)
✗ requires varlen kernels and a scheduler that handles mixed batches
WHEN TO USE Any workload with both long prompts and an ITL SLO.
→ almost always. This is on by default in modern engines.
WHEN NOT TO Pure batch workloads with no latency SLO, where you want
maximum prefill throughput.2. The picture#
WITHOUT CHUNKED PREFILL
step: 1 2 3 4 5 6 7
GPU: [d] [d] [d] [████ P ████] [d] [d] [d]
└─ 500 ms ─┘
decode sequences see: 10, 10, 10, 510, 10, 10 ms ← p99 = 510 ms
WITH CHUNKED PREFILL (chunk = 2048 tokens, 16 chunks)
step: 1 2 3 4 ... 18 19
GPU: [d] [d+p1] [d+p2] [d+p3] ... [d+p16] [d]
└ 40 ms each ┘
decode sequences see: 10, 40, 40, 40, ... 40, 10 ms ← p99 = 40 ms
the long prompt's TTFT: 16 × 40 = 640 ms instead of 510 ms
→ slightly worse for THAT request, dramatically better for everyone elseThe trade is explicit and favorable: one request’s TTFT degrades 25% so that dozens of requests’ ITL improves 12x.
3. The mechanism#
The scheduler builds a MIXED batch each iteration:
tokens = [prefill chunk for req A: 2048 tokens]
+ [decode token for req B: 1]
+ [decode token for req C: 1]
+ ...
+ [decode token for req Z: 1]
cu_seqlens = [0, 2048, 2049, 2050, ..., 2048+N]
→ a single varlen forward pass (Section IV.11)
→ the attention kernel handles the ragged shape
→ req A's chunk attends to its own previous chunks (already in KV)
plus itself, causallyThe scheduler tracks per-request progress:
type Request struct {
PromptTokenIDs []int
NumComputedTokens int // how much of the prompt is prefilled
Stage Stage // Prefill | Decode
}
type Scheduled struct {
Req *Request
Tokens int
}
func (e *Engine) schedule(tokenBudget int) (scheduled []Scheduled) {
// 1. decode requests first — they're latency-critical, 1 token each
for _, r := range e.runningDecode {
if tokenBudget < 1 {
break
}
scheduled = append(scheduled, Scheduled{r, 1})
tokenBudget--
}
// 2. fill the remaining budget with prefill chunks
for _, r := range e.runningPrefill {
chunk := min(len(r.PromptTokenIDs)-r.NumComputedTokens, tokenBudget)
if chunk <= 0 {
break
}
scheduled = append(scheduled, Scheduled{r, chunk})
r.NumComputedTokens += chunk
tokenBudget -= chunk
if r.NumComputedTokens == len(r.PromptTokenIDs) {
r.Stage = Decode // prefill complete; first token next step
}
}
return scheduled
}Decode is scheduled first because decode requests are latency-critical and cheap (1 token each); prefill fills whatever budget remains.
4. The tuning parameter#
max_num_batched_tokens is the total token budget per iteration.
SMALL (512-1024)
✓ each step is fast → excellent ITL
✗ prefill is slow (many tiny chunks, each with fixed overhead)
✗ prefill chunks may be too small to be compute-efficient
(a 512-token chunk doesn't fill the GPU well — Section V.03's
crossover is ~300 tokens)
LARGE (8192-32768)
✓ prefill throughput is high
✗ a step containing a large chunk is slow → ITL spikes
✗ approaches the unchunked behavior
TYPICAL: 2048-8192HOW TO CHOOSE
1. measure prefill throughput: tok/s for your model
2. ITL budget for the prefill contribution:
target_ITL - decode_step_time
3. max_num_batched_tokens ≈ prefill_tok_per_sec × budget_seconds
Example: prefill 28,000 tok/s, ITL SLO 50 ms, decode step 12 ms
budget = (50 - 12) ms = 38 ms
tokens = 28,000 × 0.038 = 1,064
→ but that's smaller than the compute-efficiency threshold
→ compromise at 2048, measure the actual p95 ITL, adjustMeasure, don’t just compute. The formula gives a starting point; the interaction with scheduling, CUDA graph batch sizes, and your actual traffic mix determines the right value.
5. The throughput bonus#
Chunked prefill often improves throughput, which is counterintuitive:
PREFILL alone: compute-bound (tensor cores busy, memory idle)
DECODE alone: memory-bound (memory busy, tensor cores idle)
MIXED: both resources used
→ a step containing prefill chunks AND decode tokens uses the
GPU more completely than either alone
MEASURED: 5-15% throughput improvement from mixing, on top of the
latency improvement.This is a genuine free lunch and one of the reasons chunked prefill became the default: it improves both latency and throughput.
6. Interactions#
CUDA GRAPHS (Section VII.08)
⚠ mixed batches have variable shapes → harder to capture
→ engines typically graph the decode-only steps and run mixed
steps eagerly, or capture per (decode_count, chunk_size) pair
→ check whether your engine loses graph capture when chunked
prefill is enabled
PREFIX CACHING (Section V.11)
✓ compatible; cached blocks reduce the tokens to prefill,
which reduces the number of chunks
PAGED ATTENTION
✓ required; the chunk attends to previously-computed blocks
PIPELINE PARALLELISM
⚠ chunks flow through the pipeline; more complex bookkeeping
SPECULATIVE DECODING
⚠ speculative tokens are like a mini-prefill; both compete for
the token budgetThe CUDA graph interaction is the one to verify. If enabling chunked prefill silently disables graph capture, you’ve traded 25% throughput for latency — probably still worth it, but you should know.
7. Measured effect#
Llama-3-70B, 8×H100, workload: 90% short (500 token) prompts,
10% long (24k token) prompts, 40 concurrent decodes
CONFIGURATION p50 ITL p95 ITL p99 ITL TTFT p95 Throughput
chunked off 18 ms 142 ms 680 ms 480 ms 3,900 tok/s
chunked, 8192 21 ms 58 ms 124 ms 520 ms 4,180
chunked, 2048 24 ms 34 ms 51 ms 610 ms 4,240
chunked, 512 31 ms 38 ms 47 ms 890 ms 3,850
↑ prefill
too slowThe 2048 row is the sweet spot here: p99 ITL improved 13x, TTFT degraded 27%, throughput improved 9%.
Note the 512 row: chunks that are too small hurt everything, because prefill becomes inefficient (below the compute-bound threshold) and TTFT suffers without further ITL benefit.
8. Production implications#
- Enable it. Default in recent vLLM and TensorRT-LLM. If you’re on an older version or a custom stack, this is a high-value addition.
- Tune
max_num_batched_tokensagainst your ITL SLO. Start at 2048, measure p95 ITL. - Don’t go below ~1024. Chunks that small are compute-inefficient.
- Verify CUDA graph capture still works after enabling it.
- Expect a small TTFT regression for long prompts. It’s the trade.
- It composes with long-context pool segregation (Section VIII.06) — do both.
- Re-tune when your traffic mix changes. The right budget depends on the ratio of long to short prompts.
9. Common mistakes#
Leaving it off with long prompts and an ITL SLO. The single most common cause of ITL spikes.
max_num_batched_tokens too large. Approaches unchunked behavior.
Too small. Prefill becomes inefficient; TTFT suffers with no ITL benefit.
Not verifying the CUDA graph interaction.
Scheduling prefill before decode. Decode is latency-critical and cheap; it should go first.
Not re-tuning when the traffic mix changes.
10. Hands-on exercise#
A. Demonstrate the problem. With chunked prefill off, run 20 streaming requests and inject a 24k-token prompt. Plot the streaming requests’ ITL over time. Measure the spike.
B. Fix it and measure. Enable chunked prefill and repeat. Quantify the p99 ITL improvement and the TTFT cost.
C. Sweep the budget. Reproduce the table in section 7 for your model: sweep
max_num_batched_tokens over {512, 1024, 2048, 4096, 8192, 16384} and measure p50/p95/p99 ITL,
p95 TTFT, and throughput. Where’s your sweet spot?
D. Verify the throughput bonus. Measure throughput with chunked prefill on and off, at a traffic mix where both prefill and decode are present. Do you see the 5-15% improvement?
E. CUDA graph interaction. Check your engine’s logs and profile: is graph capture still active with chunked prefill enabled? What’s the throughput difference?
F. Implement it. Add chunked prefill to your continuous-batching engine from Project 08. Verify correctness (outputs match unchunked) and measure the latency improvement.
11. Interview questions#
- What problem does chunked prefill solve? Quantify it.
- What does it cost, and for whom?
- How would you choose
max_num_batched_tokens? - Why can chunks that are too small hurt?
- Why does chunked prefill sometimes improve throughput?
- Why schedule decode before prefill in the token budget?
- How does chunked prefill interact with CUDA graphs?
12. Further reading#
- [ESTABLISHED] Agrawal et al., “Sarathi-Serve: Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve” (OSDI 2024)
- [ESTABLISHED] Agrawal et al., “SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills” (2023)
- [REFERENCE] vLLM chunked prefill documentation and implementation
- Next: 06 — Disaggregated prefill/decode