PidokuInfra

Chunked Prefill

Expert Advanced 1h Difficulty 3/5 Topic 05 of 12

Prerequisites V.03, V.09, VIII.04


1. Problem → Why → Optimization#

PROBLEM   A long prefill occupies the GPU for hundreds of milliseconds,
          during which every decoding sequence stalls. p99 ITL explodes.

WHY       Prefill and decode compete for the same GPU. A 32k-token prefill
          is ~500 ms of compute; a decode step is ~10 ms. Running the
          prefill as one unit means a 50-step stall for everyone.

OPTIMIZE  Split the prefill into chunks and process one chunk per
          iteration, alongside the decode work.

TRADE-OFFS
  ✓ p99 ITL improves dramatically (5-20x in bad cases)
  ✓ throughput improves slightly (better GPU utilization —
    prefill is compute-bound, decode is memory-bound; mixing them
    uses both resources)
  ✗ TTFT for the long prompt is slightly worse (it's spread out)
  ✗ requires varlen kernels and a scheduler that handles mixed batches

WHEN TO USE   Any workload with both long prompts and an ITL SLO.
              → almost always. This is on by default in modern engines.
WHEN NOT TO   Pure batch workloads with no latency SLO, where you want
              maximum prefill throughput.

2. The picture#

WITHOUT CHUNKED PREFILL

  step:  1    2    3    4              5    6    7
  GPU:  [d]  [d]  [d]  [████ P ████]  [d]  [d]  [d]
                        └─ 500 ms ─┘
  decode sequences see: 10, 10, 10, 510, 10, 10 ms   ← p99 = 510 ms

WITH CHUNKED PREFILL (chunk = 2048 tokens, 16 chunks)

  step:  1     2       3       4       ...    18     19
  GPU:  [d]  [d+p1]  [d+p2]  [d+p3]   ...  [d+p16]  [d]
              └ 40 ms each ┘
  decode sequences see: 10, 40, 40, 40, ... 40, 10 ms   ← p99 = 40 ms
  
  the long prompt's TTFT: 16 × 40 = 640 ms instead of 510 ms
  → slightly worse for THAT request, dramatically better for everyone else

The trade is explicit and favorable: one request’s TTFT degrades 25% so that dozens of requests’ ITL improves 12x.


3. The mechanism#

The scheduler builds a MIXED batch each iteration:

  tokens = [prefill chunk for req A: 2048 tokens]
         + [decode token for req B: 1]
         + [decode token for req C: 1]
         + ...
         + [decode token for req Z: 1]

  cu_seqlens = [0, 2048, 2049, 2050, ..., 2048+N]

  → a single varlen forward pass (Section IV.11)
  → the attention kernel handles the ragged shape
  → req A's chunk attends to its own previous chunks (already in KV)
    plus itself, causally

The scheduler tracks per-request progress:

Go
type Request struct {
	PromptTokenIDs    []int
	NumComputedTokens int   // how much of the prompt is prefilled
	Stage             Stage // Prefill | Decode
}

type Scheduled struct {
	Req    *Request
	Tokens int
}

func (e *Engine) schedule(tokenBudget int) (scheduled []Scheduled) {
	// 1. decode requests first — they're latency-critical, 1 token each
	for _, r := range e.runningDecode {
		if tokenBudget < 1 {
			break
		}
		scheduled = append(scheduled, Scheduled{r, 1})
		tokenBudget--
	}

	// 2. fill the remaining budget with prefill chunks
	for _, r := range e.runningPrefill {
		chunk := min(len(r.PromptTokenIDs)-r.NumComputedTokens, tokenBudget)
		if chunk <= 0 {
			break
		}
		scheduled = append(scheduled, Scheduled{r, chunk})
		r.NumComputedTokens += chunk
		tokenBudget -= chunk
		if r.NumComputedTokens == len(r.PromptTokenIDs) {
			r.Stage = Decode // prefill complete; first token next step
		}
	}
	return scheduled
}

Decode is scheduled first because decode requests are latency-critical and cheap (1 token each); prefill fills whatever budget remains.


4. The tuning parameter#

max_num_batched_tokens is the total token budget per iteration.

SMALL (512-1024)
  ✓ each step is fast → excellent ITL
  ✗ prefill is slow (many tiny chunks, each with fixed overhead)
  ✗ prefill chunks may be too small to be compute-efficient
    (a 512-token chunk doesn't fill the GPU well — Section V.03's
     crossover is ~300 tokens)

LARGE (8192-32768)
  ✓ prefill throughput is high
  ✗ a step containing a large chunk is slow → ITL spikes
  ✗ approaches the unchunked behavior

TYPICAL: 2048-8192
HOW TO CHOOSE

  1. measure prefill throughput: tok/s for your model
  2. ITL budget for the prefill contribution:
       target_ITL - decode_step_time
  3. max_num_batched_tokens ≈ prefill_tok_per_sec × budget_seconds

  Example: prefill 28,000 tok/s, ITL SLO 50 ms, decode step 12 ms
    budget = (50 - 12) ms = 38 ms
    tokens = 28,000 × 0.038 = 1,064
    → but that's smaller than the compute-efficiency threshold
    → compromise at 2048, measure the actual p95 ITL, adjust

Measure, don’t just compute. The formula gives a starting point; the interaction with scheduling, CUDA graph batch sizes, and your actual traffic mix determines the right value.


5. The throughput bonus#

Chunked prefill often improves throughput, which is counterintuitive:

PREFILL alone:  compute-bound (tensor cores busy, memory idle)
DECODE alone:   memory-bound (memory busy, tensor cores idle)
MIXED:          both resources used

  → a step containing prefill chunks AND decode tokens uses the
    GPU more completely than either alone

MEASURED: 5-15% throughput improvement from mixing, on top of the
          latency improvement.

This is a genuine free lunch and one of the reasons chunked prefill became the default: it improves both latency and throughput.


6. Interactions#

CUDA GRAPHS (Section VII.08)
  ⚠ mixed batches have variable shapes → harder to capture
  → engines typically graph the decode-only steps and run mixed
    steps eagerly, or capture per (decode_count, chunk_size) pair
  → check whether your engine loses graph capture when chunked
    prefill is enabled

PREFIX CACHING (Section V.11)
  ✓ compatible; cached blocks reduce the tokens to prefill,
    which reduces the number of chunks

PAGED ATTENTION
  ✓ required; the chunk attends to previously-computed blocks

PIPELINE PARALLELISM
  ⚠ chunks flow through the pipeline; more complex bookkeeping

SPECULATIVE DECODING
  ⚠ speculative tokens are like a mini-prefill; both compete for
    the token budget

The CUDA graph interaction is the one to verify. If enabling chunked prefill silently disables graph capture, you’ve traded 25% throughput for latency — probably still worth it, but you should know.


7. Measured effect#

Llama-3-70B, 8×H100, workload: 90% short (500 token) prompts,
10% long (24k token) prompts, 40 concurrent decodes

  CONFIGURATION          p50 ITL   p95 ITL   p99 ITL   TTFT p95   Throughput
  chunked off             18 ms    142 ms    680 ms    480 ms      3,900 tok/s
  chunked, 8192           21 ms     58 ms    124 ms    520 ms      4,180
  chunked, 2048           24 ms     34 ms     51 ms    610 ms      4,240
  chunked, 512            31 ms     38 ms     47 ms    890 ms      3,850
                                                                    ↑ prefill
                                                                    too slow

The 2048 row is the sweet spot here: p99 ITL improved 13x, TTFT degraded 27%, throughput improved 9%.

Note the 512 row: chunks that are too small hurt everything, because prefill becomes inefficient (below the compute-bound threshold) and TTFT suffers without further ITL benefit.


8. Production implications#

  • Enable it. Default in recent vLLM and TensorRT-LLM. If you’re on an older version or a custom stack, this is a high-value addition.
  • Tune max_num_batched_tokens against your ITL SLO. Start at 2048, measure p95 ITL.
  • Don’t go below ~1024. Chunks that small are compute-inefficient.
  • Verify CUDA graph capture still works after enabling it.
  • Expect a small TTFT regression for long prompts. It’s the trade.
  • It composes with long-context pool segregation (Section VIII.06) — do both.
  • Re-tune when your traffic mix changes. The right budget depends on the ratio of long to short prompts.

9. Common mistakes#

Leaving it off with long prompts and an ITL SLO. The single most common cause of ITL spikes.

max_num_batched_tokens too large. Approaches unchunked behavior.

Too small. Prefill becomes inefficient; TTFT suffers with no ITL benefit.

Not verifying the CUDA graph interaction.

Scheduling prefill before decode. Decode is latency-critical and cheap; it should go first.

Not re-tuning when the traffic mix changes.


10. Hands-on exercise#

A. Demonstrate the problem. With chunked prefill off, run 20 streaming requests and inject a 24k-token prompt. Plot the streaming requests’ ITL over time. Measure the spike.

B. Fix it and measure. Enable chunked prefill and repeat. Quantify the p99 ITL improvement and the TTFT cost.

C. Sweep the budget. Reproduce the table in section 7 for your model: sweep max_num_batched_tokens over {512, 1024, 2048, 4096, 8192, 16384} and measure p50/p95/p99 ITL, p95 TTFT, and throughput. Where’s your sweet spot?

D. Verify the throughput bonus. Measure throughput with chunked prefill on and off, at a traffic mix where both prefill and decode are present. Do you see the 5-15% improvement?

E. CUDA graph interaction. Check your engine’s logs and profile: is graph capture still active with chunked prefill enabled? What’s the throughput difference?

F. Implement it. Add chunked prefill to your continuous-batching engine from Project 08. Verify correctness (outputs match unchunked) and measure the latency improvement.


11. Interview questions#

  1. What problem does chunked prefill solve? Quantify it.
  2. What does it cost, and for whom?
  3. How would you choose max_num_batched_tokens?
  4. Why can chunks that are too small hurt?
  5. Why does chunked prefill sometimes improve throughput?
  6. Why schedule decode before prefill in the token budget?
  7. How does chunked prefill interact with CUDA graphs?

12. Further reading#

  • [ESTABLISHED] Agrawal et al., “Sarathi-Serve: Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve” (OSDI 2024)
  • [ESTABLISHED] Agrawal et al., “SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills” (2023)
  • [REFERENCE] vLLM chunked prefill documentation and implementation
  • Next: 06 — Disaggregated prefill/decode

↑↓ navigate↵ openesc close