Section V.12 covered the algorithm and the theory. This file covers deployment: which variant, how to tune it, how it interacts with everything else, and how to tell whether it’s helping.
1. The deployment decision#
START: is your workload decode-bound at LOW batch?
├─ NO (prefill-heavy, or batch > 64)
│ → speculative decoding will help little or hurt. Skip it.
│
└─ YES
├─ Does output often COPY from input? (summarization, editing, RAG, code
│ completion with context, structured output)
│ → N-GRAM / PROMPT LOOKUP. Free, no extra model, try it first.
│
├─ Is there a small model in the same family with the same tokenizer?
│ → DRAFT MODEL. 1.5-2.5x. Costs memory for the draft.
│
├─ Can you train auxiliary heads?
│ → EAGLE or MEDUSA. Highest acceptance, best speedups, most work.
│
└─ Does the model support multi-token prediction natively?
→ use it. (DeepSeek-V3-style MTP.)Try n-gram first. It costs nothing, requires no extra model, and for the right workload gives 2-3x. Many teams skip straight to draft models and never evaluate the free option.
2. Choosing and sizing the draft model#
Requirements:
1. SAME TOKENIZER as the target. Non-negotiable.
2. Same family / distilled from the target → higher acceptance.
3. Cost ratio c = draft_time/target_time should be < 0.1, ideally < 0.05.
Good pairings:
Llama-3-70B ← Llama-3-8B c ≈ 0.11, α ≈ 0.75 speedup ~1.9x
Llama-3-70B ← Llama-3.2-1B c ≈ 0.015, α ≈ 0.65 speedup ~2.1x
Llama-3-8B ← Llama-3.2-1B c ≈ 0.12, α ≈ 0.70 speedup ~1.6x
Qwen2.5-72B ← Qwen2.5-0.5B c ≈ 0.01, α ≈ 0.60 speedup ~1.9x
Bad pairings:
70B ← a 13B draft c too high (0.19); the draft eats the gain
Llama ← Mistral draft different tokenizers; won't work at all
Any ← an untuned random small model α too lowNote the pattern: a smaller, less accurate draft often beats a larger, more accurate one,
because c matters as much as α. Recompute the formula from Section V.12 for your candidates
before choosing.
3. Tuning k (speculation length)#
Optimal k ≈ 1/(1-α), but measure.
Fixed k:
Simple. Choose from measured α.
α=0.6 → k=2-3; α=0.75 → k=4; α=0.85 → k=6-7
Adaptive k:
Track a running acceptance rate PER REQUEST and adjust.
Code and structured output draft well (α up to 0.9) → use large k.
Creative prose drafts poorly (α ~0.5) → use small k or disable.
Simple policy:
if last 3 iterations accepted >= k-1: k = min(k+1, k_max)
if last iteration accepted <= 1: k = max(k-1, 1)Adaptive k typically buys another 10-20% over a well-chosen fixed k, and much more if your traffic is heterogeneous.
4. The batch-size interaction (the critical caveat)#
Verification cost = k+1 tokens through the target model, per sequence.
At batch B, the target's forward pass processes B×(k+1) tokens instead of B.
B=1, k=4: 5 tokens. Still memory-bound. Verification is nearly free. ✓
B=32, k=4: 160 tokens. Approaching the compute-bound regime. Costs something.
B=128, k=4: 640 tokens. Firmly compute-bound. Verification COSTS REAL TIME. ✗Measured speedup vs batch size (70B, 1B draft, α=0.7, k=4):
batch 1: 2.1x
batch 4: 1.9x
batch 16: 1.6x
batch 64: 1.2x
batch 128: 0.95x ← now a LOSS
batch 256: 0.85xProduction systems must gate speculation on batch size. A fixed configuration that helps at 3 a.m. hurts at peak. Some engines do this automatically; verify yours.
Policy: enable speculation when running_batch < threshold
threshold ≈ ridge_point / (k+1) as a starting point, then measure5. Interaction with everything else#
FEATURE INTERACTION
Continuous batching ⚠ variable accepted tokens per sequence per step
complicates the scheduler's bookkeeping
CUDA graphs ⚠ variable token counts → either capture per-count
or run verification outside the graph
Paged KV cache ⚠ must handle speculative tokens: allocate blocks for
k+1 tokens, then FREE the rejected ones
Prefix caching ✓ compatible
Quantization ✓ compatible; the draft can be quantized too
Chunked prefill ✓ orthogonal
Tensor parallelism ⚠ the draft usually runs on fewer GPUs; adds complexity
Structured decoding ⚠ the grammar constrains the draft too, or you waste
drafts on invalid tokensThe KV cache interaction is the one people underestimate. Speculative tokens write K and V into the cache. When tokens are rejected, those entries must be discarded — which means the block manager needs a rollback operation. Getting this wrong causes subtle corruption.
6. Measuring whether it’s working#
Instrument these, per request and in aggregate:
acceptance_rate = accepted_tokens / drafted_tokens
tokens_per_iteration = accepted + 1 (bonus)
draft_time_fraction = draft_time / total_time
effective_speedup = baseline_ITL / current_ITL
And the decisive one:
tokens_per_second WITH speculation vs WITHOUT, at the same batch sizeIf you cannot A/B it, you don’t know whether it’s helping. Build the ability to toggle speculation and compare on live traffic.
Warning signs:
acceptance_rate < 0.4 → draft model is poorly matched; reconsider
draft_time_fraction > 0.25 → draft is too expensive; use a smaller one
speedup < 1.1 at batch 1 → something is wrong; check the implementation
speedup < 1.0 at your p50 batch → disable it7. Performance summary#
Method Best case Typical Extra memory Engineering
n-gram lookup 3.0x 1.4-2.0x none trivial
Draft model 2.5x 1.6-2.1x draft size moderate
Medusa 2.8x 1.8-2.2x heads (~5%) high (training)
EAGLE-2 4.0x 2.2-3.0x head (~2%) high (training)
MTP (native) 2.5x 1.8-2.2x built in none (if the model has it)Best case = ideal workload, batch 1. Typical = realistic mixed traffic at moderate batch.
The gap between “best case” and “typical” is why published speedups disappoint in production. Papers benchmark at batch 1 on favorable content.
8. Production implications#
- Start with n-gram. Free, and for RAG/summarization/editing workloads it’s often the best option outright.
- Gate on batch size. Non-negotiable for a system that sees varying load.
- Budget the draft model’s memory as KV cache you gave up. A 1B draft is 2 GB — roughly 4 concurrent sequences at 4k context on an 8B target.
- Monitor acceptance rate by workload type. It varies enormously and tells you where speculation is earning its complexity.
- Verify distribution equivalence in testing: with temperature 0, speculative and non-speculative decoding must produce identical output. If they don’t, your rejection sampling is wrong and you have a silent quality bug.
- Consider it a later optimization. It’s rank 10 on the list in file 01 for a reason — significant engineering for a benefit that batching partly erases.
9. Common mistakes#
Not gating on batch size. A regression at peak load.
Choosing a draft that’s too large. c matters as much as α.
Mismatched tokenizers. Doesn’t work; sometimes fails silently with garbage.
Skipping the rejection-sampling correction. You’re now serving the draft model.
Not measuring acceptance rate. Flying blind.
Forgetting the KV rollback for rejected tokens. Subtle corruption.
Deploying it before fixing batching and quantization. Wrong order.
Believing paper speedups. Benchmark your own workload at your own batch sizes.
10. Hands-on exercise#
A. N-gram first. Implement prompt-lookup speculation: search the prompt for the current suffix, propose the continuation. Measure the speedup on (i) summarization, (ii) code completion with context, (iii) open-ended creative writing. Explain the differences.
B. Draft model selection. For a target model you can run, evaluate 2-3 candidate drafts.
Measure α and c for each. Predict the speedup with the formula, then measure it. Which
candidate wins, and does the formula predict correctly?
C. Batch gating. Measure speedup at batch 1, 4, 16, 64, 128. Find the batch size at which speculation stops helping on your hardware. Implement the gate.
D. Adaptive k. Implement per-request adaptive k. Compare to the best fixed k on a mixed workload (code + prose + summarization). How much does adaptivity buy?
E. Verify equivalence. With temperature 0, confirm that speculative and non-speculative decoding produce byte-identical output over 100 prompts. If not, find the bug.
F. KV rollback. Instrument your block manager to verify that rejected speculative tokens' KV entries are properly discarded. Construct a test that would catch a rollback bug.
11. Interview questions#
- What’s the first speculative decoding variant you’d try, and why?
- Why does a smaller, less accurate draft sometimes beat a larger, more accurate one?
- Why must speculation be gated on batch size? At what batch does it stop helping?
- What does speculative decoding do to the KV cache, and what must the block manager support?
- How would you tune k?
- How do you verify that a speculative decoding implementation is correct?
- Why do published speedups exceed what you see in production?
12. Further reading#
- [ESTABLISHED] Leviathan et al. (2022); Chen et al. (2023)
- [EMERGING] Cai et al., “Medusa” (2024); Li et al., “EAGLE-2” (2024)
- [REFERENCE] vLLM speculative decoding documentation and its batch-size heuristics
- [ESTABLISHED] DeepSeek-V3 technical report — native multi-token prediction
- Next: 14 — Pruning, distillation, architecture