Section V.12 covered the algorithm; VII.13 covered deployment. This file surveys the variants and gives an honest assessment of each.
1. The design space#
Every variant answers three questions differently:
1. WHERE DO THE DRAFT TOKENS COME FROM?
a separate model / extra heads / the model's own features /
the prompt itself / a datastore
2. WHAT SHAPE ARE THE PROPOSALS?
a chain (k tokens) / a tree (many branches)
3. WHAT DOES IT COST?
extra memory / extra training / extra compute per stepVARIANT DRAFT SOURCE SHAPE TRAINING? STATUS
Draft model a smaller LLM chain no ESTABLISHED
N-gram / lookup the prompt itself chain no ESTABLISHED
Medusa extra heads on target tree yes EMERGING
EAGLE / EAGLE-2 feature-level head tree yes EMERGING
Lookahead decoding Jacobi iteration tree no EMERGING
Self-speculative a subset of layers chain no/light RESEARCH
MTP (native) built into the model chain n/a ESTABLISHED*
REST a retrieval datastore tree no RESEARCH
* where the model provides it (e.g. DeepSeek-V3)2. N-gram / prompt lookup — try this first#
IDEA: search the prompt (and generated text so far) for the current
suffix; propose what followed it last time.
context: "...the quick brown fox jumps over the lazy dog. The quick brown"
suffix: "The quick brown"
found earlier: "the quick brown fox"
→ propose "fox jumps over"
COST: zero. No model, no memory, no training. A string search.
ACCEPTANCE RATE
summarization (output quotes input): 0.5-0.9
code completion with context: 0.4-0.8
editing / rewriting: 0.6-0.9
RAG (answer quotes retrieved docs): 0.3-0.6
open-ended creative writing: 0.05-0.15
general chat: 0.1-0.3
SPEEDUP: 1.5-3x on the favorable workloads, ~1.0x otherwise (and it
costs nothing when it doesn't help)This should be the first thing you try. It’s a config flag in most engines
(--speculative-model [ngram] or similar), costs nothing, and for the right workload it’s the
best option available.
// ngramPropose finds the last earlier occurrence of the current n-gram and proposes what followed.
func ngramPropose(context []int, n, k int) []int {
if len(context) < n {
return nil
}
suffix := context[len(context)-n:]
for i := len(context) - n - 1; i >= 0; i-- {
if slices.Equal(context[i:i+n], suffix) {
return context[i+n : min(i+n+k, len(context))]
}
}
return nil
}3. Draft model — the classic#
Covered in Sections V.12 and VII.13. Summary:
✓ works for any workload (not content-dependent like n-gram)
✓ well-understood, widely supported
✗ needs a matched small model (same tokenizer, ideally distilled)
✗ costs memory for the draft
✗ the draft's forward pass is on the critical path
SELECTION: minimize c (draft cost ratio) and maximize α (acceptance).
Section VII.13's table. Typically a 0.5-1B draft for a 70B target.4. Medusa — extra heads#
IDEA: add K extra prediction heads to the target model. Head i predicts
token t+i+1 directly from the same hidden state.
hidden state h_t
├─ lm_head → token t+1 (the normal head)
├─ medusa_1 → token t+2
├─ medusa_2 → token t+3
└─ medusa_3 → token t+4
Each head proposes its top-k candidates → a TREE of possibilities.
Verify the whole tree in one forward pass with a custom attention mask.✓ no separate model → no extra weight loading, minimal extra memory
✓ the heads are cheap (one linear layer each)
✓ tree verification gives higher effective acceptance than a chain
✗ requires TRAINING the heads (on the target model's outputs)
✗ acceptance degrades with head index (head 3 is much worse than head 1)
✗ the tree attention mask is a custom kernel
✗ head quality is task-dependent
REPORTED: 2.0-2.8x at batch 1The training requirement is the practical barrier. You need the target model’s outputs on representative data, and you must retrain when the target changes.
5. EAGLE / EAGLE-2 — feature-level drafting#
IDEA: instead of predicting TOKENS, predict the target model's
HIDDEN STATE (feature) for the next position, then decode
a token from it.
Insight: the token sequence is high-entropy and hard to predict,
but the FEATURE sequence is smoother and more predictable.
EAGLE's draft head takes (previous feature, previous token embedding)
and predicts the next feature. Small — one transformer layer.
EAGLE-2 adds: a dynamic draft tree whose shape adapts to the draft
model's confidence at each position.✓ highest reported acceptance rates (0.8-0.9)
✓ small head (~2% of target parameters)
✓ EAGLE-2's dynamic tree adapts to easy vs hard positions
✗ requires training the head
✗ needs access to the target's hidden states (an intrusive integration)
✗ tree attention kernels
✗ retraining needed when the target model changes
REPORTED: 2.5-4.0x at batch 1 — the best of the variantsEAGLE-2 currently has the strongest reported results. The integration cost is real: you need the target model’s internals, not just its outputs.
6. Lookahead decoding#
IDEA: use Jacobi iteration to generate multiple tokens in parallel,
maintaining an n-gram pool of previously-seen continuations.
No draft model, no training.
Each step: run the model on several positions simultaneously with
guessed values, refine them, and collect the n-grams that emerge
into a pool for future use.
✓ NO draft model, NO training
✓ works with any model
✗ requires more compute per step (multiple positions)
✗ acceptance builds up over the generation (cold start)
✗ complex to implement correctly
REPORTED: 1.5-2.3xAttractive because it needs no training, which removes the main barrier of Medusa and EAGLE. Less mature in production engines.
7. Multi-token prediction (native)#
Some models are TRAINED to predict multiple future tokens.
DeepSeek-V3 includes MTP modules for exactly this.
✓ no bolt-on machinery; the capability is in the model
✓ the model's own training aligned the prediction
✓ can be used for speculation at inference, or discarded
✗ only available if the model has it
STATUS: [ESTABLISHED] where the model provides it.This is probably where the field goes. Bolting speculation onto a model trained without it is inherently a workaround; training for it is cleaner. Expect more models to include MTP.
8. Tree verification#
Several variants use trees rather than chains. Worth understanding.
CHAIN: propose t+1, t+2, t+3, t+4. Verify. Accept a prefix.
→ if t+2 is wrong, t+3 and t+4 are wasted regardless of correctness
TREE: propose a branching set of candidates
t+1
┌──────┴──────┐
"the" "a"
┌──┴──┐ ┌──┴──┐
"cat" "dog" "cat" "bird"
→ verify ALL paths in one forward pass
→ accept the longest matching path
→ higher effective acceptance for the same number of target passes
THE MECHANISM: a custom attention mask where each node attends only to
its ancestors, not its siblings.
positions: [prompt...][the][a][cat][dog][cat][bird]
mask: "cat"(child of "the") attends to prompt + "the", NOT to "a"The tree costs more verification compute (more tokens through the target) but that compute was free at low batch (Section V.12). At high batch it isn’t — which is why tree methods degrade faster with batch size.
9. Honest comparison#
VARIANT batch 1 batch 32 Training? Memory Complexity Status
n-gram 1.5-3.0x 1.1-1.5x no none trivial ESTABLISHED
draft model 2.0-2.5x 1.4-1.8x no* 2-4 GB moderate ESTABLISHED
Medusa 2.0-2.8x 1.3-1.7x YES ~5% high EMERGING
EAGLE-2 2.5-4.0x 1.5-2.0x YES ~2% high EMERGING
lookahead 1.5-2.3x 1.1-1.4x no none high EMERGING
MTP (native) 1.8-2.5x 1.3-1.8x n/a built-in low ESTABLISHED*
* the draft model may itself need distillation for good acceptanceEvery column degrades with batch size (Section V.12’s batch interaction). The batch-32 column is closer to what production sees.
10. Choosing#
DECISION
□ Is your workload decode-bound at LOW batch?
NO → speculation won't help. Stop.
□ Does your output often copy from your input?
YES → N-GRAM. Free, and often the best option. Try it first.
□ Does your model have native MTP?
YES → use it.
□ Is there a small matched model available?
YES → DRAFT MODEL. The reliable general-purpose choice.
□ Can you train auxiliary heads, and is the extra 1.5x worth
weeks of work plus retraining on every model update?
YES → EAGLE-2.
NO → stop at the draft model.Most teams should stop at n-gram or a draft model. The trained variants give another 1.3-1.6x for substantially more engineering and an ongoing retraining obligation.
11. Production implications#
- Try n-gram first. Free.
- Gate on batch size (Section VII.13). Non-negotiable.
- Measure acceptance rate per workload type, continuously.
- Verify distribution equivalence at temperature 0 (Section V.12).
- Budget the draft model’s memory as KV cache you gave up.
- Trained variants create a retraining obligation. Every target model update requires retraining the heads. Factor that into the decision.
- Tree verification degrades faster with batch than chains.
- Watch for native MTP in new models — it removes the bolt-on complexity.
12. Common mistakes#
Skipping n-gram. It’s free and often best.
Not gating on batch size. A regression at peak.
Adopting a trained variant without accounting for retraining.
Believing paper speedups. They’re batch-1, on favorable content.
Not measuring acceptance rate in production.
Skipping the rejection-sampling correction. Silent quality change.
Forgetting the KV rollback for rejected tokens (Section VII.13).
13. Hands-on exercise#
A. Implement n-gram. Write the proposer from section 2. Measure acceptance and speedup on four workload types: summarization, code, RAG, creative writing.
B. Compare variants. For a model you can run, benchmark n-gram and a draft model at batch 1, 8, 32, 128. Plot speedup vs batch for each. Where do they cross 1.0?
C. Tree vs chain. Implement chain verification and a simple 2-branch tree. Measure effective acceptance (tokens accepted per target pass) for each.
D. Acceptance by content. Instrument acceptance rate by request type in a mixed workload. Which content types draft well? Does it match section 2’s table?
E. Adaptive k. Implement per-request adaptive speculation length based on recent acceptance. Compare to the best fixed k on a heterogeneous workload.
F. Verify equivalence. At temperature 0, confirm speculative and non-speculative decoding produce byte-identical output over 200 prompts. If not, find the bug in your rejection sampling.
14. Interview questions#
- Survey the speculative decoding variants. Which would you try first and why?
- Why does n-gram speculation work, and for which workloads?
- What is EAGLE’s insight about features versus tokens?
- What is tree verification and what does it cost?
- Why do all variants degrade with batch size?
- What ongoing obligation does a trained variant create?
- How would you verify a speculative decoding implementation is correct?
15. Further reading#
- [ESTABLISHED] Leviathan et al. (2022); Chen et al. (2023)
- [EMERGING] Cai et al., “Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads” (2024)
- [EMERGING] Li et al., “EAGLE” (2024) and “EAGLE-2” (2024)
- [EMERGING] Fu et al., “Lookahead Decoding” (2024)
- [ESTABLISHED] DeepSeek-V3 technical report — native multi-token prediction
- Next: 10 — Low-bit inference: FP4 and beyond