PidokuInfra

Speculative Decoding Variants

Expert Advanced 1h 15m Difficulty 4/5 Topic 09 of 12

Prerequisites V.12, VII.13

Section V.12 covered the algorithm; VII.13 covered deployment. This file surveys the variants and gives an honest assessment of each.


1. The design space#

Every variant answers three questions differently:

1. WHERE DO THE DRAFT TOKENS COME FROM?
     a separate model / extra heads / the model's own features /
     the prompt itself / a datastore

2. WHAT SHAPE ARE THE PROPOSALS?
     a chain (k tokens) / a tree (many branches)

3. WHAT DOES IT COST?
     extra memory / extra training / extra compute per step
VARIANT              DRAFT SOURCE          SHAPE   TRAINING?  STATUS
Draft model          a smaller LLM         chain   no         ESTABLISHED
N-gram / lookup      the prompt itself     chain   no         ESTABLISHED
Medusa               extra heads on target tree    yes        EMERGING
EAGLE / EAGLE-2      feature-level head    tree    yes        EMERGING
Lookahead decoding   Jacobi iteration      tree    no         EMERGING
Self-speculative     a subset of layers    chain   no/light   RESEARCH
MTP (native)         built into the model  chain   n/a        ESTABLISHED*
REST                 a retrieval datastore tree    no         RESEARCH

* where the model provides it (e.g. DeepSeek-V3)

2. N-gram / prompt lookup — try this first#

IDEA: search the prompt (and generated text so far) for the current
      suffix; propose what followed it last time.

  context: "...the quick brown fox jumps over the lazy dog. The quick brown"
  suffix:  "The quick brown"
  found earlier: "the quick brown fox"
  → propose "fox jumps over"

COST: zero. No model, no memory, no training. A string search.

ACCEPTANCE RATE
  summarization (output quotes input):     0.5-0.9
  code completion with context:            0.4-0.8
  editing / rewriting:                     0.6-0.9
  RAG (answer quotes retrieved docs):      0.3-0.6
  open-ended creative writing:             0.05-0.15
  general chat:                            0.1-0.3

SPEEDUP: 1.5-3x on the favorable workloads, ~1.0x otherwise (and it
         costs nothing when it doesn't help)

This should be the first thing you try. It’s a config flag in most engines (--speculative-model [ngram] or similar), costs nothing, and for the right workload it’s the best option available.

Go
// ngramPropose finds the last earlier occurrence of the current n-gram and proposes what followed.
func ngramPropose(context []int, n, k int) []int {
	if len(context) < n {
		return nil
	}
	suffix := context[len(context)-n:]
	for i := len(context) - n - 1; i >= 0; i-- {
		if slices.Equal(context[i:i+n], suffix) {
			return context[i+n : min(i+n+k, len(context))]
		}
	}
	return nil
}

3. Draft model — the classic#

Covered in Sections V.12 and VII.13. Summary:

  ✓ works for any workload (not content-dependent like n-gram)
  ✓ well-understood, widely supported
  ✗ needs a matched small model (same tokenizer, ideally distilled)
  ✗ costs memory for the draft
  ✗ the draft's forward pass is on the critical path

SELECTION: minimize c (draft cost ratio) and maximize α (acceptance).
  Section VII.13's table. Typically a 0.5-1B draft for a 70B target.

4. Medusa — extra heads#

IDEA: add K extra prediction heads to the target model. Head i predicts
      token t+i+1 directly from the same hidden state.

      hidden state h_t
         ├─ lm_head    → token t+1  (the normal head)
         ├─ medusa_1   → token t+2
         ├─ medusa_2   → token t+3
         └─ medusa_3   → token t+4

  Each head proposes its top-k candidates → a TREE of possibilities.
  Verify the whole tree in one forward pass with a custom attention mask.
✓ no separate model → no extra weight loading, minimal extra memory
✓ the heads are cheap (one linear layer each)
✓ tree verification gives higher effective acceptance than a chain

✗ requires TRAINING the heads (on the target model's outputs)
✗ acceptance degrades with head index (head 3 is much worse than head 1)
✗ the tree attention mask is a custom kernel
✗ head quality is task-dependent

REPORTED: 2.0-2.8x at batch 1

The training requirement is the practical barrier. You need the target model’s outputs on representative data, and you must retrain when the target changes.


5. EAGLE / EAGLE-2 — feature-level drafting#

IDEA: instead of predicting TOKENS, predict the target model's
      HIDDEN STATE (feature) for the next position, then decode
      a token from it.

  Insight: the token sequence is high-entropy and hard to predict,
  but the FEATURE sequence is smoother and more predictable.

  EAGLE's draft head takes (previous feature, previous token embedding)
  and predicts the next feature. Small — one transformer layer.

EAGLE-2 adds: a dynamic draft tree whose shape adapts to the draft
              model's confidence at each position.
✓ highest reported acceptance rates (0.8-0.9)
✓ small head (~2% of target parameters)
✓ EAGLE-2's dynamic tree adapts to easy vs hard positions

✗ requires training the head
✗ needs access to the target's hidden states (an intrusive integration)
✗ tree attention kernels
✗ retraining needed when the target model changes

REPORTED: 2.5-4.0x at batch 1 — the best of the variants

EAGLE-2 currently has the strongest reported results. The integration cost is real: you need the target model’s internals, not just its outputs.


6. Lookahead decoding#

IDEA: use Jacobi iteration to generate multiple tokens in parallel,
      maintaining an n-gram pool of previously-seen continuations.
      No draft model, no training.

  Each step: run the model on several positions simultaneously with
  guessed values, refine them, and collect the n-grams that emerge
  into a pool for future use.

✓ NO draft model, NO training
✓ works with any model
✗ requires more compute per step (multiple positions)
✗ acceptance builds up over the generation (cold start)
✗ complex to implement correctly

REPORTED: 1.5-2.3x

Attractive because it needs no training, which removes the main barrier of Medusa and EAGLE. Less mature in production engines.


7. Multi-token prediction (native)#

Some models are TRAINED to predict multiple future tokens.
DeepSeek-V3 includes MTP modules for exactly this.

✓ no bolt-on machinery; the capability is in the model
✓ the model's own training aligned the prediction
✓ can be used for speculation at inference, or discarded
✗ only available if the model has it

STATUS: [ESTABLISHED] where the model provides it.

This is probably where the field goes. Bolting speculation onto a model trained without it is inherently a workaround; training for it is cleaner. Expect more models to include MTP.


8. Tree verification#

Several variants use trees rather than chains. Worth understanding.

CHAIN: propose t+1, t+2, t+3, t+4. Verify. Accept a prefix.
  → if t+2 is wrong, t+3 and t+4 are wasted regardless of correctness

TREE: propose a branching set of candidates
                    t+1
              ┌──────┴──────┐
            "the"        "a"
           ┌──┴──┐      ┌──┴──┐
        "cat" "dog"  "cat" "bird"
  → verify ALL paths in one forward pass
  → accept the longest matching path
  → higher effective acceptance for the same number of target passes

THE MECHANISM: a custom attention mask where each node attends only to
               its ancestors, not its siblings.

  positions:  [prompt...][the][a][cat][dog][cat][bird]
  mask: "cat"(child of "the") attends to prompt + "the", NOT to "a"

The tree costs more verification compute (more tokens through the target) but that compute was free at low batch (Section V.12). At high batch it isn’t — which is why tree methods degrade faster with batch size.


9. Honest comparison#

VARIANT          batch 1   batch 32   Training?  Memory   Complexity  Status
n-gram           1.5-3.0x  1.1-1.5x   no         none     trivial     ESTABLISHED
draft model      2.0-2.5x  1.4-1.8x   no*        2-4 GB   moderate    ESTABLISHED
Medusa           2.0-2.8x  1.3-1.7x   YES        ~5%      high        EMERGING
EAGLE-2          2.5-4.0x  1.5-2.0x   YES        ~2%      high        EMERGING
lookahead        1.5-2.3x  1.1-1.4x   no         none     high        EMERGING
MTP (native)     1.8-2.5x  1.3-1.8x   n/a        built-in low         ESTABLISHED*

* the draft model may itself need distillation for good acceptance

Every column degrades with batch size (Section V.12’s batch interaction). The batch-32 column is closer to what production sees.


10. Choosing#

DECISION

□ Is your workload decode-bound at LOW batch?
    NO → speculation won't help. Stop.

□ Does your output often copy from your input?
    YES → N-GRAM. Free, and often the best option. Try it first.

□ Does your model have native MTP?
    YES → use it.

□ Is there a small matched model available?
    YES → DRAFT MODEL. The reliable general-purpose choice.

□ Can you train auxiliary heads, and is the extra 1.5x worth
  weeks of work plus retraining on every model update?
    YES → EAGLE-2.
    NO  → stop at the draft model.

Most teams should stop at n-gram or a draft model. The trained variants give another 1.3-1.6x for substantially more engineering and an ongoing retraining obligation.


11. Production implications#

  • Try n-gram first. Free.
  • Gate on batch size (Section VII.13). Non-negotiable.
  • Measure acceptance rate per workload type, continuously.
  • Verify distribution equivalence at temperature 0 (Section V.12).
  • Budget the draft model’s memory as KV cache you gave up.
  • Trained variants create a retraining obligation. Every target model update requires retraining the heads. Factor that into the decision.
  • Tree verification degrades faster with batch than chains.
  • Watch for native MTP in new models — it removes the bolt-on complexity.

12. Common mistakes#

Skipping n-gram. It’s free and often best.

Not gating on batch size. A regression at peak.

Adopting a trained variant without accounting for retraining.

Believing paper speedups. They’re batch-1, on favorable content.

Not measuring acceptance rate in production.

Skipping the rejection-sampling correction. Silent quality change.

Forgetting the KV rollback for rejected tokens (Section VII.13).


13. Hands-on exercise#

A. Implement n-gram. Write the proposer from section 2. Measure acceptance and speedup on four workload types: summarization, code, RAG, creative writing.

B. Compare variants. For a model you can run, benchmark n-gram and a draft model at batch 1, 8, 32, 128. Plot speedup vs batch for each. Where do they cross 1.0?

C. Tree vs chain. Implement chain verification and a simple 2-branch tree. Measure effective acceptance (tokens accepted per target pass) for each.

D. Acceptance by content. Instrument acceptance rate by request type in a mixed workload. Which content types draft well? Does it match section 2’s table?

E. Adaptive k. Implement per-request adaptive speculation length based on recent acceptance. Compare to the best fixed k on a heterogeneous workload.

F. Verify equivalence. At temperature 0, confirm speculative and non-speculative decoding produce byte-identical output over 200 prompts. If not, find the bug in your rejection sampling.


14. Interview questions#

  1. Survey the speculative decoding variants. Which would you try first and why?
  2. Why does n-gram speculation work, and for which workloads?
  3. What is EAGLE’s insight about features versus tokens?
  4. What is tree verification and what does it cost?
  5. Why do all variants degrade with batch size?
  6. What ongoing obligation does a trained variant create?
  7. How would you verify a speculative decoding implementation is correct?

15. Further reading#

  • [ESTABLISHED] Leviathan et al. (2022); Chen et al. (2023)
  • [EMERGING] Cai et al., “Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads” (2024)
  • [EMERGING] Li et al., “EAGLE” (2024) and “EAGLE-2” (2024)
  • [EMERGING] Fu et al., “Lookahead Decoding” (2024)
  • [ESTABLISHED] DeepSeek-V3 technical report — native multi-token prediction
  • Next: 10 — Low-bit inference: FP4 and beyond

↑↓ navigate↵ openesc close