1. The two problems#
PROBLEM A — SPEED
decode is serial and memory-bound (Section V.03).
→ speculative decoding and its variants (Section XIII.09)
→ multi-token prediction
→ parallel decoding schemes
PROBLEM B — QUALITY
which token to pick, and how sampling affects output quality.
→ sampling methods (Section V.13)
→ constrained/structured decoding
→ inference-time scaling (thinking longer)These are usually studied separately but interact: inference-time scaling makes decode speed matter much more.
2. Speed: the state of speculative decoding#
Covered in Section XIII.09. The research summary:
SETTLED
✓ the algorithm and its correctness proof (Leviathan, Chen 2023)
✓ that it trades compute for bandwidth
✓ that it degrades with batch size
✓ n-gram/lookup as a free variant
ACTIVE
~ better draft models (distillation, architecture)
~ tree construction (EAGLE-2's dynamic trees)
~ feature-level vs token-level drafting
~ making it work at high batch
DIMINISHING
- marginally better draft heads
- variants requiring training, without addressing the batch problemThe batch problem is the open question. Every variant degrades toward 1.0x speedup as batch grows, because verification compute stops being free. A method that works at batch 128 would be genuinely new.
WHY THE BATCH PROBLEM IS HARD
speculation trades COMPUTE for BANDWIDTH.
at batch 1: compute is 99% idle → the trade is free
at batch 128: compute is 40% used → the trade costs real time
→ the only escape is a draft that costs approximately nothing
(n-gram) or a target verification that's cheaper than k separate
forward passes by more than the draft costs3. Multi-token prediction — the cleaner path#
INSTEAD OF: bolting speculation onto a model trained for one token
DO: train the model to predict several tokens
DeepSeek-V3 includes MTP modules.
Gloeckle et al. (2024) show multi-token training also improves
the model's quality, not just its speed.
WHY THIS IS BETTER
✓ no separate draft model or trained heads
✓ the predictions are aligned by construction
✓ the training signal improves the base model
✗ requires training with it; can't retrofit
STATUS: shipping in some models. Likely to become standard.This is where the field should go, and the fact that multi-token training also improves quality (not just speed) makes it likely.
4. Quality: sampling#
SETTLED
✓ pure greedy is bad for open-ended generation (Holtzman 2019)
✓ nucleus (top-p) sampling as the practical default
✓ temperature as the diversity control
REFINEMENTS THAT MATTER
min-p sampling: threshold relative to the max probability
→ adapts to the model's confidence, unlike fixed top-p
→ genuinely better across confidence levels (Section V.13)
typical sampling, eta/epsilon sampling, and others
→ marginal; not widely adopted
ACTIVE
~ sampling for reasoning tasks (does temperature help or hurt
chain-of-thought?)
~ adaptive sampling based on position or contentmin-p is the one refinement worth knowing. It’s simple, better-motivated than top-p, and increasingly supported.
5. Constrained and structured decoding#
THE PROBLEM: force output to match a grammar (JSON schema, regex,
a programming language).
THE NAIVE APPROACH: mask disallowed tokens at each step.
→ 0.5-5 ms per token to compute the mask. Can dominate ITL.
THE ADVANCES
compressed FSM (SGLang) precompute masks per automaton state
jump-forward decoding when only one continuation is valid,
emit it WITHOUT running the model
XGrammar, llguidance fast mask computation, often as bitmasks
→ 0.02 ms per token, and fewer model steps
WHY JUMP-FORWARD MATTERS
in JSON, much of the output is literal structure:
{"name": "..."}
^^^^^^^^^^ the model doesn't need to generate this
→ skipping it saves time AND removes the chance of getting it wrongSTATUS: ESTABLISHED. This is a solved engineering problem with
good implementations available.
OPEN
~ constrained decoding that doesn't distort the distribution
(masking changes the probabilities — is the result still the
model's "intent"?)
~ efficient handling of very large grammars
~ constraints that depend on generated content (semantic, not syntactic)The distribution-distortion question is real and under-discussed: forcing a token the model assigned low probability to is a form of intervention, and the downstream generation is conditioned on a token the model wouldn’t have chosen.
6. Inference-time scaling — the big shift#
THE IDEA: spend more compute at inference to get better answers.
chain-of-thought, self-consistency (sample n, take the majority),
tree search, verifier-guided search, extended "thinking" before
answering
WHY IT MATTERS FOR INFERENCE ENGINEERING
it changes the WORKLOAD SHAPE fundamentally:
before: 500 input, 200 output
with reasoning: 500 input, 5,000-50,000 output
→ decode becomes overwhelmingly dominant
→ per-token cost matters far more
→ speculative decoding becomes far more valuable
→ KV cache grows much larger per request
→ latency budgets change (users accept 30 s for a hard question)CONSEQUENCES FOR THE STACK
1. DECODE OPTIMIZATION becomes the dominant concern
everything in Sections VII and XIII that speeds decode
is worth more
2. KV CACHE PRESSURE increases dramatically
a 50,000-token reasoning trace has a 50,000-token KV cache
→ GQA/MLA, KV quantization, and long-context techniques
become critical
3. BATCHING CHANGES
long generations mean sequences occupy slots for minutes
→ concurrency is limited by slot-time, not arrival rate
→ Little's Law with much larger W
4. NEW OPTIMIZATION OPPORTUNITIES
~ can you cache/reuse reasoning traces?
~ can you speculate on reasoning steps?
~ can you prune search branches early?
~ parallel sampling shares the prompt's KV (Section V.10) —
self-consistency is cheap in KV termsThis is the most consequential shift in the field for inference engineers. A workload that was 30% prefill and 70% decode becomes 3% prefill and 97% decode, and every optimization’s value changes accordingly.
7. What to watch, ranked#
1. INFERENCE-TIME SCALING ADOPTION
→ changes workload shape more than any technique changes efficiency
→ watch: product behavior, not papers
2. NATIVE MULTI-TOKEN PREDICTION IN MODELS
→ cleaner than bolted-on speculation
→ watch: model releases
3. SPECULATIVE DECODING AT HIGH BATCH
→ currently unsolved; would be significant
→ watch: methods that address the compute cost of verification
4. REASONING-TRACE REUSE / CACHING
→ if reasoning traces for similar problems can be reused,
the cost of inference-time scaling drops substantially
→ watch: early work in this area
5. CONSTRAINED DECODING WITHOUT DISTRIBUTION DISTORTION
→ a correctness question that becomes important as structured
output becomes universal8. What is unlikely to matter#
✗ Another sampling method
top-p and min-p cover the practical space.
✗ Speculative decoding variants that don't address the batch problem
They're all 1.1-1.3x at production batch sizes.
✗ Beam search revivals for open-ended generation
Holtzman's finding stands; and the KV cost is prohibitive.
✗ Constrained decoding methods slower than the existing fast ones9. Hands-on exercise#
A. Measure the workload shift. For a reasoning model versus a standard model on the same questions, measure the output-length distribution. How does the prefill/decode time split change?
B. Recompute your optimization priorities. With the reasoning workload’s shape, recompute which optimizations matter (Section VII.01’s ranking). What moved?
C. KV pressure. For a 20,000-token reasoning trace, compute the KV cache size and the resulting concurrency. Compare to a 200-token response.
D. Self-consistency KV sharing. Generate n=8 samples from one prompt with and without KV sharing (Section V.10). Measure the memory difference.
E. Constrained decoding cost. Measure ITL with a naive masking implementation and with a fast one (XGrammar or SGLang). Quantify the difference.
F. Jump-forward. For a JSON-schema-constrained generation, count how many output tokens were literal structure that jump-forward decoding could skip.
10. Interview questions#
- Why does speculative decoding degrade with batch size, and what would fix it?
- What is multi-token prediction and why is it cleaner than bolted-on speculation?
- What is jump-forward decoding?
- How does inference-time scaling change the inference workload?
- Which optimizations become more valuable with reasoning models?
- What is min-p sampling and why is it better than top-p?
- What’s the correctness concern with constrained decoding?
11. Further reading#
- [FUNDAMENTAL] Holtzman et al., “The Curious Case of Neural Text Degeneration” (2019)
- [ESTABLISHED] Leviathan et al. (2022); Chen et al. (2023)
- [EMERGING] Gloeckle et al., “Better & Faster Large Language Models via Multi-token Prediction” (2024)
- [ESTABLISHED] Zheng et al., “SGLang” — compressed FSM and jump-forward decoding
- [EMERGING] Work on inference-time compute scaling and its serving implications
- Next: 06 — MoE research