1. The problem, and the shape of the solution space#
KV_bytes = 2 × L × h_kv × d_head × bytes × S × B
Every technique attacks one term. The research divides cleanly:
TERM ATTACKED APPROACH STATUS
L (layers) cross-layer sharing EMERGING
h_kv (heads) GQA/MQA/MLA SOLVED (Section XIII.01)
bytes (precision) KV quantization ESTABLISHED
S (positions) eviction, compression, RESEARCH — the risky one
sliding window (window: ESTABLISHED)
B (sequences) better memory management SOLVED (PagedAttention)
allocation paging SOLVED
duplication prefix caching SOLVEDThree of seven are solved, two are established, and the interesting research is concentrated in the one with the worst failure mode.
2. Solved: paging and sharing#
PagedAttention (2023) eliminated fragmentation; enabled sharing
Prefix caching eliminated duplicate computation
RadixAttention (SGLang) tree-structured prefix cache
STATUS: universal. These are the foundation everything else builds on.
REMAINING WORK
~ cross-replica prefix sharing (Section IX.11) — routing solves
most of it (Section XII.07)
~ hierarchical/tiered caches (Section XIII.07)3. Established: KV quantization#
FP8 KV cache 2x, ~free on Hopper. Production standard.
INT8 KV cache 2x, needs care. Production-common.
KIVI (2024) asymmetric: per-channel K, per-token V,
2-bit demonstrated
KVQuant (2024) non-uniform quantization, outlier handling
THE KEY FINDING (independently reproduced): K is more sensitive
than V, because K passes through the exponential in softmax while
V is averaged (Section VII.12).
→ asymmetric treatment: K at higher precision than V
→ per-channel for K, per-token for V
STATUS: FP8 is standard. Sub-4-bit KV is EMERGING and workload-dependent.This line is productive and low-risk because quantization error is bounded and measurable, unlike eviction.
4. The risky area: eviction and compression#
THE PROMISE: 3-5x KV reduction by keeping only "important" tokens.
H2O (2023) keep tokens with high cumulative attention scores
SnapKV (2024) select prompt tokens using the last queries' attention
PyramidKV (2024) layer-dependent budgets (more for lower layers)
StreamingLLM (2023) first few tokens (sinks) + a recent window
Scissorhands, FastGen, and many others
THE PROBLEM (Section XIII.04, stated again because it matters):
quality loss is TASK-DEPENDENT and the failure is SILENT.
If the model needed an evicted token, it confabulates.WHY THE BENCHMARKS LOOK GOOD AND PRODUCTION DOESN'T
Common benchmarks (perplexity, summarization, most QA):
the relevant information is recent or highly attended
→ eviction based on attention scores keeps it
→ results look excellent
The failure case:
a specific fact from the middle of a long document, needed by
a question asked at the end
→ low attention during prefill (nobody asked yet)
→ evicted
→ the model confabulates an answer
→ the evaluation must SPECIFICALLY test this, and most don't.What would change the assessment:
□ a method with a BOUNDED failure mode (a guarantee about what is
retained, not a heuristic)
□ a reliable detector for "I evicted something I needed"
□ evaluation that specifically constructs the failure case
□ deployment at scale by a major provider with published resultsNone of these exist yet as of this writing. That’s why this remains RESEARCH despite three years of papers.
StreamingLLM is the exception worth separating out: its contribution — that the first few tokens act as attention sinks and must be retained — is a robust finding that has been incorporated into model training. The bounded-memory streaming application is legitimate; the claim is narrower (“the model won’t break”) than eviction’s (“you lose nothing”).
5. Emerging: cross-layer sharing#
CLA (Cross-Layer Attention, 2024) adjacent layers share K,V
YOCO (You Only Cache Once, 2024) one global KV, all layers attend to it
→ 2x (CLA with pairs) to Lx (YOCO) reduction
→ orthogonal to GQA: composes multiplicatively
→ architectural: requires training with it
STATUS: promising, limited production validation.
Watch for adoption in a major model family.If CLA-style sharing composes with GQA-8 and FP8 KV, the combined reduction is 8 × 2 × 2 = 32x versus MHA at FP16 — which would substantially change long-context economics.
6. Systems-level: tiering and disaggregation#
Mooncake (2024) KV cache as a first-class distributed store
LMCache a KV caching layer across engines and tiers
CacheGen, CacheBlend compressing KV for network transfer
InfiniGen, and others offloading with prefetch
THE UNIFYING IDEA: KV is data with a lifecycle, and it should live
in a tiered store rather than only in GPU memory.
STATUS: EMERGING. The arithmetic (Section XIII.07) constrains it:
fetching must beat recomputing, which requires a fast
interconnect.
WHAT WOULD CHANGE IT
→ coherent CPU-GPU memory (Grace-Hopper class) makes the CPU tier
genuinely fast, which changes the calculation by 16x7. What to watch, ranked#
1. CROSS-LAYER KV SHARING IN A MAJOR MODEL
→ 2x on the binding constraint, low risk, architectural
→ watch: model releases
2. COHERENT CPU-GPU MEMORY AT SCALE
→ makes tiered KV practical
→ watch: Grace-Hopper deployment, and successors
3. A BOUNDED-FAILURE KV COMPRESSION METHOD
→ would move eviction from RESEARCH to viable
→ watch: methods with guarantees, not heuristics
4. SUB-4-BIT KV QUANTIZATION VALIDATION
→ 4x instead of 2x, with the quality question answered
→ watch: independent long-context retrieval evaluations
5. STANDARDIZED KV FORMAT / INTERCHANGE
→ would enable cross-engine and cross-tier KV movement
→ watch: whether the engines converge on one8. What is unlikely to matter#
✗ Another attention-score-based eviction heuristic
The space is saturated; the failure mode is unchanged.
✗ KV compression evaluated only on summarization and perplexity
Those tasks don't exercise the failure case.
✗ Learned KV compression requiring per-model training
The training cost exceeds the benefit for most deployments.
✗ Methods requiring a custom attention kernel that doesn't exist
outside the paper9. Hands-on exercise#
A. Construct the failure case. Build an evaluation that specifically tests eviction’s failure mode: a fact placed early in a long context, with a question asked at the end that requires it. Run it against an eviction method. Does the reported “minimal quality loss” hold?
B. Compose the reductions. Compute the KV size for a hypothetical model with MHA/FP16 versus GQA-8 + CLA + FP8 KV. What’s the combined factor? What does it do to concurrency at 128k context?
C. Reproduce the K/V asymmetry. Quantize only K to 4 bits, then only V, then both. Measure quality on a retrieval task. Confirm K is more sensitive.
D. Tier arithmetic. For your hardware, compute the fetch-vs-recompute crossover (Section XIII.07) for CPU DRAM. Then recompute assuming NVLink-C2C bandwidth. How does the conclusion change?
E. Evaluate a paper. Take a recent KV compression paper. Apply the seven questions (Section XIV.01). Does it test the failure case?
10. Interview questions#
- Which KV cache problems are solved and which are open?
- Why is K more sensitive to quantization than V?
- Why have KV eviction methods not been adopted despite many papers?
- What would make eviction production-viable?
- What is cross-layer KV sharing and why does it matter?
- What hardware change would make tiered KV caching practical?
- How would you evaluate a KV compression method properly?
11. Further reading#
- [ESTABLISHED] Kwon et al., PagedAttention (2023)
- [EMERGING] Liu et al., “KIVI” (2024); Hooper et al., “KVQuant” (2024)
- [EMERGING] Brandon et al., “Cross-Layer Attention” (2024); Sun et al., “YOCO” (2024)
- [RESEARCH] Zhang et al., “H2O” (2023); Li et al., “SnapKV” (2024)
- [EMERGING] Xiao et al., “StreamingLLM” (2023) — the attention sink finding
- [EMERGING] Qin et al., “Mooncake” (2024)
- Next: 04 — Quantization research