PidokuInfra

KV Cache Optimization

Intermediate Advanced 1h 30m Difficulty 4/5 Topic 11 of 14

Prerequisites V.05, V.06, V.10


1. Problem → Why → Optimization#

PROBLEM   The KV cache limits concurrency, and at long context/high batch it
          dominates decode bandwidth — more than the model weights.
WHY       KV grows as batch × context × layers × kv_heads × head_dim,
          with no upper bound except your context limit.
OPTIMIZE  A family of techniques, each attacking a different term.

This file is the catalogue and the decision guide. Individual techniques get their own files (12 for quantization, XIII for MLA/offload/disaggregation).


2. The catalogue, organized by which term they attack#

KV_bytes = 2 × L × n_kv_heads × head_dim × bytes × context × batch
           ↑   ↑        ↑           ↑         ↑        ↑        ↑
           A   B        C           D         E        F        G

A. Store only K, not V?          → not possible; both are needed
B. Fewer layers with KV          → cross-layer KV sharing [EMERGING]
                                   YOCO, CLA: share KV across layer groups
C. Fewer KV heads                → GQA (4-8x), MQA (h x)      [ESTABLISHED]
                                   MLA (compress to latent)   [ESTABLISHED]
D. Smaller head_dim              → architecture choice; MLA compresses this
E. Fewer bytes per element       → FP8/INT4 KV quantization   [ESTABLISHED] (file 12)
F. Shorter effective context     → sliding window             [ESTABLISHED]
                                   KV eviction (H2O, SnapKV)  [EMERGING]
                                   StreamingLLM (sink+recent) [EMERGING]
G. Fewer sequences in GPU memory → offload to CPU/NVMe        [EMERGING]
                                   disaggregation             [EMERGING]

Plus, orthogonal to the formula:
   PagedAttention          eliminate fragmentation (2.5-4x effective)  [ESTABLISHED]
   Prefix caching          don't store duplicates                      [ESTABLISHED]

3. The decision guide#

START: measure your KV pressure.
  kv_bytes_per_step / weight_bytes_per_step  (Section V.06)
  and KV pool utilization

If KV utilization < 60%:
  → KV is not your problem. Optimize elsewhere.

If KV bytes < weight bytes:
  → weight quantization is the bigger lever. Do that first.

If KV bytes > weight bytes (long context and/or high batch):
  1. Are you using paged KV?           → if not, that's a 2.5-4x win. Do it.
  2. Do requests share prefixes?       → prefix caching. Often huge.
  3. Is the KV in FP16?                → FP8 KV cache. 2x, ~free. (file 12)
  4. Is the model MHA?                 → if you can switch models, GQA is 4-8x
  5. Still constrained?
     ├─ Very long context, low batch   → offload or disaggregate (XIII)
     ├─ Long context, quality tolerant → eviction/compression [EMERGING]
     └─ Otherwise                      → more GPUs

4. Simple analogy#

Managing a hotel’s guest records.

  • GQA/MQA: fewer fields per record. Structural, decided when the system was designed.
  • Quantization: shorter fields — abbreviate rather than write in full.
  • Sliding window: only keep the last 30 days of records.
  • Eviction: keep only the important records, discard the rest.
  • Paging: store records in standard-size folders so no shelf space is wasted.
  • Prefix sharing: one copy of the shared preamble, referenced by many guests.
  • Offloading: move older records to the basement (slower to fetch, but they exist).

Each is a different tradeoff between space, access speed, and information loss. Only eviction loses information.


5. Technical explanation of the main techniques#

GQA / MQA (architectural — [ESTABLISHED])#

Covered in Section III.08 and XIII.01. The lever:

n_kv_heads = h    (MHA)   →  baseline
n_kv_heads = h/8  (GQA-8) →  8x smaller KV, AND 8x higher decode attention intensity
n_kv_heads = 1    (MQA)   →  h× smaller, some quality cost

Two benefits from one change: smaller cache and less memory-bound attention. This is why every model since 2023 uses it.

Sliding window ([ESTABLISHED])#

Cap the KV cache at W tokens instead of S.

Mistral-7B, W=4096:
  context 128k, full attention:  16 GB per sequence
  context 128k, sliding window:   0.5 GB per sequence     32x

Cost: information beyond W is only accessible through the layer chain.

Modern designs interleave: e.g. 1 global layer per 5 sliding-window layers. This caps ~83% of the KV while retaining true long-range access.

Cross-layer KV sharing ([EMERGING])#

Standard:  every layer has its own K, V
CLA/YOCO:  layers 2i and 2i+1 share the same K, V

→ 2x smaller cache
→ quality cost is modest in published results
→ not yet in mainstream production models

Worth watching. If it holds up, it’s another 2x on top of GQA.

KV eviction and compression ([EMERGING] — be careful)#

H2O (Heavy-Hitter Oracle)
  Keep tokens with high cumulative attention scores; evict the rest.
  Retains ~20% of the cache with modest quality loss on some tasks.

SnapKV
  At the end of prefill, use the last few queries' attention to select
  which prompt tokens to keep. Prompt-compression at the KV level.

StreamingLLM
  Keep the first few "attention sink" tokens plus a recent window.
  Enables infinite streaming with bounded memory.
  Note: the model doesn't gain long-range ability — it just doesn't break.

PyramidKV, adaptive budgets
  Allocate more KV budget to lower layers, less to upper ones.

Honest assessment: these work well on benchmarks designed to show them working, and can fail badly on tasks requiring information from evicted tokens. The failure mode is silent — the model confabulates rather than erroring. Deploy only with task-specific validation, and prefer them for tasks where you understand the attention pattern.

The exception is StreamingLLM’s insight about attention sinks, which is well-established and now built into several models’ training.

Prefix caching ([ESTABLISHED])#

Covered in Section V.11. Not a compression technique — a deduplication one. Often the largest practical win for chat and agentic workloads.

Offloading ([EMERGING], hardware-dependent)#

Move cold KV blocks to CPU memory, fetch on demand.

Viability = PCIe bandwidth vs the KV read rate needed.
  PCIe Gen5 x16:   ~50 GB/s
  A decode step needing 30 GB of KV: 600 ms just to fetch. Not viable.
  Fetching 2 GB of cold blocks while computing: viable with prefetch.

  NVLink-C2C (Grace-Hopper): 900 GB/s → genuinely viable.

Do the bandwidth arithmetic before designing around offload. On standard PCIe it works only for cold, prefetchable blocks — not for the active working set. Section XIII.07.


6. Combining techniques#

They compose, with caveats:

Llama-3-70B baseline (MHA equivalent, FP16, full attention, contiguous):
  KV per token: 2.5 MiB
  At 8k context, 640 GB fleet, 141 GB weights: ~24 sequences

+ GQA-8:              320 KiB/token       → ~195 sequences   (8x)
+ Paged:              (eliminates 60-80% waste, already counted in "effective")
+ FP8 KV:             160 KiB/token       → ~390 sequences   (2x)
+ Prefix caching:     (60% hit → ~1.5x effective on shared-prefix workloads)
                                          → ~585 effective
+ Sliding window 4k:  caps at 4k          → context-independent

Total: ~24 → ~585 concurrent sequences. 24x.

Not all are available to you — GQA and sliding window are the model’s choice, not yours. But paged + FP8 KV + prefix caching are all yours, and together they’re typically 4-8x.


7-9. Performance, production, mistakes#

Performance summary:

TechniqueKV reductionQuality costAvailability
GQA-88x~0model architecture
MQAh×smallmodel architecture
MLA10-20x~0 (claimed)model architecture
Sliding window 4kup to 32x at long ctxtask-dependentmodel architecture
Cross-layer sharing2xsmallmodel architecture
PagedAttention2.5-4x effective0yours
FP8 KV2x~0yours
INT4 KV4xsmall-moderateyours
Prefix cachingworkload-dependent0yours
Eviction (H2O etc.)3-5xtask-dependent, riskyyours, carefully
CPU offloadcapacity, not bandwidth0yours, if the link is fast

Production:

  • Measure KV pressure first. If you’re not KV-bound, none of this matters.
  • Do the free ones: paged (already default), prefix caching, FP8 KV.
  • Treat model architecture as a selection criterion: when choosing between models of similar quality, KV size per token is a direct cost multiplier.
  • Be conservative with eviction. Silent quality failures.
  • Report KV bytes per token in your model registry alongside parameter count.

Mistakes:

  • Optimizing KV when weights dominate. Check the ratio.
  • Deploying eviction without task-specific validation.
  • Designing around CPU offload without checking PCIe bandwidth.
  • Setting max_model_len to the model’s maximum — it inflates per-sequence block table sizes and reduces available blocks.
  • Forgetting that prefix caching consumes KV blocks too — cached blocks compete with active ones.

10. Hands-on exercise#

A. Measure your pressure. For your workload, compute KV bytes vs weight bytes per decode step at your p50 and p95 (batch, context). Which dominates? Plot the crossover curve and mark your operating point.

B. Compose the techniques. For a model and GPU of your choice, compute max concurrency under: baseline, +paged, +FP8 KV, +prefix caching at 60% hit rate. Build the table from section 6.

C. Sliding window. Compare a full-attention and a sliding-window model of similar size at 32k context. Measure max concurrency and memory for each.

D. Evaluate eviction honestly. Implement or use a simple eviction scheme (keep first 4 + last N tokens). Evaluate on: a task where the answer is at the end (easy) and one where it’s in the middle of a long document (hard). Quantify the failure.

E. Offload arithmetic. For your hardware, compute the maximum KV fetch rate over PCIe. Compare to the KV read rate needed at your batch and context. Is offload viable? For what fraction of the cache?


11. Interview questions#

  1. Write the KV size formula and name a technique that attacks each term.
  2. Why does GQA give two benefits, not one?
  3. When does KV cache optimization matter more than weight quantization?
  4. What is sliding-window attention and what does it cost?
  5. Why are KV eviction methods risky in production?
  6. How would you determine whether CPU KV offload is viable on your hardware?
  7. Rank the KV optimizations you can apply without changing the model.

12. Further reading#

  • [ESTABLISHED] Ainslie et al., “GQA” (2023); Shazeer, “MQA” (2019)
  • [ESTABLISHED] Jiang et al., “Mistral 7B” (2023) — sliding window
  • [ESTABLISHED] DeepSeek-V2 technical report — MLA
  • [EMERGING] Zhang et al., “H2O” (2023); Li et al., “SnapKV” (2024); Xiao et al., “StreamingLLM” (2023)
  • [EMERGING] Brandon et al., “Reducing Transformer Key-Value Cache Size with Cross-Layer Attention” (2024)
  • Next: 12 — KV cache quantization

↑↓ navigate↵ openesc close