PidokuInfra

KV Offloading and Transfer

Expert Advanced 1h 15m Difficulty 4/5 Topic 07 of 12

Prerequisites II.09, V.05, IX.11


1. The idea#

Move KV cache blocks out of GPU memory to a cheaper, larger tier, and fetch them back when needed.

TIER            CAPACITY      BANDWIDTH           LATENCY    $/GB
GPU HBM         80-141 GB     3,350 GB/s          0.5 µs     high
CPU DRAM        0.5-2 TB      55 GB/s (PCIe5)     10 µs      medium
                              900 GB/s (NVLink-C2C)
Local NVMe      1-100 TB      3-14 GB/s           100 µs     low
Remote store    unbounded     1-50 GB/s           1 ms+      lowest

The question is always the same: is fetching cheaper than recomputing?

Diagram — KV cache as a tiered memory#

flowchart TB
  G[("GPU HBM<br/>hot: active decodes")] -->|"evict idle sessions"| C[("CPU DRAM<br/>warm")]
  C -->|"evict"| N[("Local NVMe<br/>cold")]
  N -->|"evict"| S[("Remote / shared store<br/>cross-node reuse")]
  C -.->|"restore"| G
  N -.->|"restore"| G
  S -.->|"restore"| G
  Q{"Restore time below<br/>re-prefill time?"} -->|"yes"| RST["Restore"]
  Q -->|"no"| REC["Recompute"]

  class G,C memory
  class N,S io
  class Q queue
  class RST memory
  class REC compute

2. The governing arithmetic#

RECOMPUTE COST (re-prefill the sequence)
  T_recompute = 2 × P × S / achieved_FLOPs

FETCH COST
  T_fetch = KV_bytes / tier_bandwidth
          = (2 × L × h_kv × d_head × bytes × S) / bandwidth

CROSSOVER: fetch is worth it when T_fetch < T_recompute
WORKED — Llama-3-70B, GQA-8, FP8 KV (160 KiB/token), 8×H100:
  achieved prefill: 50,000 tok/s

  S = 1,000 tokens:
    recompute:  1000/50000 = 20 ms
    KV size:    160 MB
    fetch PCIe5 (55 GB/s):   2.9 ms    ✓ 7x cheaper
    fetch NVMe (7 GB/s):     23 ms     ✗ slower than recompute
    fetch remote (2 GB/s):   80 ms     ✗ much slower

  S = 20,000 tokens:
    recompute:  400 ms
    KV size:    3.2 GB
    fetch PCIe5:             58 ms     ✓ 7x cheaper
    fetch NVMe:             457 ms     ✗ marginally slower
    fetch remote:         1,600 ms     ✗ much slower

Two conclusions:

  1. CPU DRAM over PCIe is roughly 7x cheaper than recompute — a consistent, useful margin.
  2. NVMe and remote storage are NOT faster than recompute for these parameters. They buy capacity, not speed.

That second point is frequently missed. Offloading to disk does not make things faster; it makes more things possible.


3. What offloading is actually for#

IT IS NOT FOR: making the active working set faster.
  The active sequences' KV must be in HBM — you read all of it
  every decode step.

IT IS FOR:

1. PREFIX CACHE CAPACITY
   Cached prefixes that aren't currently in use don't need to be in
   HBM. Keep them in CPU DRAM; fetch on a hit.
   → dramatically larger effective prefix cache
   → THE PRIMARY USE CASE

2. PREEMPTION ALTERNATIVE
   Instead of discarding a preempted sequence's KV (recompute later),
   swap it to CPU and swap back.
   → only worth it if fetch < recompute (see the arithmetic above)
   → vLLM defaults to recompute because the margin is thin and
     recompute is simpler

3. LONGER CONTEXT THAN HBM ALLOWS
   Keep distant KV blocks in CPU, fetch as attention needs them.
   → requires attention that can work with partially-resident KV
   → complex; limited production use

4. CROSS-REQUEST SHARING
   A shared KV store that multiple replicas can read from
   (Section IX.11).

Use case 1 is where the value is. A 2 TB CPU DRAM tier holding cached prefixes is worth far more than trying to make active sequences work from CPU.


4. Grace-Hopper changes the calculation#

STANDARD:      CPU ←PCIe Gen5 (55 GB/s)→ GPU
GRACE-HOPPER:  CPU ←NVLink-C2C (900 GB/s)→ GPU     16x faster

At 900 GB/s:
  S = 20,000, KV = 3.2 GB:
    fetch:      3.6 ms
    recompute:  400 ms
    → 111x cheaper. Offload becomes trivially worthwhile.

  And CPU memory is 480 GB per Grace chip, vs 80-141 GB of HBM.
  → the CPU tier becomes a genuine extension of GPU memory

On Grace-Hopper (and similar coherent CPU-GPU designs), KV offload stops being a compromise and becomes an architecture. This is one of the more significant hardware developments for inference, and it’s why the technique is worth understanding even if it’s marginal on your current hardware.


5. Implementation#

THE MACHINERY

1. PINNED HOST BUFFER POOL (Section II.06)
   pre-allocated, page-locked, reused
   → without pinning, transfers are 2x slower and synchronous

2. ASYNC TRANSFER ON A SEPARATE STREAM (Section VI.06)
   cudaMemcpyAsync on a copy stream
   → overlaps with compute

3. PREFETCH
   predict which blocks will be needed and fetch them early
   → for prefix cache: fetch on admission, before the request runs
   → for preemption swap: fetch when the sequence is re-admitted

4. EVICTION POLICY
   GPU → CPU: LRU over unreferenced blocks
   CPU → disk: LRU, or drop entirely
   
5. BOOKKEEPING
   block → tier mapping, refcounts across tiers, in-flight transfers
Python
class TieredBlockManager:
    def __init__(self, gpu_blocks, cpu_blocks):
        self.gpu = BlockPool(gpu_blocks)
        self.cpu = BlockPool(cpu_blocks, pinned=True)
        self.location = {}            # block_hash → ("gpu"|"cpu", block_id)
        self.copy_stream = torch.cuda.Stream()

    def get_for_prefill(self, block_hashes):
        """Return GPU blocks, fetching from CPU where needed."""
        gpu_blocks, to_fetch = [], []
        for h in block_hashes:
            tier, bid = self.location.get(h, (None, None))
            if tier == "gpu":
                gpu_blocks.append(bid); self.gpu.touch(bid)
            elif tier == "cpu":
                dst = self.gpu.allocate(1)[0]
                to_fetch.append((bid, dst))
                gpu_blocks.append(dst)
                self.location[h] = ("gpu", dst)
            else:
                return None            # cache miss; must compute
        if to_fetch:
            with torch.cuda.stream(self.copy_stream):
                for src, dst in to_fetch:
                    self.gpu.data[dst].copy_(self.cpu.data[src],
                                             non_blocking=True)
            torch.cuda.current_stream().wait_stream(self.copy_stream)
        return gpu_blocks

    def evict_to_cpu(self, n_blocks):
        """Make room in GPU by demoting unreferenced blocks."""
        victims = self.gpu.lru_unreferenced(n_blocks)
        for bid in victims:
            dst = self.cpu.allocate(1)[0]
            with torch.cuda.stream(self.copy_stream):
                self.cpu.data[dst].copy_(self.gpu.data[bid], non_blocking=True)
            self.location[self.gpu.hash_of(bid)] = ("cpu", dst)
            self.gpu.free([bid])

6. When it’s worth it#

CHECKLIST

□ Is your prefix cache hit rate limited by CAPACITY?
    measure: how many blocks are evicted per hour? What's the hit rate?
    If the cache is never full, offloading adds nothing.

□ Compute the fetch-vs-recompute margin for your parameters.
    If fetch isn't at least 3x cheaper, the complexity isn't worth it.

□ Do you have host memory to spare?
    Offloading needs pinned memory, which competes with everything else.

□ Do you have the interconnect?
    PCIe Gen4: marginal.  Gen5: workable.  NVLink-C2C: excellent.

□ Have you done the cheaper things first?
    FP8 KV cache (2x for free), GQA model selection, prefix-aware
    routing (Section XII.07).

“Have you done the cheaper things first” rules this out for most deployments. FP8 KV cache gives 2x more effective cache capacity for a config flag; offloading gives more but costs weeks.


7. The remote/distributed variant#

A SHARED KV STORE across replicas (Section IX.11)

  ✓ one copy of a popular prefix, globally
  ✓ any replica can serve any prefix
  ✗ network fetch: 1-50 GB/s, 1 ms+ latency
  ✗ must be faster than recompute — often it isn't
  ✗ consistency, eviction, and failure handling across nodes

VIABILITY
  fetch over 400G InfiniBand (50 GB/s):
    S=20,000, KV=3.2 GB → 64 ms vs 400 ms recompute. ✓ viable
  fetch over 25G Ethernet (3 GB/s):
    → 1,067 ms. ✗ not viable

→ same conclusion as disaggregation (file 06): needs RDMA-class fabric.

Prefix-aware routing (Section XII.07) achieves much of the same benefit with none of the infrastructure, by ensuring requests go to the replica that already has the prefix. Do that first.


8. Production implications#

  • Do FP8 KV cache first. 2x for a flag (Section VII.12).
  • Do prefix-aware routing first. Solves the cross-replica sharing problem cheaply.
  • Measure whether your prefix cache is capacity-limited before building this.
  • CPU DRAM offload for prefix cache is the viable use case. Not for active sequences, not to disk.
  • Compute the fetch-vs-recompute margin for your specific parameters.
  • Pinned buffer pools and async streams are required for acceptable transfer performance.
  • Grace-Hopper-class hardware changes the calculus substantially — reassess if you have it.

9. Common mistakes#

Offloading active sequences’ KV. They’re read every step; they must be in HBM.

Offloading to disk expecting speed. NVMe is slower than recomputing.

Not doing FP8 KV cache first.

Not pinning host buffers. 2x slower, synchronous.

Building a remote KV store instead of prefix-aware routing.

Not computing the fetch-vs-recompute margin.

Underestimating the bookkeeping complexity — tier tracking, refcounts, in-flight transfers, and failure handling are all real work.


10. Hands-on exercise#

A. Compute your margin. For your model, hardware, and typical prefix lengths, compute fetch-vs-recompute for CPU DRAM, NVMe, and a remote store. Which are viable?

B. Measure your cache pressure. Instrument the prefix cache: hit rate, eviction rate, and whether the pool is ever full. Is capacity actually your limit?

C. Implement the tier. Build the TieredBlockManager from section 5. Measure: fetch latency, transfer bandwidth achieved, and the hit-rate improvement from the larger effective cache.

D. Pinned vs pageable. Measure KV block transfer with pinned and pageable host buffers. Confirm the 2x difference and the synchronization behavior.

E. Prefetch. Implement prefetch-on-admission for prefix blocks. Measure the TTFT difference versus fetching lazily during prefill.

F. Compare to routing. For a multi-replica setup, compare (i) a shared KV store and (ii) prefix-aware routing. Which gives a better hit rate, and at what complexity?


11. Interview questions#

  1. When is fetching KV cheaper than recomputing it? Give the arithmetic.
  2. Why can’t you offload active sequences’ KV cache?
  3. What is KV offloading actually for?
  4. Why is NVMe offload not a speed optimization?
  5. How does Grace-Hopper change the calculation?
  6. What should you do before building KV offload?
  7. Why is prefix-aware routing often a better answer than a shared KV store?

12. Further reading#

  • [EMERGING] Qin et al., “Mooncake” (2024) — KV-centric architecture
  • [EMERGING] LMCache and vLLM KV connector documentation
  • [ESTABLISHED] Sheng et al., “FlexGen” (2023) — aggressive offloading for throughput-oriented single-GPU inference
  • [REFERENCE] NVIDIA Grace-Hopper architecture documentation
  • Next: 08 — Request scheduling and priorities

↑↓ navigate↵ openesc close