PidokuInfra

Memory Fragmentation and Allocators

Advanced 1h 15m Difficulty 4/5 Topic 05 of 11

Prerequisites II.06, V.10


1. The problem#

You have free memory. You cannot use it.

CUDA out of memory. Tried to allocate 2.00 GiB.
GPU 0 has a total capacity of 79.15 GiB of which 3.21 GiB is free.
Process has 75.94 GiB memory in use. Of the allocated memory
68.42 GiB is allocated by PyTorch, and 6.11 GiB is reserved by PyTorch
but unallocated.
                          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                          6 GB of fragmentation

“Reserved but unallocated” is fragmentation. PyTorch holds it, your tensors don’t use it, and it can’t satisfy your 2 GB request because it’s in pieces.


2. Why it happens#

PyTorch’s caching allocator:

1. cudaMalloc is SLOW (~100 µs) and SYNCHRONIZING.
   So PyTorch requests large segments and sub-allocates from them.
2. Freed tensors return to PyTorch's pool, NOT to the driver.
3. Segments are split to satisfy requests.
4. Adjacent free blocks within a segment are coalesced;
   blocks in DIFFERENT segments cannot be merged.
5. Over time with varying request sizes, segments become checkerboarded.
Segment (2 GB):  [used 512MB][free 256MB][used 768MB][free 384MB][used 128MB]
                              ↑                       ↑
                       640 MB free total, largest contiguous piece: 384 MB
                       → a 512 MB request FAILS

This is exactly the external fragmentation from Section II.06 — and it’s the same problem PagedAttention solves for the KV cache. But the KV cache is only part of your memory; the rest still uses the general allocator.


3. Simple analogy#

A car park with painted bays of fixed sizes.

Cars come and go. Eventually you have 40 free spaces scattered around, but a coach needs 6 consecutive bays and there’s no run of 6 anywhere. The car park is 60% full and you’re turning away a coach.

The fixes map to the real ones:

  • Repaint the bays uniformly (fixed-size blocks — PagedAttention).
  • Reserve a coach area (preallocate the KV pool).
  • Ask cars to park compactly (expandable_segments).
  • Empty the car park and start over (empty_cache() — disruptive).

4. Where fragmentation comes from in inference#

SOURCE                                    SEVERITY
Variable prompt lengths → variable         HIGH  ← the main one
  activation tensor sizes
Variable batch sizes                       MEDIUM
Growing KV (if not paged)                  FATAL (this is why paging exists)
LoRA adapter load/unload                   MEDIUM
Speculative decoding (variable accepted)   MEDIUM
Beam search fork/free                      MEDIUM
torch.compile recompilation buffers        LOW
Long-running processes (accumulation)      grows over time

Variable-length prefill is the dominant source. A 200-token prompt and a 20,000-token prompt allocate activation tensors differing by 100x, and the allocator sees an unpredictable sequence of sizes.


5. The fixes#

Fix 1 — preallocate the KV pool (the big one)#

Engines claim a fixed fraction of GPU memory at startup for the KV cache
and manage it themselves with fixed-size blocks.

vLLM: --gpu-memory-utilization 0.90

→ The KV cache — the largest and most dynamic consumer — is removed from
  the general allocator entirely.
→ This is why production engines don't suffer from KV fragmentation.

Fix 2 — expandable_segments#

Shell
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
Instead of fixed-size segments, use CUDA's virtual memory API
(cuMemCreate / cuMemMap) to create segments that can GROW.

→ A segment can expand to absorb a larger request rather than
  requiring a new segment.
→ Substantially reduces fragmentation for varying shapes.
→ Small overhead; occasionally incompatible with some libraries.

For inference with variable sequence lengths, this is usually a clear win.

Fix 3 — bucket your shapes#

Round activation shapes up to buckets: 128, 256, 512, 1024, 2048, ...
→ the allocator sees a small set of sizes
→ freed blocks are reusable for the next request of the same bucket
→ costs some wasted compute on the padding

This is what CUDA graph capture lists do implicitly (Section VII.08).

Fix 4 — other allocator tuning#

Shell
PYTORCH_CUDA_ALLOC_CONF=\
expandable_segments:True,\
max_split_size_mb:512,\        # don't split segments larger than this
garbage_collection_threshold:0.8   # proactively release above 80% usage

max_split_size_mb prevents large segments from being split for small requests, which preserves large contiguous regions.

Fix 5 — empty_cache(), carefully#

Python
torch.cuda.empty_cache()     # return unused segments to the driver
✓ Genuinely frees fragmented segments
✗ SYNCHRONIZES (blocks until the GPU is idle)
✗ Future allocations pay cudaMalloc cost again
✗ Can cause a latency spike

USE: once after model load, or between distinct phases.
NEVER: in the request path.

6. Diagnosing it#

Python
print(torch.cuda.memory_summary())
|===========================================================================|
|                  PyTorch CUDA memory summary                              |
|---------------------------------------------------------------------------|
|        CUDA OOMs: 3            |        cudaMalloc retries: 12             |
|===========================================================================|
| Metric                  | Cur Usage | Peak Usage | Tot Alloc  | Tot Freed  |
|-------------------------|-----------|------------|------------|------------|
| Allocated memory        |  68.42 GB |   71.20 GB |   4.21 TB  |   4.14 TB  |
| Reserved memory         |  74.53 GB |   74.53 GB |  102.3 GB  |   27.8 GB  |
| Non-releasable memory   |   6.11 GB |    8.02 GB |            |            |
| Allocations             |     18432 |      21044 |            |            |
| Segments                |       412 |        438 |            |            |
|===========================================================================|

The two numbers to read:

  • Reserved − Allocated = 6.11 GB of fragmentation.
  • cudaMalloc retries: 12 — the allocator had to free and retry. Any nonzero value means pressure.

For detail:

Python
torch.cuda.memory._record_memory_history(max_entries=100000)
# ... run the workload ...
torch.cuda.memory._dump_snapshot("snapshot.pickle")
# view at https://pytorch.org/memory_viz

The memory visualizer is excellent — it shows every allocation’s lifetime and lets you see the fragmentation pattern directly. Worth learning.


7. Under the hood — why paging solves it completely#

GENERAL ALLOCATOR                    PAGED KV CACHE
variable-size requests               fixed-size blocks (16 tokens)
external fragmentation possible      IMPOSSIBLE (all blocks interchangeable)
coalescing needed                    not needed
allocation can fail with free memory allocation succeeds if ANY blocks free

Fixed-size allocation eliminates external fragmentation by construction. That’s the whole insight of PagedAttention (Section V.10), and it’s why the KV cache — the most dynamic consumer — is the one part of inference memory that doesn’t fragment.

The activations and workspaces still use the general allocator, which is why expandable_segments still matters.


8. Production implications#

  • Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True for any workload with variable shapes. It’s a one-line change with real benefit.
  • Preallocate the KV pool — engines do this; verify gpu_memory_utilization is set.
  • Monitor reserved - allocated and cudaMalloc retries as health metrics. Growing fragmentation over a long-running process is a real failure mode.
  • Never call empty_cache() in the request path.
  • Cap max input length so the largest activation allocation is bounded.
  • Restart long-running processes periodically if fragmentation accumulates and you can’t fix it otherwise. Inelegant but effective.
  • Test with your longest realistic prompt before setting gpu_memory_utilization high.

9. Common mistakes#

Interpreting nvidia-smi memory as tensor memory. It shows the allocator’s reservation.

Raising gpu_memory_utilization to 0.97 and then OOMing on a long prefill’s activation spike.

empty_cache() in the hot loop. Synchronizes and slows everything.

Not setting expandable_segments for variable-shape workloads.

Assuming OOM means “not enough memory.” Read memory_summary() — it’s often fragmentation.

Ignoring gradual growth. A process that OOMs after 6 hours has an accumulation problem.


10. Hands-on exercise#

A. Reproduce fragmentation. Allocate many tensors of random sizes on GPU, free every other one, then request a large contiguous tensor. Trigger an OOM with plenty of free memory. Print memory_summary() and identify the fragmentation.

B. Fix it. Re-run with expandable_segments:True. Does the OOM go away? Measure the allocation overhead difference.

C. Visualize. Use _record_memory_history and the PyTorch memory visualizer on a real inference workload with variable prompt lengths. Screenshot the fragmentation pattern.

D. The bucketing effect. Run a workload with (i) arbitrary prompt lengths and (ii) lengths rounded to buckets. Compare reserved - allocated after 1,000 requests.

E. Long-run stability. Run a server for several hours with realistic variable traffic. Plot reserved - allocated over time. Does fragmentation accumulate?


11. Interview questions#

  1. Why does PyTorch cache GPU memory instead of returning it to the driver?
  2. What is “reserved but unallocated” and what causes it?
  3. What does expandable_segments do?
  4. Why does fixed-size block allocation eliminate external fragmentation?
  5. When should you call empty_cache() and when shouldn’t you?
  6. You get a CUDA OOM but the error says 5 GB free. Diagnose it.
  7. Your server OOMs after 6 hours of operation but not at startup. What’s happening?

12. Further reading#

  • [REFERENCE] PyTorch CUDA memory management documentation; PYTORCH_CUDA_ALLOC_CONF
  • [REFERENCE] PyTorch memory visualizer (pytorch.org/memory_viz)
  • [ESTABLISHED] Kwon et al., PagedAttention (SOSP 2023) §3 — the fragmentation analysis
  • Next: 06 — OOM: causes and cures

↑↓ navigate↵ openesc close