1. What is it?#
DRAM is main memory: cheap, large, and slow relative to caches. NUMA (Non-Uniform Memory Access) is what happens on multi-socket servers: each CPU socket has its own memory controller and its own attached DRAM, and reaching the other socket’s memory costs roughly 1.5-2x more latency and less bandwidth.
┌─────────────┐ UPI/Infinity Fabric ┌─────────────┐
│ Socket 0 │◄─────────────────────►│ Socket 1 │
│ 32 cores │ ~30-60 GB/s │ 32 cores │
└──────┬──────┘ └──────┬──────┘
│ ~300 GB/s, 80 ns │ ~300 GB/s, 80 ns
┌─────▼──────┐ ┌─────▼──────┐
│ 512 GB DRAM│ │ 512 GB DRAM│
│ (node 0) │ │ (node 1) │
└────────────┘ └────────────┘
Socket 0 reading node 1 memory: ~130 ns, lower bandwidthGPUs are attached to specific sockets too. Getting this wrong costs you real throughput on every host-to-device transfer.
2. Why does it exist?#
One memory controller cannot feed 128 cores. So each socket gets its own, and the sockets are connected. The result is a shared address space with non-uniform costs — convenient to program, easy to get wrong.
For inference, NUMA matters in three places:
- Host-to-device transfers (pinned buffers should be on the GPU’s local NUMA node).
- CPU-side work (tokenization pools, the engine loop) touching remote memory.
- CPU inference and CPU-offloaded KV cache, which are bandwidth-bound and therefore very sensitive to remote access.
3. Simple analogy#
A two-building office. Your team’s filing cabinet is in your building (30 seconds away). The other team’s is across the courtyard (90 seconds). Everything works either way, but if half your files end up in the wrong building, your day takes twice as long — and nothing in the org chart tells you it’s happening.
NUMA problems are exactly like that: invisible in code, visible only in a profiler or a throughput number that is mysteriously half of what you expected.
4. Tiny example#
# What does the machine look like?
numactl --hardware
lscpu | grep -i numa
# Which NUMA node is each GPU attached to?
nvidia-smi topo -m
# Look for the "NUMA Affinity" columnTypical nvidia-smi topo -m output on an 8-GPU node:
GPU0 GPU1 GPU2 GPU3 GPU4 GPU5 GPU6 GPU7 CPU Affinity NUMA
GPU0 X NV18 NV18 NV18 SYS SYS SYS SYS 0-31,64-95 0
GPU1 NV18 X NV18 NV18 SYS SYS SYS SYS 0-31,64-95 0
...
GPU4 SYS SYS SYS SYS X NV18 NV18 NV18 32-63,96-127 1Read it: GPUs 0-3 are on NUMA node 0, GPUs 4-7 on node 1. NV18 = 18 NVLink connections;
SYS = the traffic must cross the CPU interconnect. This single table tells you how to place
your processes and how tensor parallelism will perform (Section IX.08).
Now measure the penalty:
# Local: run on node 0, allocate on node 0
numactl --cpunodebind=0 --membind=0 ./bandwidth_test
# Remote: run on node 0, allocate on node 1
numactl --cpunodebind=0 --membind=1 ./bandwidth_testExpect 1.3-2x worse bandwidth and ~1.6x worse latency in the remote case.
5. Technical explanation#
DRAM basics that matter#
DRAM is organized into channels, ranks, banks, rows. A read activates a whole row (2-8 KB) into a row buffer; subsequent reads in the same row are fast (“row hit”), reads to a different row in the same bank require a precharge + activate (“row miss”, ~2-3x slower).
Practical consequences:
- Sequential access is fast; random access across banks is much slower than the raw latency number suggests.
- Bandwidth scales with channels. A CPU with 12 DDR5 channels populated has ~2x the bandwidth of one with 6. Under-populating DIMM slots is a common and expensive procurement mistake in inference hosts, especially for CPU offload workloads.
NUMA policies#
default / first-touch page is allocated on the node of the thread that FIRST WRITES it
bind allocate only on specified nodes (fail if full)
interleave round-robin pages across nodes (halves worst case, kills best case)
preferred try one node, fall backFirst-touch is the default and the source of most surprises. If a single initialization thread touches all your memory, everything lands on one node, and the other socket’s threads all run remote. Fix: parallelize initialization with the same thread affinity used later.
Pinned memory and GPU transfers#
cudaHostAlloc / torch.empty(pin_memory=True) allocates page-locked host memory, which DMA
engines can read directly. If that buffer is on the wrong NUMA node relative to the GPU’s PCIe
root complex, every transfer crosses the socket interconnect:
Correct: pinned buffer on node 0 → PCIe root on node 0 → GPU0 ~25 GB/s
Wrong: pinned buffer on node 1 → UPI → PCIe root node 0 → GPU0 ~12-18 GB/sFor weight loading (140 GB) that’s the difference between 6 s and 11 s. For per-step transfers in a CPU-offload design, it’s the difference between viable and not.
6. Under the hood#
Check where a running process’s memory actually is:
# Per-node memory usage of a process
numastat -p $(pgrep -f vllm)
# Page-level detail
cat /proc/<pid>/numa_maps | head
# Automatic NUMA balancing (kernel migrates pages; can cause latency spikes)
cat /proc/sys/kernel/numa_balancingnuma_balancing migrates pages toward the accessing thread. Good for long-running steady
workloads; a source of periodic latency spikes for latency-sensitive ones. Many HPC and
low-latency deployments disable it and pin explicitly instead.
7. Performance implications#
| Scenario | Penalty for getting NUMA wrong |
|---|---|
| Host→device weight load | 1.5-2x slower cold start |
| Pinned-buffer streaming (KV offload) | 1.5-2x less effective bandwidth |
| CPU inference (bandwidth-bound) | up to 2x slower |
| Tokenizer thread pool | 1.1-1.3x (small but free to fix) |
| NCCL over PCIe across sockets | significant; often the reason TP=8 underperforms |
8. Production implications#
- Pin one engine process per NUMA node when running multiple model replicas on a
multi-socket, multi-GPU host:Shell
numactl --cpunodebind=0 --membind=0 python -m vllm.entrypoints.openai.api_server --port 8000 ... numactl --cpunodebind=1 --membind=1 python -m vllm.entrypoints.openai.api_server --port 8001 ... - Keep a tensor-parallel group within one NUMA node / NVLink domain where possible. TP across sockets over PCIe is a common cause of disappointing multi-GPU scaling.
- Check
nvidia-smi topo -mon every new instance type. Cloud SKUs differ, and the topology determines your parallelism plan. - In Kubernetes, enable the Topology Manager with
single-numa-nodepolicy so pods get CPU, memory, and GPU from the same node (Section XII.03). - Populate all memory channels when specifying hosts.
9. Common mistakes#
Ignoring NUMA entirely on a 2-socket box. The most common one. Costs 20-50% on memory-bound host work, silently.
Initializing all memory from one thread. First-touch puts it all on one node.
Interleaving as a default “fix.” It bounds the worst case but prevents the best case. Bind properly instead.
Assuming the container sees the topology. Containers inherit the host’s NUMA layout but may be restricted to CPUs on one node while memory is allocated elsewhere.
Forgetting that GPUs have NUMA affinity too. They hang off a PCIe root complex attached to one socket.
10. Hands-on exercise#
A. Map your machine. Run numactl --hardware, lstopo --output-format txt,
nvidia-smi topo -m. Draw the topology on paper: sockets, memory, PCIe roots, GPUs, NVLink.
B. Measure the penalty. Write a simple memory-bandwidth benchmark (STREAM-like triad). Run
it with --membind=0 --cpunodebind=0 and --membind=1 --cpunodebind=0. Report the ratio.
C. Measure H2D transfer with NUMA. Allocate pinned host memory on each node (use
numactl --membind), transfer 1 GB to GPU 0, and compare bandwidth. Record in numbers.md.
D. Fix a real deployment. If you have a multi-socket GPU box, run a model server with and
without numactl binding and compare cold-start time and steady-state throughput.
11. Interview questions#
- What is NUMA and what is the typical remote-access penalty?
- What is first-touch allocation and how does it cause NUMA problems?
- How do you determine which NUMA node a GPU is attached to?
- Why does NUMA affect host-to-device transfer bandwidth?
- You are running 2 model replicas on a dual-socket 8-GPU box. How do you place them?
- When would you use
--interleaveand why is it usually not the right answer?
12. Further reading#
- [REFERENCE]
numactl(8),numastat(8),man 7 numa - [FUNDAMENTAL] Drepper, “What Every Programmer Should Know About Memory,” part 5
- [REFERENCE] Kubernetes Topology Manager documentation
- Next: 04 — SIMD and vectorization