1. What is it?#
The GPU is not in the CPU’s memory. It is a peripheral at the end of a bus. Everything that moves between them crosses PCIe, and it is moved by a DMA engine, not by the CPU.
CPU ──── PCIe Gen4 x16 ──── GPU
~25 GB/s real vs HBM3 at 3,350 GB/s
→ 134x slowerThere are faster paths — NVLink between GPUs, NVLink-C2C between Grace CPU and Hopper GPU — and knowing which path your data takes is the difference between a design that works and one that doesn’t.
2. Why does it exist?#
Modularity. A GPU is a card you plug in. That flexibility costs you a bus that is two orders of magnitude slower than local memory. Every inference architecture decision involving data movement between host and device is downstream of that number.
3. Simple analogy#
A workshop with a supply depot down the road. Inside the workshop everything is at arm’s reach (HBM). The depot is a 20-minute drive (PCIe). You would never drive to the depot for each screw; you’d bring a whole crate once and keep it on the shelf.
That is exactly the rule: load weights once, keep them resident, and only ship small things (token IDs) across the bus.
4. Tiny example#
import torch, time
def h2d(size_mb, pinned):
n = size_mb * 1024 * 1024 // 4
src = torch.empty(n, dtype=torch.float32, pin_memory=pinned)
dst = torch.empty(n, dtype=torch.float32, device='cuda')
for _ in range(3): dst.copy_(src, non_blocking=pinned)
torch.cuda.synchronize()
t0 = time.perf_counter()
for _ in range(10): dst.copy_(src, non_blocking=pinned)
torch.cuda.synchronize()
dt = (time.perf_counter()-t0)/10
print(f"{size_mb:6d} MB pinned={pinned!s:5s} {dt*1e3:8.2f} ms {size_mb/1024/dt:7.2f} GB/s")
for mb in [1, 4, 16, 64, 256, 1024]:
h2d(mb, False)
h2d(mb, True)Typical results (PCIe Gen4 x16):
1 MB pinned=False 0.35 ms 2.79 GB/s
1 MB pinned=True 0.09 ms 10.85 GB/s
1024 MB pinned=False 92.10 ms 10.86 GB/s
1024 MB pinned=True 41.20 ms 24.27 GB/sTwo lessons:
- Pinned memory roughly doubles bandwidth (no staging-buffer copy).
- Small transfers are dominated by fixed overhead (~5-10 µs per transfer). Batch them.
Now the decisive comparison for inference design:
Reading 14 GB of weights from HBM: 14/3350 = 4.2 ms
Reading 14 GB of weights over PCIe: 14/24 = 583 ms139x. This is why weights live in HBM permanently, and why “just offload the model to CPU RAM and stream it” produces a system that generates ~1.7 tokens/sec.
5. Technical explanation#
PCIe generations#
| Gen | Per-lane | x16 bidirectional (theoretical) | x16 realistic one-way |
|---|---|---|---|
| 3.0 | 0.985 GB/s | 15.75 GB/s each way | ~12 GB/s |
| 4.0 | 1.97 GB/s | 31.5 GB/s each way | ~24 GB/s |
| 5.0 | 3.94 GB/s | 63 GB/s each way | ~50 GB/s |
| 6.0 | 7.56 GB/s | 121 GB/s each way | ~100 GB/s |
Realistic is ~75-80% of theoretical due to encoding and protocol overhead. Check your actual link:
nvidia-smi -q | grep -A4 "GPU Link Info"
lspci -vv -s <bus:dev.fn> | grep -E "LnkCap|LnkSta"Watch for downgraded links. A GPU physically in an x16 slot may negotiate x8 or x4 due to bifurcation or a riser. That halves or quarters your transfer bandwidth and is invisible unless you look.
DMA#
The CPU does not copy data to the GPU. It programs a DMA engine (a “copy engine” on the GPU) with source, destination, and length; the engine moves it while the CPU does other work.
Requirements and consequences:
- The source must be pinned (page-locked), or CUDA stages through an internal pinned buffer, costing an extra copy.
- Copies on a separate stream overlap with compute. An H100 has multiple copy engines, so H2D, D2H, and compute can all run concurrently.
cudaMemcpyAsyncfrom pageable memory is not actually async.
stream = torch.cuda.Stream()
with torch.cuda.stream(stream):
gpu_buf.copy_(pinned_buf, non_blocking=True) # overlaps with default-stream computeGPU-to-GPU paths#
Path Bandwidth (per GPU) Latency
NVLink 4 (H100, 18 links) 900 GB/s bidir ~2 µs
NVLink 3 (A100, 12 links) 600 GB/s bidir ~2 µs
PCIe Gen5 x16 peer-to-peer ~50 GB/s ~5-10 µs
PCIe via host bounce ~25 GB/s, 2 hops ~20 µs
InfiniBand NDR (400 Gb/s) 50 GB/s ~2-3 µs + switchThis table decides your parallelism strategy. Tensor parallelism does an AllReduce per transformer block — dozens per token. On NVLink that is a few hundred microseconds total; over PCIe it can exceed the compute time entirely (Section IX.09).
nvidia-smi topo -m # shows NV#, PIX, PXB, PHB, SYS between every GPU pairLegend: NV# = NVLink with # links; PIX = same PCIe switch; PXB = multiple switches;
PHB = via host bridge; SYS = across the CPU interconnect (worst).
Special paths worth knowing#
- GPUDirect P2P — GPU A reads GPU B’s memory directly over NVLink/PCIe without host bounce.
- GPUDirect RDMA — the NIC DMAs straight into GPU memory. Essential for multi-node.
- GPUDirect Storage — NVMe → GPU memory directly.
- NVLink-C2C (Grace Hopper) — 900 GB/s CPU↔GPU, coherent. This changes the economics of CPU offload dramatically: KV offload to host memory becomes genuinely practical (Section XIII.07).
6. Under the hood#
# Live PCIe utilization
nvidia-smi dmon -s t # rxpci, txpci columns in MB/s
# Peer-to-peer capability and measured bandwidth
# (from cuda-samples)
./p2pBandwidthLatencyTest
./bandwidthTest --memory=pinned --mode=range --start=1024 --end=1073741824 --increment=1048576In Nsight Systems, host-to-device copies appear on their own rows. If you see copies on the critical path between kernels every step, you have a design problem — find what is being moved and eliminate it.
7. Performance implications#
Where PCIe bites in inference:
| Pattern | Cost | Verdict |
|---|---|---|
| Weight load at startup | 140 GB / 24 GB/s = 6 s | fine, once |
| Token IDs per request | ~4 KB | negligible |
| Logits D2H every step (for CPU sampling) | B×V×4 bytes = 8 MB at B=16,V=128k | 1 ms/step — significant. Sample on GPU. |
| CPU-offloaded KV cache | GB per step | usually fatal over PCIe; viable over NVLink-C2C |
| Weight streaming (model > VRAM) | 140 GB/step | ~1.7 tok/s. Only for hobby use. |
| Multi-GPU AllReduce over PCIe | 10s of MB × 2 per layer | often dominates; use NVLink |
The logits one catches people. Any per-step device→host transfer plus synchronization adds directly to ITL. Modern engines keep sampling entirely on the GPU for exactly this reason.
8. Production implications#
- Verify link width and generation on every new node type. Downgraded links are common and silent.
- Check
nvidia-smi topo -mbefore choosing a TP degree. TP=8 across two PCIe islands with no NVLink is usually worse than TP=4 within one island plus data parallelism. - Use a pinned buffer pool for any recurring transfer.
- Keep sampling, stopping criteria, and logit processing on the GPU.
- For multi-node, require GPUDirect RDMA. Without it, every cross-node byte bounces through host memory, roughly halving effective bandwidth and adding latency.
- In Kubernetes, ensure the GPUs assigned to one pod share an NVLink domain (Topology Manager, Section XII.03).
9. Common mistakes#
Designing around CPU offload without checking PCIe bandwidth. Do the arithmetic first.
Non-pinned transfers in the hot path. Halves bandwidth and blocks.
Per-step D2H copies with synchronization. Adds milliseconds to ITL.
Assuming all 8 GPUs in a box are equally connected. They often are not.
Ignoring that transfers share bandwidth. Concurrent H2D and D2H on Gen4 each get their direction’s bandwidth, but multiple GPUs behind one PCIe switch share the uplink.
Forgetting that NCCL uses these same paths. Your “network” performance in Section IX is determined by this file.
10. Hands-on exercise#
A. Measure your bus. Run the pinned/pageable benchmark from section 4. Determine your
realistic PCIe bandwidth and the fixed per-transfer overhead (intercept of a linear fit). Record
in numbers.md.
B. Verify link status. Check nvidia-smi -q | grep -A4 "GPU Link Info" and lspci -vv.
Is every GPU at full width and generation?
C. Map the topology. Run nvidia-smi topo -m and draw the GPU interconnect graph. Which
GPUs would you group for TP=4?
D. Prove the offload argument. Compute the tokens/sec ceiling for a 70B model whose weights live in host RAM and stream over PCIe Gen4 x16 each step. Then, if you have the hardware, measure it with an offloading runtime and compare.
E. Overlap. Write code that transfers a large buffer on a side stream while a matmul runs on the default stream. Verify overlap in Nsight Systems (or by timing: total < sum of parts).
11. Interview questions#
- What is the bandwidth of PCIe Gen4 x16 and how does it compare to HBM3?
- Why is pinned memory faster for H2D transfers? What does it cost?
- What is DMA and why does it let transfers overlap with compute?
- Explain
nvidia-smi topo -moutput and how it affects tensor parallelism decisions. - Why is streaming model weights from host RAM impractical? Show the arithmetic.
- Why should sampling happen on the GPU?
- What is GPUDirect RDMA and why does multi-node inference need it?
12. Further reading#
- [REFERENCE] NVIDIA CUDA Best Practices Guide, “Data Transfer Between Host and Device”
- [REFERENCE] Mark Harris, “How to Optimize Data Transfers in CUDA C/C++”
- [REFERENCE] NVIDIA NVLink and NVSwitch technical briefs
- [REFERENCE] cuda-samples:
bandwidthTest,p2pBandwidthLatencyTest - Next: 10 — OS scheduling