The idea in one minute#
When work spans several GPUs, they must exchange data, and the wires between them become part of the computation. There are three tiers, each roughly ten times slower than the last: NVLink between GPUs in one server, PCIe to the host, and the network (InfiniBand or Ethernet) between servers.
Which GPUs share which wires — the topology — decides whether a multi-GPU job is fast or quietly throttled.
A picture#
flowchart TB
subgraph N1["Server 1"]
direction LR
G1["GPU 0"] <-->|"NVLink<br/>hundreds of GB/s"| G2["GPU 1"]
G1 <-->|"PCIe"| C1["CPU + RAM"]
G2 <-->|"PCIe"| C1
NIC1["Network card<br/>RDMA"] <--> C1
end
subgraph N2["Server 2"]
direction LR
G3["GPU 0"] <-->|"NVLink"| G4["GPU 1"]
NIC2["Network card<br/>RDMA"]
end
NIC1 <-->|"InfiniBand or Ethernet<br/>25 to 100 GB/s"| SW["Switch"] <--> NIC2
class G1,G2,G3,G4 compute
class C1 neutral
class NIC1,NIC2,SW ioHow it really works#
The tiers#
| Link | Connects | Rough speed | Notes |
|---|---|---|---|
| HBM | A GPU to its own memory | 3,000–8,000 GB/s | For scale |
| NVLink | GPU ↔ GPU in a server (or rack) | 300–900 GB/s per GPU (H100: 900) | Direct, no host involved |
| PCIe 5.0 x16 | GPU ↔ host | ~64 GB/s | Shared with everything else on the bus |
| InfiniBand / RoCE Ethernet | Server ↔ server | 200–800 Gbit/s = 25–100 GB/s per port | Uses RDMA |
Even NVLink is several times slower than a GPU’s own memory. Any design that moves data between GPUs every step is paying for it.
NVLink and NVSwitch#
NVLink is NVIDIA’s direct GPU-to-GPU connection. In an eight-GPU server, NVSwitch chips connect every GPU to every other at full speed, so all pairs are equally close. Recent rack-scale systems extend this to dozens of GPUs acting as one NVLink domain.
Consumer cards and many cloud instances have no NVLink. Their GPUs talk through PCIe via the host, an order of magnitude slower. A multi-GPU technique that works beautifully on a data-center server can be useless on four gaming cards for this reason alone.
RDMA#
Normal networking copies data through the operating system on both ends. RDMA (remote direct memory access) lets a network card read and write application memory directly, and with GPUDirect RDMA, GPU memory directly — the bytes go GPU → network card → wire → network card → GPU without touching either CPU. InfiniBand provides this natively; RoCE provides it over Ethernet.
Topology traps inside a server#
- NUMA. A two-socket server has two CPUs, each with its own RAM and its own PCIe slots. A process on socket 0 feeding a GPU attached to socket 1 pushes all its data across the inter-socket link. Pin the process to the CPU that owns the GPU.
- PCIe switches. Several GPUs may hang off one switch and share its uplink to the host. Loading models into all of them at once is slower than the per-GPU figure suggests.
- Degraded links. A badly seated card can negotiate x8 or an older generation. It works, slowly, forever. Check the negotiated link (IV.05).
nvidia-smi topo -m prints the matrix of how each GPU pair is connected.
The cost of a transfer#
time = latency + bytes ÷ bandwidthFor large messages bandwidth dominates. For many small messages the fixed latency does — microseconds on NVLink, tens of microseconds across a network. Multi-GPU inference sends small messages very often, which is why the tier you are on matters so much.
Code#
// links.go — what does it cost to move data between GPUs on each tier?
package main
import "fmt"
type Link struct {
Name string
GBps float64 // bandwidth, GB/s
LatencyUs float64 // fixed cost per message, microseconds
}
func (l Link) Seconds(bytes float64) float64 { return l.LatencyUs*1e-6 + bytes/(l.GBps*1e9) }
func main() {
links := []Link{
{"NVLink (H100)", 450, 2}, // realistic share of the 900 GB/s total
{"PCIe via host", 25, 10},
{"100 GbE RDMA", 11, 20},
{"10 GbE TCP", 1.1, 100},
}
// A 70B model split across 2 GPUs exchanges activations twice per layer, per token:
const layers, messagesPerLayer = 80, 2
const hidden, batch, bytesPerVal = 8192, 32, 2
msg := float64(hidden * batch * bytesPerVal) // one activation tensor: 512 KiB
fmt.Printf("message %.0f KiB, %d messages per token\n\n", msg/1024, layers*messagesPerLayer)
for _, l := range links {
perToken := float64(layers*messagesPerLayer) * l.Seconds(msg)
fmt.Printf("%-16s %7.2f ms of communication per token\n", l.Name, perToken*1e3)
}
fmt.Println("\nCompute per token for this model is on the order of 20-40 ms.")
}On NVLink the communication is noise. Through the host it is a tax of 15–25%. Over ordinary networking it is several times the compute. Same model, same GPUs, same code.
Remember this#
- Three tiers: NVLink (in-server), PCIe (to host), network (between servers) — each ~10x slower.
- Without NVLink, GPU-to-GPU traffic goes through PCIe and the host.
- RDMA moves data between machines without CPU copies.
- Check topology: NUMA placement, shared PCIe switches, degraded links.
Try it#
- Run
links.go. Reduce the batch to 1. Which term ofSecondsdominates now? - (GPU server) Run
nvidia-smi topo -mandnumactl --hardware. Draw your machine. - Why might four GPUs without NVLink serve a split model more slowly than one larger GPU?
Check yourself#
- Order HBM, NVLink, PCIe and 100 GbE by bandwidth.
- What does RDMA avoid?
- Give two topology problems that slow a multi-GPU server without any error message.