PidokuInfra

Interconnects and Topology

Advanced 55 min Difficulty 4/5 Topic 01 of 04

Prerequisites II.05, IV.01

The idea in one minute#

When work spans several GPUs, they must exchange data, and the wires between them become part of the computation. There are three tiers, each roughly ten times slower than the last: NVLink between GPUs in one server, PCIe to the host, and the network (InfiniBand or Ethernet) between servers.

Which GPUs share which wires — the topology — decides whether a multi-GPU job is fast or quietly throttled.

A picture#

flowchart TB
  subgraph N1["Server 1"]
    direction LR
    G1["GPU 0"] <-->|"NVLink<br/>hundreds of GB/s"| G2["GPU 1"]
    G1 <-->|"PCIe"| C1["CPU + RAM"]
    G2 <-->|"PCIe"| C1
    NIC1["Network card<br/>RDMA"] <--> C1
  end
  subgraph N2["Server 2"]
    direction LR
    G3["GPU 0"] <-->|"NVLink"| G4["GPU 1"]
    NIC2["Network card<br/>RDMA"]
  end
  NIC1 <-->|"InfiniBand or Ethernet<br/>25 to 100 GB/s"| SW["Switch"] <--> NIC2
  class G1,G2,G3,G4 compute
  class C1 neutral
  class NIC1,NIC2,SW io

How it really works#

The tiers#

LinkConnectsRough speedNotes
HBMA GPU to its own memory3,000–8,000 GB/sFor scale
NVLinkGPU ↔ GPU in a server (or rack)300–900 GB/s per GPU (H100: 900)Direct, no host involved
PCIe 5.0 x16GPU ↔ host~64 GB/sShared with everything else on the bus
InfiniBand / RoCE EthernetServer ↔ server200–800 Gbit/s = 25–100 GB/s per portUses RDMA

Even NVLink is several times slower than a GPU’s own memory. Any design that moves data between GPUs every step is paying for it.

NVLink is NVIDIA’s direct GPU-to-GPU connection. In an eight-GPU server, NVSwitch chips connect every GPU to every other at full speed, so all pairs are equally close. Recent rack-scale systems extend this to dozens of GPUs acting as one NVLink domain.

Consumer cards and many cloud instances have no NVLink. Their GPUs talk through PCIe via the host, an order of magnitude slower. A multi-GPU technique that works beautifully on a data-center server can be useless on four gaming cards for this reason alone.

RDMA#

Normal networking copies data through the operating system on both ends. RDMA (remote direct memory access) lets a network card read and write application memory directly, and with GPUDirect RDMA, GPU memory directly — the bytes go GPU → network card → wire → network card → GPU without touching either CPU. InfiniBand provides this natively; RoCE provides it over Ethernet.

Topology traps inside a server#

  • NUMA. A two-socket server has two CPUs, each with its own RAM and its own PCIe slots. A process on socket 0 feeding a GPU attached to socket 1 pushes all its data across the inter-socket link. Pin the process to the CPU that owns the GPU.
  • PCIe switches. Several GPUs may hang off one switch and share its uplink to the host. Loading models into all of them at once is slower than the per-GPU figure suggests.
  • Degraded links. A badly seated card can negotiate x8 or an older generation. It works, slowly, forever. Check the negotiated link (IV.05).

nvidia-smi topo -m prints the matrix of how each GPU pair is connected.

The cost of a transfer#

time = latency + bytes ÷ bandwidth

For large messages bandwidth dominates. For many small messages the fixed latency does — microseconds on NVLink, tens of microseconds across a network. Multi-GPU inference sends small messages very often, which is why the tier you are on matters so much.

Code#

Go
// links.go — what does it cost to move data between GPUs on each tier?
package main

import "fmt"

type Link struct {
	Name      string
	GBps      float64 // bandwidth, GB/s
	LatencyUs float64 // fixed cost per message, microseconds
}

func (l Link) Seconds(bytes float64) float64 { return l.LatencyUs*1e-6 + bytes/(l.GBps*1e9) }

func main() {
	links := []Link{
		{"NVLink (H100)", 450, 2}, // realistic share of the 900 GB/s total
		{"PCIe via host", 25, 10},
		{"100 GbE RDMA", 11, 20},
		{"10 GbE TCP", 1.1, 100},
	}
	// A 70B model split across 2 GPUs exchanges activations twice per layer, per token:
	const layers, messagesPerLayer = 80, 2
	const hidden, batch, bytesPerVal = 8192, 32, 2
	msg := float64(hidden * batch * bytesPerVal) // one activation tensor: 512 KiB

	fmt.Printf("message %.0f KiB, %d messages per token\n\n", msg/1024, layers*messagesPerLayer)
	for _, l := range links {
		perToken := float64(layers*messagesPerLayer) * l.Seconds(msg)
		fmt.Printf("%-16s %7.2f ms of communication per token\n", l.Name, perToken*1e3)
	}
	fmt.Println("\nCompute per token for this model is on the order of 20-40 ms.")
}

On NVLink the communication is noise. Through the host it is a tax of 15–25%. Over ordinary networking it is several times the compute. Same model, same GPUs, same code.

Remember this#

  • Three tiers: NVLink (in-server), PCIe (to host), network (between servers) — each ~10x slower.
  • Without NVLink, GPU-to-GPU traffic goes through PCIe and the host.
  • RDMA moves data between machines without CPU copies.
  • Check topology: NUMA placement, shared PCIe switches, degraded links.

Try it#

  1. Run links.go. Reduce the batch to 1. Which term of Seconds dominates now?
  2. (GPU server) Run nvidia-smi topo -m and numactl --hardware. Draw your machine.
  3. Why might four GPUs without NVLink serve a split model more slowly than one larger GPU?

Check yourself#

  1. Order HBM, NVLink, PCIe and 100 GbE by bandwidth.
  2. What does RDMA avoid?
  3. Give two topology problems that slow a multi-GPU server without any error message.

↑↓ navigate↵ openesc close