1. What is it?#
Running one model instance across more than one physical machine. It is a step change in operational complexity, and the first question should always be whether you can avoid it.
2. Can you avoid it?#
Before going multi-node, exhaust these:
1. QUANTIZE. FP8 halves the model. INT4 quarters it.
A 405B model at FP16 needs 810 GB (2 nodes). At FP8, 405 GB (1 node with
8×80GB... tight). At INT4, 203 GB (comfortably 1 node).
2. USE A BIGGER-MEMORY GPU. H200 has 141 GB vs H100's 80 GB.
8×H200 = 1,128 GB — a 405B FP16 model fits on one node.
3. USE A SMALLER MODEL. Is the 405B actually better for your task than
the 70B? Measure it.
4. REDUCE max_model_len. Less KV per sequence, more room for weights.
Only if all four fail: go multi-node.Most teams that go multi-node didn’t need to. The operational cost — network configuration, failure handling, deployment coordination, debugging across machines — is substantial and permanent.
3. Simple analogy#
A surgical team versus two surgical teams in different hospitals operating on the same patient.
One team in one room: coordination is trivial.
Two teams in two buildings, coordinating over a phone line, on the same patient, in real time: possible, occasionally necessary, and everything that can go wrong now has twice the surface area.
4. The canonical layout#
2 nodes × 8 GPUs, model too large for one node:
TP=8 WITHIN each node (over NVLink — fast collectives)
PP=2 ACROSS nodes (over InfiniBand — small activation sends)
Node 0: layers 0-62, GPUs 0-7, TP group {0..7}
Node 1: layers 63-125, GPUs 8-15, TP group {8..15}
Per token:
within node: 2 AllReduces × 63 layers over NVLink
across node: 1 activation send (batch × d × bytes)Why not TP=16? Section IX.09 example 3: the cross-node AllReduce makes it unusable.
Why not PP=16? Bubbles. (16-1)/(M+15) requires M ≫ 16 microbatches.
5. What breaks when you go multi-node#
Networking#
REQUIRED:
✓ RDMA (InfiniBand or properly-configured RoCE)
✓ GPUDirect RDMA enabled and verified
✓ Consistent interface naming across nodes
✓ NCCL configured for the right HCAs
✓ MTU consistency (9000 for RoCE typically)
VERIFY BEFORE DEPLOYING:
ib_write_bw between every node pair
nccl-tests all_reduce_perf across nodes
NCCL_DEBUG=INFO showing GDRDMAA single misconfigured node in a cluster makes the whole job slow, and the symptom is “everything is slower” rather than “node 3 is broken.” Test pairwise.
Launch and coordination#
The ranks must start together, find each other, and initialize NCCL.
Options:
torchrun --nnodes=2 --nproc_per_node=8 --rdzv_backend=c10d \
--rdzv_endpoint=head:29500 serve.py
Ray (what vLLM uses for multi-node):
ray start --head on node 0
ray start --address=head:6379 on node 1
then vllm serve --tensor-parallel-size 8 --pipeline-parallel-size 2
Kubernetes: LeaderWorkerSet, or a JobSet, or an MPI operatorStartup coordination is the most common source of multi-node failures. A rank that doesn’t start, or starts late, or can’t reach the rendezvous, hangs the whole job.
Failure handling#
SINGLE NODE: a GPU fails → the process crashes → restart the process
MULTI-NODE: a GPU fails → that rank crashes → all other ranks HANG
(NCCL has no fault tolerance)
→ need a timeout, then kill and restart ALL ranks
→ which means reloading the model on every node: minutesSet NCCL_TIMEOUT and supervise the whole group. Without a timeout, a failure becomes a hang
that holds 16 GPUs indefinitely.
Practical supervision:
1. NCCL_TIMEOUT=600 (or lower)
2. TORCH_NCCL_ASYNC_ERROR_HANDLING=1 (fail fast on error)
3. A supervisor (Kubernetes Job/LeaderWorkerSet, or Ray) that restarts
ALL ranks when ANY fails
4. Readiness gating so traffic doesn't arrive during the restart
5. Alerting on restart frequencyDeployment#
Rolling a multi-node deployment means:
- draining in-flight requests (minutes)
- stopping all ranks
- starting all ranks with the new version
- loading the model on every node (minutes)
- warming up
→ 10-20 minutes of unavailability for that instance
Mitigation: run at least 2 multi-node instances, roll them one at a time.
Which doubles your minimum footprint.This is why multi-node has a high minimum cost. You can’t run one instance of a multi-node model in production; you need at least two for availability.
6. Under the hood — the Kubernetes shape#
# LeaderWorkerSet (the emerging standard for multi-node inference)
apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
name: llama-405b
spec:
replicas: 2 # two multi-node instances
leaderWorkerTemplate:
size: 2 # 2 pods per instance (2 nodes)
restartPolicy: RecreateGroupOnPodRestart # ← critical: all-or-nothing
leaderTemplate:
spec:
containers:
- name: vllm-leader
resources:
limits:
nvidia.com/gpu: 8
rdma/hca: 1 # RDMA device plugin
workerTemplate:
spec:
containers:
- name: vllm-worker
resources:
limits:
nvidia.com/gpu: 8
rdma/hca: 1RecreateGroupOnPodRestart encodes the all-or-nothing property: if one pod dies, recreate the
whole group. Without it, Kubernetes restarts one pod and the others hang.
Gang scheduling matters too: all pods of an instance must be schedulable simultaneously, or you get partial deployments that occupy GPUs and never become ready. Use a scheduler that supports it (Volcano, Kueue, or Kubernetes’ native pod scheduling gates).
7. Performance#
Llama-3-405B FP16, 2 nodes × 8 H100, InfiniBand NDR:
Configuration ITL (b=32) Throughput Notes
TP=8 × PP=2 34 ms 940 tok/s the canonical layout
TP=16 — — unusable (comms dominate)
1 node, FP8, TP=8 28 ms 1,140 tok/s ← BETTER, and single-node!
The FP8 single-node option is faster AND operationally simpler.That last row is the point of this file. Before accepting multi-node complexity, check whether quantization makes it unnecessary. In this example it does, and it wins on both performance and operations.
8. Production implications#
- Avoid multi-node if you can. Quantize, use bigger-memory GPUs, or use a smaller model.
- If you must: TP within nodes, PP across. Never TP across nodes.
- Verify RDMA and GPUDirect before deploying, pairwise between all nodes.
- Set NCCL timeouts and supervise all-or-nothing.
- Run at least 2 instances for availability during rolls.
- Use gang scheduling. Partial placements waste GPUs.
- Expect longer cold starts — every node loads its shard, and they must all finish.
- Test the failure path: kill one rank and verify the group restarts cleanly within your timeout.
- Monitor per-node health. One degraded node slows the whole instance, and the symptom is global.
9. Common mistakes#
Going multi-node without trying quantization.
TP across nodes. Unusable overhead.
No NCCL timeout. A failure becomes an indefinite hang.
Restarting one pod instead of the group. Remaining ranks hang.
No gang scheduling. Partial placements hold GPUs and never serve.
Not verifying GPUDirect RDMA. Silent 2x slowdown.
Running one instance. No availability during deployments.
Assuming the network is fine because one pair works. Test all pairs.
10. Hands-on exercise#
A. The avoidance check. For a model you think needs multi-node, compute whether FP8 or INT4 lets it fit on one node. Compare the performance and operational complexity of both options.
B. Verify the fabric. If you have multi-node access: run ib_write_bw between every node
pair, then all_reduce_perf across nodes. Is any pair slower? Confirm GDRDMA with
NCCL_DEBUG=INFO.
C. Measure both layouts. Run TP=8×PP=2 and (if it completes) TP=16. Measure ITL and throughput. Confirm the prediction from Section IX.09.
D. Test the failure path. Kill one rank of a multi-node job. Measure: how long until the group notices, how long until it restarts, how long until it serves again. Is your timeout set appropriately?
E. Cold start. Measure multi-node cold start end to end. Which node is slowest? Is the model load parallel across nodes?
11. Interview questions#
- What should you try before going multi-node?
- What’s the canonical multi-node layout and why?
- What happens when one rank of a multi-node job fails?
- Why do you need at least two multi-node instances in production?
- What is gang scheduling and why does multi-node inference need it?
- How would you verify a multi-node fabric before deploying?
- Why might a single-node FP8 deployment beat a two-node FP16 one?
12. Further reading#
- [REFERENCE] vLLM distributed serving documentation (multi-node with Ray)
- [REFERENCE] Kubernetes LeaderWorkerSet; Kueue and Volcano for gang scheduling
- [REFERENCE] NVIDIA multi-node deployment guides
- Next: 11 — Distributed KV cache