PidokuInfra
09Distributed Inference
On this page

Distributed Inference

AdvancedModule 0912 topics~14h 45m

Topics, in order

01 Why Distribute (and When Not To)Using more than one GPU for a single model. There are exactly three reasons to do it, and one very common bad reason. Intermediate 1h 02 Data ParallelismRun N complete copies of the model on N GPUs (or N groups of GPUs). Each request goes to one replica. There is no communication between replicas. Intermediate 45 min 03 Tensor ParallelismThe dominant form of parallelism in LLM inference. Understand the layout and you understand why it needs NVLink. Advanced 2h 04 Pipeline ParallelismSplit the model by layers across GPUs. GPU 0 holds layers 0-19, GPU 1 holds 20-39, and so on. Activations flow from one stage to the next. Advanced 1h 15m 05 Expert ParallelismMixture of Experts as a model is covered in Section XIII.02. This file is about the distributed systems problem it creates. Advanced 1h 15m 06 Sequence and Context ParallelismSplitting the sequence dimension across GPUs, rather than the weight or layer dimension. Advanced 1h 07 Collectives and NCCLCollectives are communication operations involving all ranks in a group. NCCL (NVIDIA Collective Communications Library) is the implementation everything uses. Advanced 1h 15m 08 Interconnects: NVLink, PCIe, InfiniBand, RDMAThe physical links between GPUs, and between nodes. These numbers determine which parallelism strategies are viable. Advanced 1h 15m 09 Communication Cost MathPure arithmetic, like Section V.06. This is how you decide a parallelism strategy on paper before spending money on hardware. Advanced 1h 30m 10 Multi-Node InferenceRunning one model instance across more than one physical machine. It is a step change in operational complexity, and the first question should always be whether you can avoid it. Advanced 1h 15m 11 Distributed KV CacheWhere the KV cache lives when a model instance spans multiple GPUs — and, increasingly, when you deliberately move KV between GPUs, nodes, or storage tiers. Advanced 1h 12 When More GPUs Make Things WorseThe counterintuitive file. Read it before your next capacity request. Advanced 1h 15m

About this module

Goal: understand what happens when one GPU isn’t enough — and, equally important, when adding GPUs makes things worse.

The central lesson of this section: parallelism is not free. Every form of it trades communication for computation, and whether that trade is good depends on your interconnect, your model, and your batch size. The arithmetic in file 09 is what separates people who can size a distributed deployment from people who guess.

Files#

#FileLevelTime
01Why distribute (and when not to)Intermediate60 min
02Data parallelismIntermediate45 min
03Tensor parallelism ★Advanced120 min
04Pipeline parallelismAdvanced75 min
05Expert parallelismAdvanced75 min
06Sequence and context parallelismAdvanced60 min
07Collectives and NCCLAdvanced75 min
08Interconnects: NVLink, PCIe, InfiniBand, RDMAAdvanced75 min
09Communication cost math ★Advanced90 min
10Multi-node inferenceAdvanced75 min
11Distributed KV cacheAdvanced60 min
12When more GPUs make things worse ★Advanced75 min

The thread#

flowchart TD
  N0["The model doesn't fit, or is too slow on one GPU<br/><b>01</b>"]
  N1["Replicate it (02) — helps throughput, not latency"]
  N2["Split each layer's weights (03) — helps latency, costs an AllReduce per layer"]
  N3["Or split by layers (04) — cheap communication, but bubbles"]
  N4["Or split by experts (05) — for MoE, with All-to-All"]
  N5["Or split by sequence (06) — for very long context"]
  N6["All of which run on collectives (07) over an interconnect<br/><b>08</b>"]
  N7["Whose cost you can compute<br/><b>09</b>"]
  N8["And which gets much worse across nodes<br/><b>10</b>"]
  N9["Where the KV cache also has to live somewhere<br/><b>11</b>"]
  N10["And where, past a point, more GPUs make everything slower<br/><b>12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10

  class N0,N1,N2 neutral
  class N3,N4 io
  class N5,N6 queue
  class N7,N8 compute
  class N9,N10 memory

↑↓ navigate↵ openesc close