PidokuInfra
08Inference Serving Systems
On this page

Inference Serving Systems

IntermediateModule 0814 topics~16h 30m

Topics, in order

01 Anatomy of an Inference ServerThe components every LLM inference server has, and how they fit together. Intermediate 1h 15m 02 APIs: HTTP, gRPC, StreamingThe contract between clients and your inference service: protocol, schema, streaming semantics, errors, and limits. Intermediate 1h 03 Queues, Scheduling, and Admission ControlMore production incidents are caused by bad queueing than by bad kernels. This is where tail latency lives. Advanced 1h 30m 04 Batching in ServersSection V covered what continuous batching is. This file covers the server-side knobs: what they control, how to tune them, and what happens when you get them wrong. Advanced 1h 05 Backpressure, Timeouts, and RetriesThe three mechanisms that determine whether your service degrades gracefully or collapses. Advanced 1h 15m 06 Load Balancing and RoutingAnd all four destroy prefix caching, because consecutive turns of the same conversation land on different replicas with cold caches. Advanced 1h 15m 07 Autoscaling and Cold StartsBoth halves of autoscaling are broken: the signal and the reaction time. Fixing the signal is easy; fixing the reaction time is a capacity-planning problem, not an engineering one. Advanced 1h 15m 08 Model Lifecycle: Loading, Warmup, UnloadingEverything that happens to a model between "it exists in storage" and "it is serving traffic correctly," and back again. Intermediate 1h 09 Multi-Model ServingServing more than one model from shared infrastructure. Three distinct problems, often confused: Advanced 1h 15m 10 Versioning, Canary Deployments, and A/B TestingHow you change what's serving without breaking anything, and how you find out whether the change was good. Intermediate 1h 11 vLLM ArchitectureRead this file with the source open. vLLM is the reference implementation of everything in Section V, and reading it is the fastest way to make those concepts concrete. Advanced 1h 30m 12 SGLang ArchitectureAn LLM serving system whose distinguishing contributions are RadixAttention (a radix-tree prefix cache) and a frontend language for structured LLM programs. Advanced 1h 13 TGI, Triton, and ONNX RuntimeThree more serving systems, each designed for a different problem. Understanding what each optimized for is more useful than a feature checklist. Advanced 1h 15m 14 Choosing a Serving StackYou need to choose, and the honest answer is that for most teams, most of the time, the decision matters less than how well you tune whatever you choose. A well-tuned vLLM beats a … Intermediate 1h

About this module

Goal: understand how models become services — and, more importantly, understand why real inference servers are built the way they are.

This section does not merely describe vLLM, SGLang, TGI, and Triton. It explains the design decisions each made, what they optimized for, and what they gave up. By the end you should be able to read an inference server’s source and predict its behavior under load.

Files#

#FileLevelTime
01Anatomy of an inference serverIntermediate75 min
02APIs: HTTP, gRPC, streamingIntermediate60 min
03Queues, scheduling, admission control ★Advanced90 min
04Batching in serversAdvanced60 min
05Backpressure, timeouts, retriesAdvanced75 min
06Load balancing and routingAdvanced75 min
07Autoscaling and cold startsAdvanced75 min
08Model lifecycle: loading, warmup, unloadingIntermediate60 min
09Multi-model servingAdvanced75 min
10Versioning, canary, A/BIntermediate60 min
11vLLM architecture ★Advanced90 min
12SGLang architectureAdvanced60 min
13TGI, Triton, ONNX RuntimeAdvanced75 min
14Choosing a serving stackIntermediate60 min

The thread#

flowchart TD
  N0["A model becomes a service (01) with an API<br/><b>02</b>"]
  N1["Requests arrive faster than they can be served, so you queue<br/><b>03</b>"]
  N2["And decide who to admit, and when to say no<br/><b>03, 05</b>"]
  N3["And group them for efficiency<br/><b>04</b>"]
  N4["Across many replicas, chosen carefully<br/><b>06</b>"]
  N5["Whose number changes with load — slowly<br/><b>07</b>"]
  N6["Each of which must load models (08), possibly several<br/><b>09</b>"]
  N7["In versions you can roll forward and back safely<br/><b>10</b>"]
  N8["All of which real systems implement in specific ways<br/><b>11-13</b>"]
  N9["And you have to pick one<br/><b>14</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8,N9 memory

↑↓ navigate↵ openesc close