Goal: understand how models become services — and, more importantly, understand why real inference servers are built the way they are.
This section does not merely describe vLLM, SGLang, TGI, and Triton. It explains the design decisions each made, what they optimized for, and what they gave up. By the end you should be able to read an inference server’s source and predict its behavior under load.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | Anatomy of an inference server | Intermediate | 75 min |
| 02 | APIs: HTTP, gRPC, streaming | Intermediate | 60 min |
| 03 | Queues, scheduling, admission control ★ | Advanced | 90 min |
| 04 | Batching in servers | Advanced | 60 min |
| 05 | Backpressure, timeouts, retries | Advanced | 75 min |
| 06 | Load balancing and routing | Advanced | 75 min |
| 07 | Autoscaling and cold starts | Advanced | 75 min |
| 08 | Model lifecycle: loading, warmup, unloading | Intermediate | 60 min |
| 09 | Multi-model serving | Advanced | 75 min |
| 10 | Versioning, canary, A/B | Intermediate | 60 min |
| 11 | vLLM architecture ★ | Advanced | 90 min |
| 12 | SGLang architecture | Advanced | 60 min |
| 13 | TGI, Triton, ONNX Runtime | Advanced | 75 min |
| 14 | Choosing a serving stack | Intermediate | 60 min |
The thread#
flowchart TD N0["A model becomes a service (01) with an API<br/><b>02</b>"] N1["Requests arrive faster than they can be served, so you queue<br/><b>03</b>"] N2["And decide who to admit, and when to say no<br/><b>03, 05</b>"] N3["And group them for efficiency<br/><b>04</b>"] N4["Across many replicas, chosen carefully<br/><b>06</b>"] N5["Whose number changes with load — slowly<br/><b>07</b>"] N6["Each of which must load models (08), possibly several<br/><b>09</b>"] N7["In versions you can roll forward and back safely<br/><b>10</b>"] N8["All of which real systems implement in specific ways<br/><b>11-13</b>"] N9["And you have to pick one<br/><b>14</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 class N0,N1 neutral class N2,N3 io class N4,N5 queue class N6,N7 compute class N8,N9 memory