PidokuInfra
12Inference Platform Engineering
On this page

Inference Platform Engineering

ExpertModule 1211 topics~13h

Topics, in order

01 Platform Overview: Control Plane vs Data PlaneThe critical property: the data plane must keep working when the control plane is unavailable. This is the difference between a platform that degrades and one that fails. Advanced 1h 02 Model RegistryThe authoritative record of what models exist, what versions of each, where their artifacts live, and how to serve them. Intermediate 1h 03 Kubernetes for GPUsMost inference platforms run on Kubernetes. Most Kubernetes GPU configurations are subtly wrong in ways that cost 20-40% of performance. Advanced 1h 30m 04 GPU Scheduling and Model PlacementThis is a bin-packing problem with expensive moves and a moving target. The expensive-moves property is what makes it different from ordinary Kubernetes scheduling. Advanced 1h 30m 05 Multi-Tenancy and QuotasSection XI.09 covered isolation (security) and XI.10 covered limits (enforcement). This file is about the platform-level design: how tenants share a cluster. Advanced 1h 15m 06 Routing StrategiesSection VIII.06 covered load balancing between replicas of one model. This file covers the platform's routing layer, which also decides which model and which pool. Advanced 1h 15m 07 Prefix-Aware RoutingBuilding prefix caching and then routing randomly is one of the most common self-inflicted wounds in LLM platform engineering. The fix is routing logic, not more caching. Advanced 1h 08 Inference GatewaysThe single front door for all inference traffic. Everything common to every request lives here, so it doesn't have to live in every engine. Advanced 1h 15m 09 Cascades and FallbacksThey use similar machinery and are often confused. Build both; know which you're building. Advanced 1h 10 Cluster and Capacity ManagementThe operational management of a GPU fleet: keeping nodes healthy, keeping capacity allocated sensibly, and knowing when to buy more. Advanced 1h 11 Building the Platform: A RoadmapThe order matters more than the components. This is the file to read before you start building. Advanced 1h 15m

About this module

Goal: build the system that lets many teams serve many models on shared hardware, safely and efficiently.

Section XI was about running a service well. This section is about running a platform — which is a different problem, with the central insight from Section IX.12: consolidating traffic is often worth more than any kernel optimization.

Files#

#FileLevelTime
01Platform overview: control plane vs data planeAdvanced60 min
02Model registryIntermediate60 min
03Kubernetes for GPUs ★Advanced90 min
04GPU scheduling and model placementAdvanced90 min
05Multi-tenancy and quotasAdvanced75 min
06Routing strategiesAdvanced75 min
07Prefix-aware routingAdvanced60 min
08Inference gatewaysAdvanced75 min
09Cascades and fallbacksAdvanced60 min
10Cluster and capacity managementAdvanced60 min
11Building the platform: a roadmap ★Advanced75 min

The thread#

flowchart TD
  N0["A platform separates what changes slowly from what happens per request<br/><b>01</b>"]
  N1["It needs to know what models exist<br/><b>02</b>"]
  N2["And where to run them, on scarce hardware<br/><b>03, 04</b>"]
  N3["Shared fairly between teams<br/><b>05</b>"]
  N4["With requests routed intelligently<br/><b>06, 07</b>"]
  N5["Through a front door that handles everything common<br/><b>08</b>"]
  N6["With fallbacks when the primary path fails or is too expensive<br/><b>09</b>"]
  N7["On a cluster whose capacity is actively managed<br/><b>10</b>"]
  N8["Built incrementally, in the right order<br/><b>11</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8 memory

Checkpoint F (part 2)#

  1. What belongs in the control plane and what in the data plane? Why does the distinction matter?
  2. Design the model placement algorithm for a cluster with 200 GPUs and 40 models.
  3. How do you attribute cost to tenants, and what’s the term everyone forgets?
  4. Why does prefix-aware routing matter more than load balance for some workloads?
  5. What’s the first thing you’d build, and what’s the last?

↑↓ navigate↵ openesc close