PidokuInfra
11Production Inference Engineering
On this page

Production Inference Engineering

AdvancedModule 1111 topics~13h 30m

Topics, in order

01 SLOs, SLIs, and Latency BudgetsFor LLM services, the SLI definitions require more care than for ordinary services, because "latency" is not one number. Intermediate 1h 15m 02 Capacity PlanningHow many GPUs do I need? And its harder cousin: how many will I need in six months, and when do I have to order them? Advanced 1h 30m 03 Cost Per TokenBecause it's the number that connects every engineering decision to the business, and because it is the only fair way to compare configurations, models, and providers. Advanced 1h 15m 04 Autoscaling in ProductionSection VIII.07 covered the mechanics. This file covers the operational policy: what to actually configure, and what autoscaling can and cannot do for you. Advanced 1h 05 Fault Tolerance and Failure ModesThe two that are LLM-specific and most damaging: NCCL hangs and quality regressions. Both are hard to detect and both hold resources while broken. Advanced 1h 15m 06 Multi-Region and Disaster RecoveryThree distinct reasons, with different implications: Advanced 1h 07 Model RolloutsSection VIII.10 covered the mechanics of versioning and canary. This file covers the operational process: what a safe model release actually looks like end to end. Intermediate 1h 08 Observability for LLM ServicesSection X.10 covered the metrics catalogue and dashboards. This file covers what's specific to LLM services and what a generic APM setup will miss. Intermediate 1h 09 Security and IsolationAn inference service has an unusual attack surface: the input is arbitrary text that the system is designed to act on, and the compute is expensive and shared. Advanced 1h 15m 10 Abuse Prevention, Rate Limiting, and QuotasEvery LLM rate limiting scheme must be denominated in the resource that's actually scarce, which is a combination of tokens and memory-time. Advanced 1h 11 Scenario: 70B Model, 10,000 Concurrent Users, Fixed BudgetThe capstone. This is the system design interview question for inference engineering roles, and it is a real design exercise. Work through it yourself before reading the solution. Expert 2h

About this module

Goal: run an inference service that meets commitments, costs what you planned, and doesn’t wake you at 3 a.m.

Sections I-X taught you how inference works and how to make it fast. This section is about running it: SLOs, capacity, cost, reliability, observability, and security. It is less about GPUs and more about judgment.

Files#

#FileLevelTime
01SLOs, SLIs, and latency budgets ★Intermediate75 min
02Capacity planning ★Advanced90 min
03Cost per token ★Advanced75 min
04Autoscaling in productionAdvanced60 min
05Fault tolerance and failure modesAdvanced75 min
06Multi-region and disaster recoveryAdvanced60 min
07Model rolloutsIntermediate60 min
08Observability for LLM servicesIntermediate60 min
09Security and isolationAdvanced75 min
10Abuse prevention, rate limiting, quotasAdvanced60 min
11Scenario: 70B, 10,000 users, fixed budget ★Expert120 min

The thread#

flowchart TD
  N0["Promise something specific<br/><b>01</b>"]
  N1["Size the fleet to keep the promise<br/><b>02</b>"]
  N2["Know what it costs and why<br/><b>03</b>"]
  N3["Adjust as load changes — slowly<br/><b>04</b>"]
  N4["Survive failures (05) and regions going away<br/><b>06</b>"]
  N5["Change the model without breaking anything<br/><b>07</b>"]
  N6["See what's happening<br/><b>08</b>"]
  N7["Keep tenants isolated (09) and abusers out<br/><b>10</b>"]
  N8["Then do all of it at once, under constraints<br/><b>11</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8 memory

Checkpoint F (part 1)#

  1. Write an SLO for an LLM chat service. Why those specific metrics and thresholds?
  2. Size a fleet for 10,000 concurrent users on a 70B model. Show every step.
  3. Derive cost per million tokens and name the five terms that dominate it.
  4. Your service is over capacity and cannot scale in time. What happens?
  5. How do you attribute cost to individual tenants in a shared cluster?

↑↓ navigate↵ openesc close