1. The definitions#
SLI Service Level INDICATOR — a measurement. "p95 TTFT over 5 minutes."
SLO Service Level OBJECTIVE — a target for an SLI. "p95 TTFT < 800 ms."
SLA Service Level AGREEMENT — a contractual promise with consequences.For LLM services, the SLI definitions require more care than for ordinary services, because “latency” is not one number.
2. Why LLM SLOs are different#
ORDINARY SERVICE LLM SERVICE
one latency metric TTFT and ITL are separate experiences
requests are comparable cost varies 1000x by request shape
success is 2xx success includes "the answer was good"
latency is roughly stationary latency depends on prompt and output lengthA single “p99 latency < 2 s” SLO is meaningless for an LLM service. It will be violated by every long generation and satisfied by every short one, telling you nothing about health.
3. A well-formed LLM SLO#
SERVICE: chat-completions, model llama-3-70b, standard tier
AVAILABILITY
99.9% of requests return a non-5xx response
measured over a rolling 30 days
LATENCY — TIME TO FIRST TOKEN
p95 TTFT < 800 ms for requests with ≤ 2,000 prompt tokens
p95 TTFT < 3,000 ms for requests with 2,000-16,000 prompt tokens
p95 TTFT < 12,000 ms for requests with > 16,000 prompt tokens
measured over rolling 5-minute windows
LATENCY — INTER-TOKEN
p95 ITL < 50 ms
p99 ITL < 120 ms
measured over rolling 5-minute windows
THROUGHPUT (capacity commitment)
the service accepts ≥ 40 requests/second at the above latencies
GOODPUT (the composite)
≥ 97% of requests meet BOTH their TTFT and ITL objectives
QUALITY
the deployed model version's evaluation suite score is within 2% of
the reference, measured at each deployment
EXCLUSIONS
requests exceeding published limits (max input 32k, max output 4k)
requests rejected with 429 (over quota) — counted separatelyFour things that make this good:
- TTFT is bucketed by prompt length. Without this, one user pasting a book violates your SLO.
- ITL has both p95 and p99. The p99 catches blocking events (Section X, case 2).
- Goodput is the composite metric. It can’t be gamed by making everyone uniformly slow.
- Quality is an SLO. For an LLM service, “it responded” is not “it worked.”
4. Choosing the numbers#
Don’t invent thresholds. Derive them.
TTFT#
What does the user experience?
< 200 ms feels instant
200-500 ms feels responsive
0.5-1 s noticeable but acceptable for a "thinking" interaction
1-3 s requires a loading indicator; users tolerate it if warned
> 3 s users leave or retry
Then check what's ACHIEVABLE:
TTFT_floor = tokenize + prefill(p95_prompt_length) + network_rtt
For a 70B model, 2,000-token prompt: ~150 ms prefill + 50 ms network
→ 800 ms p95 leaves 600 ms of queueing headroom. Reasonable.
Set the SLO where the achievable meets the acceptable, with headroom.ITL#
Human reading speed: 200-300 words/min ≈ 4-5 words/sec ≈ 5-7 tokens/sec
→ 140-200 ms per token is "reading speed"
But: users don't read as it generates; they wait for a chunk then read.
And perceived smoothness matters.
< 30 ms feels like fast typing; smooth
30-60 ms smooth
60-100 ms slightly choppy but fine
> 150 ms visibly slow
→ p95 ITL < 50 ms is a good target for chat.
→ For agentic/batch workloads, ITL barely matters; E2E does.Availability#
99.9% = 43 minutes/month of downtime. Achievable with N+1 and good practices.
99.95% = 22 minutes/month. Needs multi-AZ and careful deployments.
99.99% = 4.3 minutes/month. Needs multi-region and a lot of engineering.
For a GPU service, be honest: GPU failures, driver issues, and long cold
starts make 99.99% expensive. Promise 99.9% and deliver it.5. The latency budget#
Decompose the SLO into a budget you can hold each component to:
p95 TTFT budget: 800 ms
DNS + TCP + TLS (cached) 20 ms
Network RTT (regional) 40 ms
Gateway (auth, rate limit, route) 15 ms
Tokenization 5 ms
Queue wait 300 ms ← the flex
Prefill (p95 prompt = 1,800 tok) 180 ms
First sample 2 ms
SSE framing + first write 8 ms
─────────────────────────────────────────
Total 570 ms
Headroom 230 msEach line is owned by someone and monitored. When the SLO is violated, the budget tells you which line grew.
The “queue wait” line is the flex — it’s what absorbs load variation, and it’s what capacity planning (file 02) sizes for.
6. Error budgets#
SLO: 99.9% availability over 30 days
→ error budget = 0.1% = 43.2 minutes of "down" per month
Consumption:
a 15-minute incident → 35% of the budget
0.05% baseline error rate → 50% of the budget continuously
a bad deploy rolled back
after 8 minutes → 19%
POLICY
budget > 50% remaining → ship freely, take risks
budget 20-50% → ship carefully, prioritize reliability work
budget < 20% → feature freeze; only reliability work
budget exhausted → no deploys except fixesThe error budget converts “how reliable should we be?” from an argument into arithmetic. It also gives you a principled reason to say no to a risky deployment.
For LLM services, consider a quality error budget too: a budget for how much quality regression you’ll accept from optimization work, spent deliberately.
7. Measuring the SLIs correctly#
✓ Measure from the CLIENT side, through the real ingress (Section X, case 8)
✓ Use histograms, not averages
✓ Bucket TTFT by prompt length
✓ Exclude requests that violated published limits
✓ Count 429s separately (they're a healthy response to over-quota, not a failure)
✓ Compute goodput per request, not per metric
✓ Use rolling windows matched to the SLO's window# Goodput: requests meeting BOTH objectives
sum(rate(requests_total{ttft_ok="true", itl_ok="true"}[5m]))
/ sum(rate(requests_total{excluded="false"}[5m]))Computing ttft_ok and itl_ok per request (as labels, or in structured logs) is more work than
computing them from separate histograms — and it’s the only way to get goodput right, because a
request can meet p95 TTFT while violating p95 ITL.
8. Tiering#
Different customers need different promises, and different promises need different infrastructure:
TIER TTFT p95 ITL p95 Availability Max context Priority
free 3 s 120 ms 99.0% 8k low
standard 800 ms 50 ms 99.9% 32k normal
premium 400 ms 30 ms 99.95% 128k high
batch n/a n/a 99.9% 128k lowest
(E2E < 24h)Tiering is how you serve everyone economically. The free tier runs at high batch on oversubscribed capacity; premium runs at low batch on reserved capacity.
Implementation: priority classes in the scheduler (Section VIII.03), separate replica pools for premium, and admission control that sheds free-tier load first.
9. Production implications#
- Write the SLO down and publish it. Internally at minimum; to customers if you have an SLA.
- Instrument goodput. It’s the metric that aligns engineering with product.
- Build the latency budget and monitor each line.
- Use the error budget to govern deployment risk.
- Include quality in the SLO. For LLMs, “it responded” is insufficient.
- Publish limits (max input, max output, rate limits) and exclude limit-violating requests from the SLO.
- Review the SLO quarterly against what users actually experience and complain about.
10. Common mistakes#
A single latency SLO. Meaningless for LLM services.
Unbucketed TTFT. Violated by every long prompt.
No ITL SLO. The user experience of streaming is ITL, not E2E.
Availability only. Ignores the quality dimension entirely.
Measuring server-side. Misses ingress problems.
Counting 429s as failures. They’re a healthy response to over-quota.
SLOs set by aspiration rather than measurement. Promise what you can deliver.
No error budget policy. Then reliability arguments are political rather than quantitative.
11. Hands-on exercise#
A. Write the SLO. For a service you work on (or a hypothetical chat product), write the full SLO in the format of section 3. Justify every threshold from either user experience or achievability.
B. Build the latency budget. Decompose your TTFT SLO into components. Measure each on a real system. Where’s your headroom? Which component is closest to its budget?
C. Implement goodput. Compute per-request ttft_ok and itl_ok, emit them, and build the
goodput metric. Compare goodput to raw throughput under load — do they diverge?
D. Error budget. Compute your current error budget consumption for the last 30 days. At the current rate, will you exhaust it? What policy would you set?
E. Tier design. Design a three-tier offering for a model you serve. For each tier, specify the SLO and the infrastructure implication (batch size, priority, dedicated capacity).
12. Interview questions#
- Why is a single latency SLO wrong for an LLM service?
- Write an SLO for a chat service and justify each threshold.
- What is goodput and why is it a better target than throughput?
- Why must TTFT be bucketed by prompt length?
- What is an error budget and how would you use it?
- Should a quality metric be in the SLO? How would you define it?
- How do you implement service tiers in an inference system?
13. Further reading#
- [FUNDAMENTAL] Google SRE Book and The Site Reliability Workbook, SLO chapters
- [FUNDAMENTAL] Dean & Barroso, “The Tail at Scale”
- [REFERENCE] Published SLAs from LLM API providers — read what they actually promise
- Next: 02 — Capacity planning