PidokuInfra

SLOs, SLIs, and Latency Budgets

Advanced Intermediate 1h 15m Difficulty 3/5 Topic 01 of 11

Prerequisites I.05, VIII.03


1. The definitions#

SLI  Service Level INDICATOR   — a measurement. "p95 TTFT over 5 minutes."
SLO  Service Level OBJECTIVE   — a target for an SLI. "p95 TTFT < 800 ms."
SLA  Service Level AGREEMENT   — a contractual promise with consequences.

For LLM services, the SLI definitions require more care than for ordinary services, because “latency” is not one number.


2. Why LLM SLOs are different#

ORDINARY SERVICE               LLM SERVICE
one latency metric             TTFT and ITL are separate experiences
requests are comparable        cost varies 1000x by request shape
success is 2xx                 success includes "the answer was good"
latency is roughly stationary  latency depends on prompt and output length

A single “p99 latency < 2 s” SLO is meaningless for an LLM service. It will be violated by every long generation and satisfied by every short one, telling you nothing about health.


3. A well-formed LLM SLO#

SERVICE: chat-completions, model llama-3-70b, standard tier

AVAILABILITY
  99.9% of requests return a non-5xx response
  measured over a rolling 30 days

LATENCY — TIME TO FIRST TOKEN
  p95 TTFT < 800 ms   for requests with ≤ 2,000 prompt tokens
  p95 TTFT < 3,000 ms for requests with 2,000-16,000 prompt tokens
  p95 TTFT < 12,000 ms for requests with > 16,000 prompt tokens
  measured over rolling 5-minute windows

LATENCY — INTER-TOKEN
  p95 ITL < 50 ms
  p99 ITL < 120 ms
  measured over rolling 5-minute windows

THROUGHPUT (capacity commitment)
  the service accepts ≥ 40 requests/second at the above latencies

GOODPUT (the composite)
  ≥ 97% of requests meet BOTH their TTFT and ITL objectives

QUALITY
  the deployed model version's evaluation suite score is within 2% of
  the reference, measured at each deployment

EXCLUSIONS
  requests exceeding published limits (max input 32k, max output 4k)
  requests rejected with 429 (over quota) — counted separately

Four things that make this good:

  1. TTFT is bucketed by prompt length. Without this, one user pasting a book violates your SLO.
  2. ITL has both p95 and p99. The p99 catches blocking events (Section X, case 2).
  3. Goodput is the composite metric. It can’t be gamed by making everyone uniformly slow.
  4. Quality is an SLO. For an LLM service, “it responded” is not “it worked.”

4. Choosing the numbers#

Don’t invent thresholds. Derive them.

TTFT#

What does the user experience?
  < 200 ms   feels instant
  200-500 ms feels responsive
  0.5-1 s    noticeable but acceptable for a "thinking" interaction
  1-3 s      requires a loading indicator; users tolerate it if warned
  > 3 s      users leave or retry

Then check what's ACHIEVABLE:
  TTFT_floor = tokenize + prefill(p95_prompt_length) + network_rtt
  For a 70B model, 2,000-token prompt: ~150 ms prefill + 50 ms network
  → 800 ms p95 leaves 600 ms of queueing headroom. Reasonable.

Set the SLO where the achievable meets the acceptable, with headroom.

ITL#

Human reading speed: 200-300 words/min ≈ 4-5 words/sec ≈ 5-7 tokens/sec
                     → 140-200 ms per token is "reading speed"

But: users don't read as it generates; they wait for a chunk then read.
     And perceived smoothness matters.

  < 30 ms   feels like fast typing; smooth
  30-60 ms  smooth
  60-100 ms slightly choppy but fine
  > 150 ms  visibly slow

→ p95 ITL < 50 ms is a good target for chat.
→ For agentic/batch workloads, ITL barely matters; E2E does.

Availability#

99.9%  = 43 minutes/month of downtime. Achievable with N+1 and good practices.
99.95% = 22 minutes/month. Needs multi-AZ and careful deployments.
99.99% = 4.3 minutes/month. Needs multi-region and a lot of engineering.

For a GPU service, be honest: GPU failures, driver issues, and long cold
starts make 99.99% expensive. Promise 99.9% and deliver it.

5. The latency budget#

Decompose the SLO into a budget you can hold each component to:

p95 TTFT budget: 800 ms

  DNS + TCP + TLS (cached)          20 ms
  Network RTT (regional)            40 ms
  Gateway (auth, rate limit, route) 15 ms
  Tokenization                       5 ms
  Queue wait                       300 ms   ← the flex
  Prefill (p95 prompt = 1,800 tok) 180 ms
  First sample                       2 ms
  SSE framing + first write          8 ms
  ─────────────────────────────────────────
  Total                            570 ms
  Headroom                         230 ms

Each line is owned by someone and monitored. When the SLO is violated, the budget tells you which line grew.

The “queue wait” line is the flex — it’s what absorbs load variation, and it’s what capacity planning (file 02) sizes for.


6. Error budgets#

SLO: 99.9% availability over 30 days
→ error budget = 0.1% = 43.2 minutes of "down" per month

Consumption:
  a 15-minute incident         → 35% of the budget
  0.05% baseline error rate    → 50% of the budget continuously
  a bad deploy rolled back
    after 8 minutes            → 19%

POLICY
  budget > 50% remaining  → ship freely, take risks
  budget 20-50%           → ship carefully, prioritize reliability work
  budget < 20%            → feature freeze; only reliability work
  budget exhausted        → no deploys except fixes

The error budget converts “how reliable should we be?” from an argument into arithmetic. It also gives you a principled reason to say no to a risky deployment.

For LLM services, consider a quality error budget too: a budget for how much quality regression you’ll accept from optimization work, spent deliberately.


7. Measuring the SLIs correctly#

✓ Measure from the CLIENT side, through the real ingress (Section X, case 8)
✓ Use histograms, not averages
✓ Bucket TTFT by prompt length
✓ Exclude requests that violated published limits
✓ Count 429s separately (they're a healthy response to over-quota, not a failure)
✓ Compute goodput per request, not per metric
✓ Use rolling windows matched to the SLO's window
PromQL
# Goodput: requests meeting BOTH objectives
sum(rate(requests_total{ttft_ok="true", itl_ok="true"}[5m]))
/ sum(rate(requests_total{excluded="false"}[5m]))

Computing ttft_ok and itl_ok per request (as labels, or in structured logs) is more work than computing them from separate histograms — and it’s the only way to get goodput right, because a request can meet p95 TTFT while violating p95 ITL.


8. Tiering#

Different customers need different promises, and different promises need different infrastructure:

TIER          TTFT p95   ITL p95   Availability   Max context   Priority
free            3 s       120 ms      99.0%           8k          low
standard      800 ms       50 ms      99.9%          32k          normal
premium       400 ms       30 ms      99.95%        128k          high
batch          n/a         n/a        99.9%         128k          lowest
              (E2E < 24h)

Tiering is how you serve everyone economically. The free tier runs at high batch on oversubscribed capacity; premium runs at low batch on reserved capacity.

Implementation: priority classes in the scheduler (Section VIII.03), separate replica pools for premium, and admission control that sheds free-tier load first.


9. Production implications#

  • Write the SLO down and publish it. Internally at minimum; to customers if you have an SLA.
  • Instrument goodput. It’s the metric that aligns engineering with product.
  • Build the latency budget and monitor each line.
  • Use the error budget to govern deployment risk.
  • Include quality in the SLO. For LLMs, “it responded” is insufficient.
  • Publish limits (max input, max output, rate limits) and exclude limit-violating requests from the SLO.
  • Review the SLO quarterly against what users actually experience and complain about.

10. Common mistakes#

A single latency SLO. Meaningless for LLM services.

Unbucketed TTFT. Violated by every long prompt.

No ITL SLO. The user experience of streaming is ITL, not E2E.

Availability only. Ignores the quality dimension entirely.

Measuring server-side. Misses ingress problems.

Counting 429s as failures. They’re a healthy response to over-quota.

SLOs set by aspiration rather than measurement. Promise what you can deliver.

No error budget policy. Then reliability arguments are political rather than quantitative.


11. Hands-on exercise#

A. Write the SLO. For a service you work on (or a hypothetical chat product), write the full SLO in the format of section 3. Justify every threshold from either user experience or achievability.

B. Build the latency budget. Decompose your TTFT SLO into components. Measure each on a real system. Where’s your headroom? Which component is closest to its budget?

C. Implement goodput. Compute per-request ttft_ok and itl_ok, emit them, and build the goodput metric. Compare goodput to raw throughput under load — do they diverge?

D. Error budget. Compute your current error budget consumption for the last 30 days. At the current rate, will you exhaust it? What policy would you set?

E. Tier design. Design a three-tier offering for a model you serve. For each tier, specify the SLO and the infrastructure implication (batch size, priority, dedicated capacity).


12. Interview questions#

  1. Why is a single latency SLO wrong for an LLM service?
  2. Write an SLO for a chat service and justify each threshold.
  3. What is goodput and why is it a better target than throughput?
  4. Why must TTFT be bucketed by prompt length?
  5. What is an error budget and how would you use it?
  6. Should a quality metric be in the SLO? How would you define it?
  7. How do you implement service tiers in an inference system?

13. Further reading#

  • [FUNDAMENTAL] Google SRE Book and The Site Reliability Workbook, SLO chapters
  • [FUNDAMENTAL] Dean & Barroso, “The Tail at Scale”
  • [REFERENCE] Published SLAs from LLM API providers — read what they actually promise
  • Next: 02 — Capacity planning

↑↓ navigate↵ openesc close