PidokuInfra

SLIs, SLOs and Error Budgets

Foundations Beginner 45 min Difficulty 2/5 Topic 03 of 04

Prerequisites 01, 02

The idea in one minute#

An SLI (service level indicator) is a measurement of what users experience, written as a ratio: good events ÷ all events. An SLO (service level objective) is a target for that ratio over a window: 99.5% of requests succeed, over 30 days. The error budget is what is left: the 0.5% you are allowed to fail.

This turns “is the service good?” into arithmetic, and gives you a principled answer to “should we page someone?”: page when the budget is being spent too fast.

An analogy#

A monthly mobile data plan. The plan is the SLO. Usage is the SLI. The remaining gigabytes are the error budget. You do not panic at every megabyte; you worry when the rate of use means you will run out before the month ends.

A picture#

flowchart TB
  EV["Every request"] --> G{"Good?<br/>succeeded and fast enough"}
  G -->|"yes"| GOOD["good count"]
  G -->|"no"| BAD["bad count"]
  GOOD --> SLI["SLI = good / total"]
  BAD --> SLI
  SLI --> CMP{"SLI vs SLO<br/>over 30 days"}
  CMP --> BUD["Error budget left"]
  BUD -->|"plenty"| SHIP["Ship features, take risk"]
  BUD -->|"burning fast"| PAGE["Page someone"]
  BUD -->|"gone"| FREEZE["Slow down, fix reliability"]
  class EV neutral
  class G,CMP queue
  class GOOD,BAD,SLI,BUD memory
  class PAGE,FREEZE warn
  class SHIP compute

How it really works#

Writing an SLI#

A good SLI is a ratio of events, measured as close to the user as you can get.

KindGood eventExample
AvailabilityRequest returned a non-5xx responsegood = status < 500
LatencyRequest finished within a thresholdgood = duration < 300 ms
QualityResponse was complete and correctgood = not truncated, no fallback used
FreshnessData newer than a thresholdgood = age < 60 s

Two rules:

  • Count events, do not average measurements. “99% of requests under 300 ms” is an SLI. “Average latency under 300 ms” hides the slow tail (lesson 04).
  • Exclude what the user caused. A 400 for a malformed request is not your failure. A 429 because you ran out of capacity is.

Choosing the target#

SLOAllowed bad time in 30 days
99%7 h 12 min
99.5%3 h 36 min
99.9%43 min
99.95%21.6 min
99.99%4.3 min

Each extra nine costs roughly ten times more engineering. Choose the lowest target your users would not notice, not the highest you can imagine. 100% is never the right answer: it forbids all change.

An SLA is a contract with penalties. Keep the internal SLO stricter than the SLA so you learn about trouble before you owe money.

Error budget and burn rate#

error budget   = 1 − SLO                         (99.9% → 0.1% of requests)
burn rate      = observed error ratio ÷ error budget

Burn rate 1 means you will spend exactly the budget over the window. Burn rate 14.4 means a 30-day budget is gone in 50 hours — or 2% of it in one hour. Alerting on burn rate (IV.03) is how you get paged for real problems and left alone for blips.

SLOs for LLM serving, briefly#

Streaming generation needs more than one latency SLI, because the user experiences two different waits:

SLIGood event
Time to first token (TTFT)First token within, say, 500 ms
Time per output token (TPOT)Each token within, say, 50 ms on average for that request
AvailabilityStream completed without a server error

A request is good only if it met all of them. The fraction of requests that did is called goodput when expressed as a rate, and it is the number a serving fleet is sized against. V.03 builds on this.

Code#

Go
// budget.go — SLI, error budget and burn rate from raw counts.
package main

import "fmt"

type Window struct {
	Name        string
	Total, Good float64
	Hours       float64
}

func main() {
	const slo = 0.999 // 99.9% over 30 days
	const windowHours = 30 * 24.0
	budget := 1 - slo

	fmt.Printf("SLO %.2f%%  → error budget %.2f%% of requests, or %.0f min of total outage per 30 days\n\n",
		slo*100, budget*100, budget*windowHours*60)

	windows := []Window{
		{"quiet hour", 360000, 359900, 1},
		{"bad deploy, 1 h", 360000, 354600, 1},
		{"slow leak, 6 h", 2160000, 2153500, 6},
	}
	fmt.Println("window              SLI       error ratio  burn rate  budget used  exhausts in")
	for _, w := range windows {
		sli := w.Good / w.Total
		errRatio := 1 - sli
		burn := errRatio / budget
		used := burn * w.Hours / windowHours
		exhaust := "never at this rate"
		if burn > 1 {
			exhaust = fmt.Sprintf("%.1f days", windowHours/burn/24)
		}
		fmt.Printf("%-18s  %.4f%%  %9.3f%%  %8.1fx  %10.1f%%  %s\n",
			w.Name, sli*100, errRatio*100, burn, used*100, exhaust)
	}
	fmt.Println("\nPage when burn rate is high over a short AND a longer window (lesson IV.03).")
}

Remember this#

  • SLI = good events ÷ total events, measured near the user.
  • SLO = a target for the SLI over a window. Error budget = 1 − SLO.
  • Burn rate says how fast the budget is being spent; it is what you alert on.
  • LLM serving needs several SLIs at once (TTFT, per-token speed, completion); a request is good only if it meets all of them.

Try it#

  1. Run budget.go. Change the SLO to 99.5%. Which windows still look alarming?
  2. Write an SLI for a service you use, as a precise good/total definition. Which requests did you exclude, and why?
  3. A team says “our SLO is 100%”. Write two sentences explaining what that would forbid.

Check yourself#

  1. Why is an SLI a ratio of events and not an average?
  2. How many minutes of total outage does 99.9% allow in 30 days?
  3. What does a burn rate of 6 mean?

↑↓ navigate↵ openesc close