PidokuInfra

Alerting on SLOs

Intermediate 55 min Difficulty 3/5 Topic 03 of 05

Prerequisites I.03, II.03

The idea in one minute#

An alert should mean: a human must act now, and here is what users are experiencing. That rules out most alerts people write. “CPU above 80%” is a cause that may or may not hurt anyone. “1% of requests are failing” is a symptom.

The dependable way to alert on symptoms is the error budget burn rate: page when the budget is being spent fast enough to run out soon, checked over a long window (so it is significant) and a short window (so it stops firing when the problem stops).

An analogy#

A smoke detector that sounds whenever the oven is on gets its battery removed. One that sounds only for real smoke gets obeyed. Every false page teaches people to ignore the next one; alert quality is a safety property.

A picture#

flowchart TB
  SLI["Error ratio<br/>from recording rules"] --> W1["Long window: 1 h<br/>burn rate above 14.4?"]
  SLI --> W2["Short window: 5 min<br/>burn rate above 14.4?"]
  W1 --> AND{"Both true?"}
  W2 --> AND
  AND -->|"yes"| PAGE["Page<br/>2% of the monthly budget gone in an hour"]
  AND -->|"no"| NEXT["Check the slower pair<br/>6 h and 30 min at 6x"]
  NEXT -->|"yes"| PAGE
  NEXT -->|"no"| TICKET["3 d and 6 h at 1x: open a ticket"]
  class SLI memory
  class W1,W2,AND,NEXT queue
  class PAGE warn
  class TICKET neutral

How it really works#

Symptoms, not causes#

Alert onNot on
Error ratio above budget burnA single 500
p99 latency SLI burning budgetCPU, memory or GPU utilization
Queue depth growing while goodput fallsQueue depth above a fixed number
Disk will be full in four hours (predict_linear)Disk at 80%
A scrape target has been down for minutesOne failed scrape

Causes belong on dashboards. The exception is an imminent, certain failure — a disk filling, a certificate expiring — where the “cause” is the symptom arriving on a schedule.

Why a simple threshold fails#

error_ratio > 0.001 for 5m on a 99.9% SLO either fires for harmless blips (short for) or takes an hour to notice a total outage (long for). The trouble is that one threshold cannot express both “a lot of budget, quickly” and “a steady leak”.

Multi-window, multi-burn-rate#

Burn rate = error ratio ÷ error budget (I.03). For a 30-day SLO:

SeverityBudget consumedLong windowShort windowBurn rate
Page2%1 h5 min14.4
Page5%6 h30 min6
Ticket10%3 d6 h1
YAML
groups:
  - name: slo-api
    rules:
      - record: job:slo_errors:ratio_rate5m
        expr: |
          sum(rate(http_requests_total{job="api",code=~"5.."}[5m]))
          / sum(rate(http_requests_total{job="api"}[5m]))
      - record: job:slo_errors:ratio_rate1h
        expr: |
          sum(rate(http_requests_total{job="api",code=~"5.."}[1h]))
          / sum(rate(http_requests_total{job="api"}[1h]))
      - alert: ErrorBudgetFastBurn
        expr: |
          job:slo_errors:ratio_rate1h > (14.4 * 0.001)
          and job:slo_errors:ratio_rate5m > (14.4 * 0.001)
        labels: { severity: page }
        annotations:
          summary: "API is burning its 30-day error budget 14x too fast"
          runbook: https://runbooks.example.internal/api/error-budget

The long window makes it significant; the short window makes it reset quickly once the problem is fixed. Tools such as Sloth and Pyrra generate these rules from a short SLO definition, and OpenSLO is a vendor-neutral format for writing the definition.

Low traffic#

With ten requests an hour, one failure is a 10% error ratio. Options: lengthen the windows, require a minimum request count, combine related services into one SLO, or add synthetic probes so there is always traffic to measure.

Routing and hygiene#

  • Severity: page (now, wakes someone), ticket (this week), info (dashboard only).
  • Group: one notification for “40 pods of service X are failing”, not forty.
  • Inhibit: if the cluster is down, suppress alerts for everything inside it.
  • Every page has a runbook link and names the user-visible impact.
  • Review: for each page, was action required? If not, fix or delete the alert. Track pages per on-call shift; more than a couple is a problem to engineer away.

Alerts worth having beyond SLOs#

AlertWhy
Dead man’s switchThe alerting pipeline itself is alive (IV.01)
up == 0 for N minutesA target vanished; its SLO alert would be silent
Absent metricInstrumentation was removed or renamed
Certificate / quota / disk exhaustion predictedCertain, scheduled failures
Telemetry volume anomalyA cardinality or log explosion before the bill

Code#

Simulate a month with three incidents and see which alerting rule catches what.

Go
// burn.go — threshold alert vs multi-window burn-rate alert on a simulated error ratio.
package main

import "fmt"

const (
	slo    = 0.999
	budget = 1 - slo
)

// window returns the mean error ratio over the last n minutes ending at t.
func window(errs []float64, t, n int) float64 {
	lo := t - n + 1
	if lo < 0 {
		lo = 0
	}
	s := 0.0
	for i := lo; i <= t; i++ {
		s += errs[i]
	}
	return s / float64(t-lo+1)
}

func main() {
	const minutes = 30 * 24 * 60
	errs := make([]float64, minutes)
	type incident struct {
		name       string
		start, len int
		ratio      float64
	}
	incidents := []incident{
		{"blip: 5% errors for 6 min", 5000, 6, 0.05},
		{"outage: 30% errors for 40 min", 15000, 40, 0.30},
		{"slow leak: 0.4% errors for 2 days", 25000, 2 * 24 * 60, 0.004},
	}
	for _, in := range incidents {
		for i := in.start; i < in.start+in.len; i++ {
			errs[i] = in.ratio
		}
	}

	firstFire := func(rule func(t int) bool, from, to int) int {
		for t := from; t < to; t++ {
			if rule(t) {
				return t - from
			}
		}
		return -1
	}
	rules := []struct {
		name string
		f    func(t int) bool
	}{
		{"threshold >0.1% for 5m", func(t int) bool {
			for i := t - 4; i <= t; i++ {
				if i < 0 || errs[i] <= budget {
					return false
				}
			}
			return true
		}},
		{"fast burn 14.4x (1h & 5m)", func(t int) bool {
			return window(errs, t, 60) > 14.4*budget && window(errs, t, 5) > 14.4*budget
		}},
		{"slow burn 6x (6h & 30m)", func(t int) bool {
			return window(errs, t, 360) > 6*budget && window(errs, t, 30) > 6*budget
		}},
		{"ticket 1x (3d & 6h)", func(t int) bool {
			return window(errs, t, 4320) > budget && window(errs, t, 360) > budget
		}},
	}

	fmt.Printf("%-36s", "incident")
	for _, r := range rules {
		fmt.Printf("  %-26s", r.name)
	}
	fmt.Println()
	for _, in := range incidents {
		fmt.Printf("%-36s", in.name)
		for _, r := range rules {
			d := firstFire(r.f, in.start, in.start+in.len+4500)
			if d < 0 {
				fmt.Printf("  %-26s", "silent")
			} else {
				fmt.Printf("  %-26s", fmt.Sprintf("fires after %d min", d))
			}
		}
		used := in.ratio * float64(in.len) / float64(minutes) / budget
		fmt.Printf("  budget used: %.0f%%\n", used*100)
	}
}

Read the table by row. The blip uses under 1% of the budget, yet the fixed threshold pages for it; the burn-rate rules stay silent. The outage is caught by the fast-burn rule within a couple of minutes. The slow leak consumes over a quarter of the budget: the fixed threshold pages for it at once — at 3 a.m., for something that can wait until morning — while the burn-rate rules turn it into a ticket.

One property worth knowing: a total outage trips the fast-burn rule in under a minute, because 100% errors for one minute is already 1.7% over the hour. That is the intended behaviour, not noise.

Remember this#

  • Page on symptoms users feel, at a rate that threatens the SLO.
  • Burn rate over a long and a short window: significant, and quick to reset.
  • Every page: actionable, urgent, with a runbook. Review and delete the rest.
  • Also alert on the pipeline itself and on absence.

Try it#

  1. Run burn.go. Change the SLO to 99.5% and see which rules change behaviour.
  2. Write the three burn-rate rules for a latency SLI (“99% of requests under 500 ms”).
  3. Take five alerts from a real system and classify each as symptom or cause. Which would you delete?

Check yourself#

  1. Why does one fixed threshold fail for both blips and slow leaks?
  2. What does each of the two windows contribute?
  3. What three things must every page have?

↑↓ navigate↵ openesc close