PidokuInfra

Dashboards

Intermediate 40 min Difficulty 2/5 Topic 02 of 05

Prerequisites II.02, II.03

The idea in one minute#

A dashboard is an answer to a question, laid out so that a tired person can read it in ten seconds. Good dashboards are layered: one overview that says whether users are happy, then one screen per service showing the RED metrics, then resource detail showing USE. You start at the top and descend only where something looks wrong.

Most bad dashboards share one fault: they show everything that can be measured rather than what someone needs to decide.

An analogy#

A hospital monitor shows heart rate, blood pressure and oxygen — big, at the top. It does not show the full blood panel; that is a separate report you request when the vital signs are off.

A picture#

flowchart TB
  L1["Level 1: Service health<br/>SLO status, error budget, are users affected"] --> L2["Level 2: One service, RED<br/>rate, errors, duration by route or model"]
  L2 --> L3["Level 3: Resources, USE<br/>CPU, memory, queue depth, GPU"]
  L3 --> L4["Level 4: Drill-down<br/>traces, logs, profiles for one instance"]
  class L1 queue
  class L2 compute
  class L3 memory
  class L4 neutral

How it really works#

Layout rules#

  1. Symptoms at the top, causes below. The top row is what users feel.
  2. Read left to right as a sentence: traffic → errors → latency → saturation.
  3. One question per panel, stated in the title: “Error ratio by route”, not “Errors”.
  4. Units and thresholds on the graph. Draw the SLO line.
  5. Same time range and same colours everywhere. Errors are always the same colour.
  6. Mark deploys and config changes as annotations. Most incidents begin at one.
  7. Variables, not copies: one dashboard with a service or model drop-down, not twenty near-identical ones.

The service dashboard#

RowPanelsQuery shape
TrafficRequests/s by routesum by (route) (rate(requests_total[$__rate_interval]))
ErrorsError ratio, with the SLO lineerrors ÷ total
Durationp50, p95, p99; a heatmap beside ithistogram_quantile
SaturationIn-flight, queue depth, waitinggauges
DependenciesThe same RED row for each downstream callclient-side metrics
ResourcesCPU, memory, restarts, throttlingUSE

Choosing a visualization#

You want to showUse
A trendTime series line
A distribution over timeHeatmap
A current value against a limitStat or gauge with thresholds
RankingBar chart or table sorted by value
Many instances at onceA table, or a status grid — never forty overlapping lines
Parts of a whole over timeStacked area, only when the parts genuinely add up

Avoid pie charts for anything that changes, dual y-axes, and averaged latency.

Traps#

  • Averages across instances hide the one that is broken. Show max, or p99 across instances, or a top-5 table.
  • A graph that is flat at zero: is the service healthy, or is the metric missing? Show up and scrape health.
  • Stacking latencies or percentiles. They do not add.
  • Auto-scaled axes make noise look like an incident. Fix the y-axis from zero for ratios.
  • Too long a rate window smooths away a short spike; too short makes noise. Use the dashboard’s interval-aware variable.

Dashboards as code#

Keep dashboards in version control and generate them — Grafana’s JSON model, Jsonnet/Grafonnet, Terraform, or the Grafana Foundation SDK (which has a Go builder). Reasons: review, reuse of a standard service row across teams, and no more “who changed this panel?”.

The mixin pattern packages dashboards, recording rules and alerts for one piece of software together. Before building a dashboard for Kubernetes, a database or an inference engine, look for an existing mixin or the vendor’s published dashboard and adapt it.

When a dashboard is the wrong tool#

Dashboards answer questions you anticipated. During a novel incident you need ad-hoc queries over events and traces (I.01). If every incident ends with “let me build a dashboard for that”, you are accumulating screens nobody will open again. Build a dashboard when a question recurs.

Code#

A dashboard is data. This program generates a standard RED row for any list of services, which is all “dashboards as code” means.

Go
// dash.go — generate a RED dashboard definition from a list of services.
package main

import (
	"encoding/json"
	"fmt"
	"os"
)

type Panel struct {
	Title string `json:"title"`
	Type  string `json:"type"`
	Unit  string `json:"unit"`
	Expr  string `json:"expr"`
	Row   int    `json:"row"`
}

func redRow(row int, svc string) []Panel {
	sel := fmt.Sprintf(`{service=%q}`, svc)
	return []Panel{
		{svc + ": requests/s by route", "timeseries", "reqps",
			fmt.Sprintf(`sum by (route) (rate(http_requests_total%s[$__rate_interval]))`, sel), row},
		{svc + ": error ratio (SLO line at 0.1%)", "timeseries", "percentunit",
			fmt.Sprintf(`sum(rate(http_requests_total{service=%q,code=~"5.."}[$__rate_interval])) / sum(rate(http_requests_total%s[$__rate_interval]))`, svc, sel), row},
		{svc + ": p99 duration", "timeseries", "s",
			fmt.Sprintf(`histogram_quantile(0.99, sum by (route) (rate(http_request_duration_seconds%s[$__rate_interval])))`, sel), row},
		{svc + ": duration heatmap", "heatmap", "s",
			fmt.Sprintf(`sum(rate(http_request_duration_seconds%s[$__rate_interval]))`, sel), row},
		{svc + ": in flight (saturation)", "timeseries", "short",
			fmt.Sprintf(`sum(http_in_flight_requests%s)`, sel), row},
	}
}

func main() {
	services := []string{"gateway", "inference", "retrieval"}
	var panels []Panel
	for i, s := range services {
		panels = append(panels, redRow(i, s)...)
	}
	dash := map[string]any{
		"title":       "Service health (generated)",
		"annotations": []string{"deploys", "config changes"},
		"variables":   []string{"cluster", "namespace"},
		"panels":      panels,
	}
	enc := json.NewEncoder(os.Stdout)
	enc.SetIndent("", "  ")
	enc.Encode(dash)
	fmt.Fprintf(os.Stderr, "%d panels for %d services, all identical in shape\n", len(panels), len(services))
}

Remember this#

  • Layer dashboards: user-facing health → service RED → resource USE → drill-down.
  • One question per panel; symptoms above causes; mark deploys.
  • Never average across instances or stack percentiles.
  • Keep dashboards in version control and generate the repetitive parts.

Try it#

  1. Run dash.go and add a model variable to each query.
  2. Open a dashboard you use. Delete (on paper) every panel you have never acted on. What is left?
  3. Sketch the Level 1 dashboard for a product with three services. It may have at most six panels.

Check yourself#

  1. What goes on the top row of a service dashboard?
  2. Why is averaging a metric across instances dangerous?
  3. When should you not build a new dashboard?

↑↓ navigate↵ openesc close