The idea in one minute#
For any request-serving component, start with three metrics — Rate, Errors, Duration (RED): how many requests, how many failed, how long they took. For any resource — a CPU, a queue, a GPU — use Utilization, Saturation, Errors (USE): how busy, how much is waiting, how many faults.
Instrument at the boundary: one piece of middleware that wraps every handler. Hand-placed counters scattered through the code are how services end up half-measured.
An analogy#
A restaurant manager needs few numbers. For the dining room: customers per hour, complaints, how long people wait for food (RED). For the kitchen: how many burners are in use, how many orders are stacked on the rail, how many dishes were sent back (USE). Everything else is detail you look up when one of those six looks wrong.
A picture#
flowchart LR
REQ["Request"] --> MW["Metrics middleware<br/>start timer, in_flight +1"]
MW --> H["Your handler"]
H --> MW2["Middleware, on the way out<br/>requests_total +1 by route and code<br/>observe duration, in_flight -1"]
MW2 --> RESP["Response"]
MW2 --> REG[("Registry")]
REG -->|"/metrics"| SCR["Prometheus or<br/>OTel Collector"]
class REQ,RESP neutral
class MW,MW2 queue
class H compute
class REG memory
class SCR ioHow it really works#
RED for services, USE for resources#
| Method | Metric | Type | Typical name |
|---|---|---|---|
| RED | Rate | Counter | http_requests_total{route,code} |
| RED | Errors | The same counter, filtered | ...{code=~"5.."} |
| RED | Duration | Histogram | http_request_duration_seconds{route} |
| USE | Utilization | Gauge or counter of busy time | process_cpu_seconds_total, GPU busy time |
| USE | Saturation | Gauge | Queue length, requests waiting |
| USE | Errors | Counter | Device errors, OOM kills |
Google’s “four golden signals” are the same idea merged: latency, traffic, errors, saturation.
Saturation is the one people forget. Utilization tells you a resource is busy; saturation tells you work is waiting for it. A queue that grows is the earliest warning you get, and for an LLM server the number of waiting requests is the best autoscaling signal there is (V.05).
What to label#
Label with what you will group by: route (the template, e.g. /v1/users/:id, never the raw
path), method, code, and for inference model. Keep every label’s value set small and
fixed. Record the status class or code, not the error message.
Things every Go service should also expose#
The Go client library registers these for free:
go_goroutines,go_gc_duration_seconds,go_memstats_*— runtime health.process_cpu_seconds_total,process_resident_memory_bytes,process_open_fds.- A build-info gauge fixed at 1 with labels for version and commit, so every graph can be
split by release:
myservice_build_info{version="1.8.2",commit="a3f9c1"} 1.
With the real library#
In production use github.com/prometheus/client_golang (or the OpenTelemetry metrics SDK,
III.03). The middleware is the same shape as the hand-written one below:
var (
reqs = promauto.NewCounterVec(prometheus.CounterOpts{
Name: "http_requests_total", Help: "Requests by route and code.",
}, []string{"route", "code"})
dur = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Help: "Request duration.",
NativeHistogramBucketFactor: 1.1, // native histogram (lesson 04)
}, []string{"route"})
)
func instrument(route string, next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
start := time.Now()
rec := &statusRecorder{ResponseWriter: w, code: 200}
next.ServeHTTP(rec, r)
reqs.WithLabelValues(route, strconv.Itoa(rec.code)).Inc()
dur.WithLabelValues(route).Observe(time.Since(start).Seconds())
})
}
// http.Handle("/metrics", promhttp.Handler())Common mistakes#
| Mistake | Consequence |
|---|---|
| Timing only successful requests | Fast failures make latency look better during an outage |
| Raw URL path as a label | Unbounded series (lesson 05) |
| Measuring inside the handler, after the queue | Queue time is invisible; the slowest part is missed |
| Milliseconds in one metric, seconds in another | Every dashboard needs a conversion and someone gets it wrong |
A gauge set from many goroutines without Inc/Dec | Lost updates |
Code#
A complete instrumented HTTP server using only the standard library. Run it, then
curl localhost:8080/work a few times and curl localhost:8080/metrics.
// server.go — RED middleware and a /metrics endpoint, no dependencies.
package main
import (
"fmt"
"log"
"math/rand"
"net/http"
"sort"
"sync"
"sync/atomic"
"time"
)
var bounds = []float64{0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5}
type routeStats struct {
codes map[int]uint64
buckets []uint64
sum float64
count uint64
}
var (
mu sync.Mutex
stats = map[string]*routeStats{}
inFlight atomic.Int64
)
type recorder struct {
http.ResponseWriter
code int
}
func (r *recorder) WriteHeader(c int) { r.code = c; r.ResponseWriter.WriteHeader(c) }
func instrument(route string, next http.HandlerFunc) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
start := time.Now()
inFlight.Add(1)
rec := &recorder{ResponseWriter: w, code: http.StatusOK}
next(rec, r)
inFlight.Add(-1)
secs := time.Since(start).Seconds()
mu.Lock()
defer mu.Unlock()
s := stats[route]
if s == nil {
s = &routeStats{codes: map[int]uint64{}, buckets: make([]uint64, len(bounds)+1)}
stats[route] = s
}
s.codes[rec.code]++
s.buckets[sort.SearchFloat64s(bounds, secs)]++
s.sum += secs
s.count++
}
}
func metrics(w http.ResponseWriter, _ *http.Request) {
mu.Lock()
defer mu.Unlock()
fmt.Fprintln(w, "# TYPE http_requests_total counter")
for route, s := range stats {
for code, n := range s.codes {
fmt.Fprintf(w, "http_requests_total{route=%q,code=\"%d\"} %d\n", route, code, n)
}
}
fmt.Fprintf(w, "# TYPE http_in_flight_requests gauge\nhttp_in_flight_requests %d\n", inFlight.Load())
fmt.Fprintln(w, "# TYPE http_request_duration_seconds histogram")
for route, s := range stats {
cum := uint64(0)
for i, ub := range bounds {
cum += s.buckets[i]
fmt.Fprintf(w, "http_request_duration_seconds_bucket{route=%q,le=\"%g\"} %d\n", route, ub, cum)
}
fmt.Fprintf(w, "http_request_duration_seconds_bucket{route=%q,le=\"+Inf\"} %d\n", route, s.count)
fmt.Fprintf(w, "http_request_duration_seconds_sum{route=%q} %g\n", route, s.sum)
fmt.Fprintf(w, "http_request_duration_seconds_count{route=%q} %d\n", route, s.count)
}
}
func work(w http.ResponseWriter, _ *http.Request) {
time.Sleep(time.Duration(10+rand.Intn(80)) * time.Millisecond)
if rand.Float64() < 0.03 {
http.Error(w, "upstream failed", http.StatusBadGateway)
return
}
fmt.Fprintln(w, "ok")
}
func main() {
http.HandleFunc("/work", instrument("/work", work))
http.HandleFunc("/metrics", metrics)
log.Println("listening on :8080")
log.Fatal(http.ListenAndServe(":8080", nil))
}Remember this#
- RED for services: rate, errors, duration. USE for resources: utilization, saturation, errors.
- Instrument once, at the boundary, in middleware.
- Time every request, including the failures and the time spent queued.
- Saturation — work waiting — is the earliest and most useful warning.
Try it#
- Run
server.goand send 200 requests with a shell loop. Read the bucket lines: roughly what is the median? - Add a
build_infogauge with aversionlabel. - Add a semaphore that admits only four concurrent requests, and a gauge for the number waiting. Load the server and watch the saturation gauge. Which metric moved first: waiting, or p99 duration?
Check yourself#
- What do RED and USE stand for, and when do you use each?
- Why must failed requests be timed too?
- Why is the route template used as a label rather than the request path?