PidokuInfra

Metric Types and the Data Model

Basic Beginner 45 min Difficulty 2/5 Topic 01 of 05

Prerequisites I.02

The idea in one minute#

A time series is a stream of (timestamp, value) pairs identified by a name and a set of labels: http_requests_total{route="/v1/chat", code="200"}. Every distinct combination of label values is its own series.

There are three types worth knowing. A counter only goes up (requests served). A gauge goes up and down (requests in flight). A histogram is a set of counters, one per bucket, that records a distribution (request duration). Almost everything else is built from these.

An analogy#

A car has an odometer (counter: total kilometres, only increases, you subtract two readings to get distance travelled), a speedometer (gauge: the value right now) and, if you kept a tally of how many trips fell into “under 10 km”, “under 50 km”, “under 200 km”, a histogram of trip lengths.

A picture#

flowchart TB
  subgraph PROC["Your process"]
    C["counter<br/>requests_total"]
    G["gauge<br/>in_flight"]
    H["histogram<br/>duration buckets + sum + count"]
  end
  PROC -->|"GET /metrics every 15 s"| SCR["Scraper"]
  SCR --> TS[("Time series database<br/>one series per name + label set")]
  TS --> Q["Query<br/>rate, sum, quantile"]
  class C,G,H memory
  class SCR io
  class TS memory
  class Q queue

How it really works#

The data model#

http_requests_total{route="/v1/chat", code="200"}  41873   @ 09:00:15
http_requests_total{route="/v1/chat", code="200"}  42260   @ 09:00:30
└───── name ──────┘└────────── labels ──────────┘  value     timestamp
  • The name says what is measured. By convention it carries the unit and, for counters, the suffix _total: request_duration_seconds, tokens_generated_total.
  • Labels say which one. They are how you slice: by route, status, model, replica.
  • Use base units: seconds, bytes, joules — never milliseconds or megabytes.

The types#

TypeBehaviourQuery it withExample
CounterMonotonic; resets to 0 on restartrate() — never the raw valueRequests, errors, tokens, bytes
GaugeAny value at any timeThe value, or avg_over_time, max_over_timeQueue depth, memory used, temperature
HistogramCounters per bucket, plus _sum and _counthistogram_quantile()Durations, sizes
SummaryPercentiles computed inside the processRead directlyAvoid: cannot be aggregated across instances (I.04)

Why counters and not “requests per second” gauges? A counter loses nothing between scrapes: whatever interval you look at, the difference of two readings is exact. A rate gauge computed inside the process reflects only the instant it was read, and a missed scrape loses data forever. Export counters; compute rates at query time.

The exposition format#

A process serves its current values as plain text at /metrics:

# HELP http_requests_total Requests handled, by route and status code.
# TYPE http_requests_total counter
http_requests_total{route="/v1/chat",code="200"} 42260
http_requests_total{route="/v1/chat",code="500"} 17
# TYPE http_in_flight_requests gauge
http_in_flight_requests 9
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{le="0.1"} 30112
http_request_duration_seconds_bucket{le="0.5"} 41020
http_request_duration_seconds_bucket{le="+Inf"} 42277
http_request_duration_seconds_sum 9120.4
http_request_duration_seconds_count 42277

This is the Prometheus text format; OpenMetrics is its standardized successor and adds exemplars. Histogram buckets are cumulative: le="0.5" counts everything at or below 0.5 s.

Pull and push#

Pull (scrape)Push
Who initiatesThe collector fetches /metricsThe process sends data out
Used byPrometheusOpenTelemetry (OTLP), StatsD
StrengthA failed scrape is a signal (up == 0); no client configWorks for short-lived jobs and behind NAT
WeaknessNeeds service discovery and reachabilityA silent process looks the same as a healthy idle one

The two have converged: Prometheus 3 accepts OTLP pushes natively, and the OpenTelemetry Collector can scrape Prometheus endpoints (III.03, IV.01).

Naming rules that save you later#

  1. One metric, one meaning, one unit.
  2. Put the unit in the name: _seconds, _bytes, _total.
  3. Labels are for dimensions you will aggregate over or filter by, with a small, bounded set of values. Never a user ID, request ID, prompt or raw URL (lesson 05).
  4. Do not encode a label’s value in the name: requests_total{code="500"}, not requests_500_total.

Code#

All three types and the text format, in the standard library.

Go
// metrics.go — a counter, a gauge and a histogram, exposed in Prometheus text format.
package main

import (
	"fmt"
	"math/rand"
	"os"
	"sort"
	"strings"
	"sync"
)

type Counter struct {
	mu sync.Mutex
	v  map[string]float64 // label string → value
}

func (c *Counter) Add(labels string, n float64) {
	c.mu.Lock()
	defer c.mu.Unlock()
	if c.v == nil {
		c.v = map[string]float64{}
	}
	c.v[labels] += n // a counter never decreases
}

type Histogram struct {
	mu      sync.Mutex
	bounds  []float64 // upper bounds, ascending
	buckets []uint64  // non-cumulative counts; the last one is +Inf
	sum     float64
	count   uint64
}

func NewHistogram(bounds ...float64) *Histogram {
	return &Histogram{bounds: bounds, buckets: make([]uint64, len(bounds)+1)}
}

func (h *Histogram) Observe(x float64) {
	h.mu.Lock()
	defer h.mu.Unlock()
	i := sort.SearchFloat64s(h.bounds, x) // first bound >= x
	h.buckets[i]++
	h.sum += x
	h.count++
}

func main() {
	requests := &Counter{}
	inFlight := 0.0
	duration := NewHistogram(0.05, 0.1, 0.25, 0.5, 1, 2.5)

	rng := rand.New(rand.NewSource(3))
	for i := 0; i < 5000; i++ {
		code := "200"
		if rng.Float64() < 0.01 {
			code = "500"
		}
		requests.Add(`route="/v1/chat",code="`+code+`"`, 1)
		duration.Observe(0.03 + rng.ExpFloat64()*0.12)
	}
	inFlight = 9

	var b strings.Builder
	b.WriteString("# TYPE http_requests_total counter\n")
	keys := make([]string, 0, len(requests.v))
	for k := range requests.v {
		keys = append(keys, k)
	}
	sort.Strings(keys)
	for _, k := range keys {
		fmt.Fprintf(&b, "http_requests_total{%s} %g\n", k, requests.v[k])
	}
	fmt.Fprintf(&b, "# TYPE http_in_flight_requests gauge\nhttp_in_flight_requests %g\n", inFlight)

	b.WriteString("# TYPE http_request_duration_seconds histogram\n")
	cum := uint64(0)
	for i, ub := range duration.bounds {
		cum += duration.buckets[i]
		fmt.Fprintf(&b, "http_request_duration_seconds_bucket{le=\"%g\"} %d\n", ub, cum)
	}
	fmt.Fprintf(&b, "http_request_duration_seconds_bucket{le=\"+Inf\"} %d\n", duration.count)
	fmt.Fprintf(&b, "http_request_duration_seconds_sum %g\n", duration.sum)
	fmt.Fprintf(&b, "http_request_duration_seconds_count %d\n", duration.count)

	fmt.Fprint(os.Stdout, b.String())
}

Count the lines of output: that is the number of time series this process creates. One histogram with six bounds is nine series — for a single label combination.

Remember this#

  • A series is a name plus a label set. Every new label value is a new series.
  • Counters go up, gauges move freely, histograms are bucketed counters.
  • Export counters and compute rates at query time.
  • Base units, unit in the name, bounded label values.

Try it#

  1. Run metrics.go. Add a model label with three values to the histogram. How many series does the process expose now?
  2. Add a second counter, tokens_generated_total, and increment it by a random number per request. What would rate() of it mean?
  3. Why are the histogram buckets cumulative in the output but stored non-cumulatively in the program? What does each choice make cheap?

Check yourself#

  1. What identifies a time series?
  2. Why is a counter better than a requests-per-second gauge?
  3. What do _bucket, _sum and _count hold?

↑↓ navigate↵ openesc close