The idea in one minute#
Between your process and a graph there is a pipeline with four stages: collect (scrape or receive), process (enrich, filter, sample), store (a database shaped for that signal), and query (a language, dashboards, alert rules). Each signal wants a different storage engine, because each has a different shape: metrics are many small numeric series, logs are large volumes of text, traces are trees looked up by ID.
The design questions are always the same: where is the buffer when a stage is down, how long is data kept and at what resolution, and who watches the watcher.
An analogy#
A city’s water system: intake, treatment, reservoirs, taps. Nobody at the tap thinks about the reservoir until it is empty. Capacity, redundancy and pressure are engineered at each stage — and a treatment plant that fails silently is worse than one that fails loudly.
A picture#
flowchart TB
subgraph SRC["Sources"]
APP["Apps<br/>SDKs"]
EXP["Exporters<br/>node, GPU, kube-state"]
K8S["Kubernetes<br/>events, container logs"]
end
SRC --> AG["Agent per node<br/>scrape, tail, receive"]
AG --> BUF["Buffer<br/>queue or Kafka"]
BUF --> GWC["Gateway collectors<br/>enrich, redact, sample"]
GWC --> M[("Metrics TSDB")]
GWC --> L[("Log store")]
GWC --> T[("Trace store")]
GWC --> P[("Profile store")]
M --> QRY["Query layer"]
L --> QRY
T --> QRY
P --> QRY
QRY --> D["Dashboards"]
QRY --> AL["Alert rules"]
AL --> AM["Alert router<br/>dedupe, group, silence"]
AM --> PG["Pager, chat, ticket"]
class APP,EXP,K8S neutral
class AG,GWC io
class BUF queue
class M,L,T,P memory
class QRY,AL,AM queue
class D,PG neutralHow it really works#
Stage 1: collect#
| Source | How it is collected |
|---|---|
| Application metrics | Scrape /metrics, or OTLP push |
| Host | node_exporter (CPU, memory, disk, network) |
| Containers | cAdvisor, built into the kubelet |
| Kubernetes object state | kube-state-metrics (desired vs actual replicas, pod phase, resource requests) |
| GPUs | A device exporter such as NVIDIA’s DCGM exporter (V.02) |
| Logs | An agent tailing container stdout |
| Traces | OTLP from SDKs or from eBPF instrumentation |
Service discovery tells the collector what exists. In Kubernetes, the Prometheus Operator’s
ServiceMonitor and PodMonitor objects declare “scrape everything with this label”.
Stage 2: process#
Done in the agent or gateway collector (III.03): attach Kubernetes and cloud metadata, drop or rename noisy labels, redact sensitive fields, sample traces, route by tenant or signal. Every transformation done here is one that does not have to be done in every application.
Stage 3: store#
| Signal | Storage shape | Common systems |
|---|---|---|
| Metrics | Time-series database: compressed chunks per series plus a label index | Prometheus; for scale and long retention Grafana Mimir, Thanos, Cortex, VictoriaMetrics |
| Logs | Label-indexed chunks, or an inverted index, or columns | Grafana Loki; Elasticsearch / OpenSearch; ClickHouse |
| Traces | Blocks in object storage looked up by trace ID, or columns | Grafana Tempo; Jaeger (v2 is built on the OTel Collector); ClickHouse |
| Profiles | Stack samples, columnar | Grafana Pyroscope; Parca |
| Everything as wide events | A column store | ClickHouse-based stacks, Honeycomb, and others |
A single Prometheus is a fine starting point: one binary, local disk, about two weeks of retention, millions of series. You outgrow it when you need high availability, more than one cluster, or long retention. Then you add remote write to a horizontally-scaled store that keeps blocks in object storage — the pattern shared by Mimir, Thanos and Cortex.
Stage 4: query, visualize, alert#
- Grafana is the usual front end for all four signals; most backends also have their own UI.
- Alert rules run on a schedule against the metrics store; Alertmanager (or the equivalent) deduplicates, groups, silences and routes notifications.
- Recording rules pre-compute expensive queries (II.03).
Reliability of the pipeline itself#
| Risk | Mitigation |
|---|---|
| Backend down, data lost | Agents buffer on disk; a queue (Kafka) between agents and backends for large systems |
| A burst overwhelms a stage | memory_limiter and bounded queues; drop the least valuable data first |
| One team’s cardinality explosion takes down everyone | Per-tenant limits on series and ingestion rate |
| The monitoring system fails silently | A dead man’s switch: an alert that always fires; an external service pages when it stops arriving |
| Monitoring shares the fate of what it monitors | Run it in a separate cluster or account, or use an external probe |
Retention and resolution#
Not all data deserves the same lifetime:
raw metrics, 15-30 s resolution 2-4 weeks
downsampled (5 min, 1 h) 1-2 years capacity planning, year-over-year
raw logs / wide events 3-30 days shortened by volume
sampled traces 7-14 days
SLO recording rules a year or more tiny, and the record of how you didBuild or buy#
| Self-hosted open source | Managed open source | Commercial platform | |
|---|---|---|---|
| Examples | Prometheus, Loki, Tempo, Grafana, ClickHouse | Grafana Cloud, cloud-provider Prometheus services | Datadog, Dynatrace, New Relic, Honeycomb, Splunk, Chronosphere |
| Cost | Hardware and engineers | Usage-based, moderate | Usage-based, often the largest line after compute |
| Effort | High | Low | Lowest |
| Lock-in | Low | Low if you use OTLP and PromQL | Lower than it used to be if you instrument with OTel |
Whatever you choose, instrument with open standards (OTel, Prometheus exposition). That keeps the choice of backend reversible, which is the only durable protection against a bill you did not expect.
Code#
A pipeline in miniature: a scraper that parses the text format, a ring-buffer store, and a query — enough to see why a failed scrape is itself a signal.
// pipeline.go — scrape → store → query, with `up` recorded for every scrape.
package main
import (
"bufio"
"fmt"
"strconv"
"strings"
)
type Sample struct {
T int
V float64
}
type Store map[string][]Sample // series → samples
func (s Store) append(series string, t int, v float64) {
s[series] = append(s[series], Sample{t, v})
if len(s[series]) > 240 { // retention: keep the newest 240 samples
s[series] = s[series][1:]
}
}
// scrape parses Prometheus text exposition and writes samples; it always records `up`.
func scrape(store Store, target string, t int, body string, ok bool) {
up := 0.0
if ok {
up = 1
sc := bufio.NewScanner(strings.NewReader(body))
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
i := strings.LastIndexByte(line, ' ')
v, err := strconv.ParseFloat(line[i+1:], 64)
if err != nil {
continue
}
name := line[:i]
// the collector adds a target label to everything it scrapes
if strings.HasSuffix(name, "}") {
name = name[:len(name)-1] + `,instance="` + target + `"}`
} else {
name += `{instance="` + target + `"}`
}
store.append(name, t, v)
}
}
store.append(`up{instance="`+target+`"}`, t, up)
}
func rate(s []Sample, window int) float64 {
if len(s) < 2 {
return 0
}
last := s[len(s)-1]
first := last
for i := len(s) - 1; i >= 0 && last.T-s[i].T <= window; i-- {
first = s[i]
}
if last.T == first.T {
return 0
}
return (last.V - first.V) / float64(last.T-first.T)
}
func main() {
store := Store{}
requests := 0.0
for t := 0; t <= 300; t += 15 { // a scrape every 15 s
requests += 15 * 40 // the app serves 40 requests/s
body := fmt.Sprintf("# TYPE http_requests_total counter\nhttp_requests_total{code=\"200\"} %g\n", requests)
ok := t < 180 || t > 240 // the target is unreachable for a minute
scrape(store, "api-0", t, body, ok)
}
fmt.Println("t(s) up")
for _, s := range store[`up{instance="api-0"}`] {
if s.V == 0 {
fmt.Printf("%4d 0 ← alert on this: `up == 0`\n", s.T)
}
}
series := `http_requests_total{code="200",instance="api-0"}`
fmt.Printf("\nsamples stored: %d (gap during the outage)\n", len(store[series]))
fmt.Printf("rate over the last 2 minutes: %.1f req/s — the counter loses nothing across the gap\n",
rate(store[series], 120))
}Remember this#
- Four stages: collect, process, store, query. Each signal has its own storage shape.
- Start with one Prometheus; add remote write and object storage when you need HA or retention.
- Buffer between stages, limit per tenant, and monitor the monitoring with a dead man’s switch.
- Instrument with open standards so the backend remains a choice.
Try it#
- Run
pipeline.go. Change the app’s counter to a requests-per-second gauge. What is lost during the gap? - Draw the pipeline for a system you operate. Mark every point where data could be lost and where it is buffered.
- Write a retention table like the one above for your own signals, with a reason for each row.
Check yourself#
- Why do metrics, logs and traces use different storage engines?
- What is a dead man’s switch alert?
- What is the single most effective protection against backend lock-in?