The idea in one minute#
Observability is moving in two directions at once. The general field is consolidating on open standards, adding a fourth signal, instrumenting from the kernel, and shifting from pre-aggregated metrics toward wide events in column stores — all under pressure from a bill that grows faster than the systems it watches. AI infrastructure is pulling the field toward new units — tokens, joules, dollars per task — and toward treating output quality as telemetry.
And one loop closes on itself: AI systems are now both the thing being observed and, more and more, the thing doing the observing.
A picture#
flowchart TB
subgraph GEN["The general field"]
G1["Standards converge<br/>OTel + Prometheus"]
G2["Fourth signal<br/>profiles"]
G3["Zero-code<br/>eBPF"]
G4["Wide events<br/>column stores"]
G5["Cost discipline<br/>pipelines, own your storage"]
end
subgraph AIO["Observability for AI"]
A1["Tokens, not requests"]
A2["Power and energy<br/>as first-class metrics"]
A3["Deeper engine telemetry<br/>KV cache, goodput"]
A4["Standard GenAI traces<br/>agents, tools, MCP"]
A5["Quality as a signal"]
A6["Heterogeneous hardware"]
end
subgraph AIB["AI for observability"]
B1["Agents that investigate"]
B2["Telemetry exposed to<br/>models through MCP"]
end
GEN --> AIO
AIO --> AIB
AIB -->|"needs clean, standard,<br/>joined data"| GEN
class G1,G2,G3,G4,G5 compute
class A1,A2,A3,A4,A5,A6 memory
class B1,B2 queueHow it really works#
Each shift below is given as: what it was, what it is becoming, why, and what to do. Dates are as of 3 October 2026.
Part 1 — the general field#
1. Instrumentation has a standard. Was: each vendor’s agent and SDK. Becoming: OpenTelemetry everywhere, with Prometheus as a peer rather than a rival — Prometheus 3 ingests OTLP and UTF-8 names natively, and its native histograms (stable since 3.8) are the same design as OTel’s exponential histograms. OTel’s declarative configuration reached stable in 2026. Why: lock-in through instrumentation became unacceptable once the bill grew. Do: new code uses the OTel API. Keep Prometheus exposition where it already works.
2. Profiles became the fourth signal. Was: profiling as a tool you ran during an incident. Becoming: always-on, fleet-wide, linked to traces. OpenTelemetry Profiles entered public alpha in March 2026 with an eBPF profiler in the Collector. Why: compute is the largest cost line, and “which code is spending it” was the one question the other three signals could not answer. Do: turn on continuous profiling for anything expensive; expect the format to settle.
3. Instrumentation moved into the kernel. Was: a library added to every service. Becoming: eBPF agents producing traces and RED metrics for anything that speaks HTTP, gRPC or SQL. OpenTelemetry eBPF Instrumentation (OBI) reached beta in April 2026 and is working toward 1.0. Why: coverage. The services nobody instruments are the ones that cause incidents. Do: eBPF for the baseline, SDKs where you need business attributes.
4. From three pillars to wide events. Was: metrics, logs and traces in three databases, joined in your head. Becoming: one wide event per unit of work in a column store, with metrics derived at query time and high-cardinality fields welcome — the idea marketed as “observability 2.0”. ClickHouse-based back ends are now common inside both open-source and commercial products. Why: columnar storage and object stores made scanning billions of rows cheap; unknown questions need dimensions that metrics cannot afford. Do: keep metrics for alerting and long-term trends. Add wide events for investigation. They complement each other.
5. Cost became a design constraint. Was: collect everything. Becoming: telemetry pipelines that filter, sample and route before storage; tiered retention; customer-owned storage (the vendor queries data in your bucket); per-team budgets. Why: telemetry volume grows with traffic × services × detail; budgets do not. Do: IV.04. Measure the bill per team.
6. SLOs replaced thresholds as the alerting unit. Burn-rate alerting (IV.03) is now what the tools generate by default, with OpenSLO as a common definition format.
Part 2 — observability for AI#
7. The unit of work changed. Was: requests per second, CPU percent. Becoming: tokens per second split three ways (input, cached, output), TTFT and per-token latency, KV-cache pressure, goodput (V.01, V.03). Why: a request is no longer a unit of anything.
8. Engines export much more of their internals. A 2024 engine reported a handful of gauges. vLLM 0.30 reports queue time by reason, KV block lifetimes and reuse gaps, prefix-cache hits by source, estimated FLOPs and bytes per GPU, and transfer metrics for disaggregated serving. The Kubernetes inference stack depends on this: the Gateway API Inference Extension’s endpoint picker routes on scraped engine metrics, so those metrics are now part of the control loop, not just the dashboard. Why: schedulers and routers need them; so do you. Do: treat engine metric names as an interface that changes per release, and test your dashboards on upgrade.
9. Power and energy became headline metrics. Was: a facilities concern. Becoming: tokens per watt and per megawatt on vendor launch slides and in capacity plans. Racks draw well over 100 kW and newer ones far more; sites are limited by megawatts, not floor space. Why: power is the binding constraint on how many tokens a site can sell. Do: export the energy counter, compute tokens per joule (V.06), watch throttling.
10. Hardware got heterogeneous, and the rack became the unit.
Fleets now mix GPU generations — Hopper, Blackwell, Rubin from autumn 2026 — with AMD’s
rack-scale systems, cloud accelerators and inference-specific parts such as NVIDIA’s Groq-based
LPX racks. Telemetry differs per vendor and per generation; liquid cooling adds coolant
temperature, flow and leak detection to the list; DRA changes how allocation is counted.
Do: normalize into your own schema (power, energy, temperature, memory, errors, link health)
with accelerator_type as an attribute, and compare like with like.
11. LLM tracing standardized — without stabilizing. The OTel GenAI conventions cover model calls, tools, agents, retrieval, MCP and evaluation events, and every major backend reads them. They are also all still marked Development and moved to their own repository in June 2026 (V.04). Do: adopt, pin, normalize in the Collector.
12. Agents made the task the unit, and traces very large. One task is tens to hundreds of spans with growing contexts. Per-call metrics miss what matters: steps per task, cost per task, loops, tool failure rates, where the time went. Do: roll up per task; sample by outcome; store content by reference.
13. Quality became a signal. Evaluations run continuously and attach to traces; canaries are gated on scores as well as latency (V.07). “LLM observability” tools are being absorbed into general platforms — ClickHouse’s acquisition of Langfuse in January 2026 is the clearest example.
14. Benchmarks became continuous. Public, continuously re-run benchmarks (InferenceX) and MLPerf Inference’s serving-style load generator give reference curves of throughput against latency per accelerator and engine. Do: compare your production tokens-per-GPU against them. A large gap is occupancy or configuration, and now you can tell.
Part 3 — AI for observability#
15. Agents that investigate. Vendors and open-source projects ship “AI SRE” agents that take an alert, query metrics, logs and traces, form hypotheses and write up a probable cause — the loop in IV.05, automated. Datadog’s Bits AI SRE (launched December 2025) is one commercial example; most platforms have an equivalent.
16. Telemetry exposed to models. Observability back ends publish MCP servers so that coding and operations agents can query dashboards, run PromQL, search traces and read profiles as tools.
What to make of these two:
- They work exactly as well as the data is standard, joined and well-named. An agent cannot
follow a trace that breaks at a queue, or guess that
lat_msanddurationare the same thing. Everything in modules I–IV is now also preparation for machine readers. - They are good at the tedious middle of an investigation — slicing by every dimension, correlating a deploy with a graph. They are not accountable for the decision. Keep a human on mitigation, and audit what an agent with production access may do.
- They generate load on your telemetry store, and cost in tokens. Observe the observer: trace the agent’s own model and tool calls with the conventions from V.04.
What is not changing#
- A counter, a histogram, a span and a stack sample are still the four primitives.
- Percentiles still cannot be averaged. Cardinality is still a product.
- Alert on what users feel. Mitigate first.
- The most valuable field is still the one you forgot to record.
Code#
A maturity check to apply to anything new before depending on it: stability, adoption, and whether you have a way out.
// adopt.go — a small decision rule for adopting something from the frontier.
package main
import "fmt"
type Tech struct {
Name string
Stability string // "stable", "beta", "alpha", "development"
MultipleImplementations bool // more than one vendor or project reads/writes it
Reversible bool // can you switch away without re-instrumenting?
OnCriticalPath bool // does an outage or rename break paging?
}
func advise(t Tech) string {
switch {
case t.Stability == "stable":
return "adopt"
case t.OnCriticalPath && !t.Reversible:
return "wait: unstable, on the paging path, and hard to undo"
case t.MultipleImplementations && t.Reversible:
return "adopt, pin the version, normalize names in the Collector"
case t.MultipleImplementations:
return "pilot on one service; write the exit plan first"
default:
return "watch: single implementation and not yet stable"
}
}
func main() {
// Statuses as of 3 October 2026 (see lesson 03 for sources to re-check).
techs := []Tech{
{"OTLP traces / metrics / logs", "stable", true, true, true},
{"Prometheus native histograms", "stable", true, true, true},
{"OTel declarative configuration", "stable", true, true, false},
{"OTel Profiles signal", "alpha", true, true, false},
{"OTel eBPF Instrumentation (OBI)", "beta", true, true, false},
{"OTel GenAI semantic conventions", "development", true, true, false},
{"Alerts keyed on gen_ai.* attribute names", "development", true, false, true},
{"One engine's newest metric names", "development", false, true, false},
}
for _, t := range techs {
fmt.Printf("%-42s %-12s → %s\n", t.Name, t.Stability, advise(t))
}
}Remember this#
- General field: OTel + Prometheus convergence, profiles, eBPF, wide events, cost discipline.
- For AI: tokens and joules as units, deeper engine telemetry, rack-scale and mixed hardware, standard-but-unstable GenAI traces, quality as a signal.
- AI for observability depends on clean, standard, joined telemetry — and must itself be observed.
- The primitives and the pitfalls have not moved.
Try it#
- Run
adopt.go. Add two technologies you are considering and argue with the advice. - Pick one shift and write the opposite case: what would have to be true for it to reverse?
- Take a dashboard you own. Which panels would an investigating agent be unable to interpret without tribal knowledge? Fix the names.
Check yourself#
- Why are wide events and metrics complementary rather than rivals?
- Why did engine metrics become part of the control loop?
- What do AI investigation agents require of your telemetry?