The idea in one minute#
Two lists. First, the metrics to watch on an AI serving platform — about two dozen, layer by layer, each with the direction that means trouble. Second, the projects and feeds to watch, with their status on 3 October 2026 and where to re-check it.
Use the first list as a checklist against your own dashboards. Use the second on a schedule: once a quarter, re-read the statuses and update what you depend on.
A picture#
flowchart LR U["User experience<br/>TTFT, TPOT, errors, goodput"] --> E["Engine<br/>waiting, KV usage, preemptions, cache hits"] E --> H["Hardware<br/>power, throttling, errors, links"] E --> P["Platform<br/>pending, cold start, outliers, routing"] P --> C["Cost<br/>$ per M tokens, occupancy, tokens per joule"] U --> Q["Quality<br/>finish reasons, parse failures, eval scores"] class U queue class E memory class H io class P compute class C neutral class Q queue
How it really works#
The metrics to watch#
Start from the top: only descend when a row above is unhealthy.
User experience — alert on these
| Metric | Trouble looks like | Lesson |
|---|---|---|
| TTFT p95 / p99 per model and workload class | Above SLO; a step change after a deploy | V.03 |
| TPOT p95, and ITL p99 | TPOT rising with load; ITL spikes = stalls | V.03 |
Error ratio, by error.type | Burn-rate alert (IV.03) | IV.03 |
| Goodput: share of requests meeting all SLOs | Falling while throughput still rises = past capacity | V.03 |
| 429 / shed rate per tenant | Rising for paying tenants | V.05 |
Engine — the first place to look
| Metric | Trouble looks like | Lesson |
|---|---|---|
| Requests waiting | Sustained above zero for interactive traffic | V.03 |
| KV-cache usage | Persistently above ~0.9 | V.03 |
| Preemptions | Any sustained rate | V.03 |
| Prefix-cache hit rate | A drop after a deploy, scale-up or prompt change | V.03, V.05 |
| Output and prompt tokens/s per replica | One replica far from its peers | V.05 |
| Queue time share of total latency | Growing | V.03 |
| Tokens per engine step (batch fill) | Low while requests wait = misconfiguration | V.03 |
Hardware — is the device healthy and working
| Metric | Trouble looks like | Lesson |
|---|---|---|
| Power draw vs limit; energy counter | A GPU far below its siblings; drawing idle power while allocated | V.02 |
| Thermal / power throttling; clock frequency | Any throttling time; clocks below peers | V.02 |
| Xid events | Any hardware Xid (48, 63/64, 74, 79, 94/95) | V.02 |
| ECC and row-remap counters | Rising on one device | V.02 |
| NVLink and PCIe error counters | Non-zero growth | V.02 |
| Temperature (GPU and memory); coolant where liquid-cooled | Trend upward | V.02 |
| Tensor-pipe and memory-interface activity | Low while “utilization” reads 100% | V.02 |
Platform — is capacity arriving and being shared fairly
| Metric | Trouble looks like | Lesson |
|---|---|---|
| Pods pending for a GPU; time in pending | Minutes | V.05 |
| Cold-start duration, by stage | Longer than a burst your queue can absorb | V.05 |
| Desired vs ready replicas | A persistent gap | V.05 |
| Slowest replica vs fleet (TTFT, tokens/s) | Ratio well above 1 | V.05 |
| Router decision latency; spill rate | Rising | V.05 |
| Tokens by tenant vs quota | One tenant crowding the rest | V.05 |
Cost and efficiency — review weekly
| Metric | Trouble looks like | Lesson |
|---|---|---|
| $ per million output tokens per model | Above target or rising | V.06 |
| Occupancy; allocation ratio | Low occupancy with high allocation = paying for idle | V.06 |
| Tokens per joule; site power vs cap | Falling; power near the cap | V.06 |
| Idle GPU-hours | Any large number | V.06 |
| Telemetry volume and series count per team | A step up | II.05, IV.04 |
Quality — gate rollouts on these
| Metric | Trouble looks like | Lesson |
|---|---|---|
| Finish-reason mix; output length distribution | A shift after any change | V.07 |
| Structured-output and tool-call failures | Rising | V.07 |
| Corrupted / NaN outputs | Any | V.07 |
| Evaluation scores, canary vs baseline | A significant drop | V.07 |
| Which model actually answered (fallback share) | Rising silently | V.07 |
Projects and standards: status on 3 October 2026#
| Thing | Status | Why you care | Re-check at |
|---|---|---|---|
| OpenTelemetry traces, metrics, logs | Stable | The default instrumentation | opentelemetry.io/status |
| OTel declarative configuration | Stable (2026) | One config file for SDKs | opentelemetry.io/blog |
| OTel Profiles | Public alpha (March 2026) | The fourth signal | opentelemetry.io/blog |
| OTel eBPF Instrumentation (OBI) | Beta (April 2026); 1.0 targeted | Zero-code traces and metrics | The OBI repository’s releases |
| OTel GenAI semantic conventions | All Development; own repository since June 2026 | Standard LLM, agent and tool spans | open-telemetry/semantic-conventions-genai |
| Prometheus | 3.x series; 3.15 released 24 September 2026; native histograms stable since 3.8; OTLP ingestion built in | The metrics baseline | prometheus.io/blog, GitHub releases |
| vLLM metrics | v0.30 (22 September 2026) | Names change between releases | docs.vllm.ai → Usage → Metrics |
| Gateway API Inference Extension | InferencePool v1 stable; endpoint picker now developed in llm-d | Routing on engine metrics | The project’s releases page |
| llm-d | CNCF sandbox since March 2026 | Kubernetes-native distributed inference, with tracing | llm-d.ai/blog |
| Kubernetes DRA | Core stable since 1.34; device taints stable in 1.37 (August 2026); partitionable devices still pre-GA | How GPUs are requested and counted | kubernetes.io/blog release posts |
| NVIDIA DCGM / dcgm-exporter | Mature; field list grows with each GPU generation | Hardware telemetry | The exporter’s default-counters.csv |
| LLM observability tools | Consolidating into general platforms (Langfuse → ClickHouse, January 2026) | Where prompts and evals live | Each project’s changelog |
| InferenceX; MLPerf Inference | Continuously re-run; MLPerf v6.0 results April 2026 | Reference curves to compare against | Their sites |
Feeds worth a regular look#
| Feed | For |
|---|---|
| OpenTelemetry blog and the status page | Signal and convention stability |
| Prometheus blog and release notes | Query and storage changes |
| CNCF and Kubernetes blogs (release posts) | DRA, Gateway API, scheduling |
| vLLM blog and release notes; SGLang releases | New engine metrics, tracing, disaggregation |
| llm-d blog | Routing, tracing and scheduling in practice |
| NVIDIA developer blog and GPU Operator / DCGM release notes | New hardware fields, health tooling |
| Grafana Labs and ClickHouse engineering blogs | Storage and query engines for telemetry |
| KubeCon + CloudNativeCon and SREcon talks | What operators actually run |
Vendor blogs report their best case. Read them for mechanisms, and verify numbers yourself.
A quarterly review, in six questions#
- Did any metric or attribute we alert on get renamed upstream?
- Is anything we depend on still below “stable”? Did it move?
- Which dashboards did nobody open? Delete them.
- Which pages needed no action? Fix or delete them.
- What did the last incident lack? Was it added?
- What did telemetry cost, per team, against the infrastructure it watches?
Evaluating a new tool or signal#
Ask, in order: which job does it do (lesson 01)? Does it speak OTLP / PromQL, or its own format? What is its stability label? Can it be joined to what you have — same resource attributes? What does it cost at your volume? What is the exit?
Code#
Turn the first list into something executable: a health roll-up that names the first unhealthy layer, which is where an investigation should start.
// watch.go — evaluate a snapshot of the key metrics and point at the first unhealthy layer.
package main
import "fmt"
type Check struct {
Layer, Name string
Value float64
Bad func(float64) bool
Unit string
}
func main() {
gt := func(limit float64) func(float64) bool { return func(v float64) bool { return v > limit } }
lt := func(limit float64) func(float64) bool { return func(v float64) bool { return v < limit } }
// A snapshot. In production each value is the result of one query.
checks := []Check{
{"1 user", "TTFT p95", 0.92, gt(0.5), "s"},
{"1 user", "TPOT p95", 0.031, gt(0.040), "s"},
{"1 user", "goodput", 0.91, lt(0.99), "ratio"},
{"2 engine", "requests waiting (max replica)", 2, gt(8), ""},
{"2 engine", "KV-cache usage", 0.71, gt(0.90), "ratio"},
{"2 engine", "preemptions per min", 0, gt(0), ""},
{"2 engine", "prefix-cache hit rate", 0.22, lt(0.50), "ratio"},
{"3 hardware", "throttling seconds per min", 0, gt(0), "s"},
{"3 hardware", "hardware Xid events (1 h)", 0, gt(0), ""},
{"4 platform", "pods pending for a GPU", 0, gt(0), ""},
{"4 platform", "slowest replica TTFT vs fleet", 1.2, gt(2), "x"},
{"5 quality", "finish_reason=length share", 0.012, gt(0.03), "ratio"},
}
first := ""
for _, c := range checks {
status := "ok"
if c.Bad(c.Value) {
status = "BAD"
if first == "" || (c.Layer > "1 user" && first == "1 user") {
first = c.Layer
}
}
fmt.Printf("%-11s %-34s %8.3f %-6s %s\n", c.Layer, c.Name, c.Value, c.Unit, status)
}
fmt.Printf("\nStart investigating at: %s\n", first)
fmt.Println("Here: users see slow first tokens, the queue is short, the hardware is healthy,")
fmt.Println("and the prefix cache stopped hitting — look for a prompt or routing change, not for capacity.")
}Remember this#
- About two dozen metrics cover an AI serving platform; read them top-down from user experience.
- Each has a direction that means trouble; write it next to the panel.
- Standards and tools carry stability labels. Re-check them quarterly, from primary sources.
- Evaluate anything new by job, standard, stability, joinability, cost and exit.
Try it#
- Run
watch.go. Change the snapshot so the cause is hardware throttling. Which values move together? - Tick off the metric tables against your own dashboards. What is missing from each layer?
- Do the six-question quarterly review for a system you know.
Check yourself#
- In what order do you read the layers during an incident, and why?
- Which three items in the status table were not yet stable in October 2026?
- What six questions do you ask of a new tool?