PidokuInfra

What to Watch

Expert 45 min Difficulty 2/5 Topic 03 of 04

Prerequisites V, 01, 02

The idea in one minute#

Two lists. First, the metrics to watch on an AI serving platform — about two dozen, layer by layer, each with the direction that means trouble. Second, the projects and feeds to watch, with their status on 3 October 2026 and where to re-check it.

Use the first list as a checklist against your own dashboards. Use the second on a schedule: once a quarter, re-read the statuses and update what you depend on.

A picture#

flowchart LR
  U["User experience<br/>TTFT, TPOT, errors, goodput"] --> E["Engine<br/>waiting, KV usage, preemptions, cache hits"]
  E --> H["Hardware<br/>power, throttling, errors, links"]
  E --> P["Platform<br/>pending, cold start, outliers, routing"]
  P --> C["Cost<br/>$ per M tokens, occupancy, tokens per joule"]
  U --> Q["Quality<br/>finish reasons, parse failures, eval scores"]
  class U queue
  class E memory
  class H io
  class P compute
  class C neutral
  class Q queue

How it really works#

The metrics to watch#

Start from the top: only descend when a row above is unhealthy.

User experience — alert on these

MetricTrouble looks likeLesson
TTFT p95 / p99 per model and workload classAbove SLO; a step change after a deployV.03
TPOT p95, and ITL p99TPOT rising with load; ITL spikes = stallsV.03
Error ratio, by error.typeBurn-rate alert (IV.03)IV.03
Goodput: share of requests meeting all SLOsFalling while throughput still rises = past capacityV.03
429 / shed rate per tenantRising for paying tenantsV.05

Engine — the first place to look

MetricTrouble looks likeLesson
Requests waitingSustained above zero for interactive trafficV.03
KV-cache usagePersistently above ~0.9V.03
PreemptionsAny sustained rateV.03
Prefix-cache hit rateA drop after a deploy, scale-up or prompt changeV.03, V.05
Output and prompt tokens/s per replicaOne replica far from its peersV.05
Queue time share of total latencyGrowingV.03
Tokens per engine step (batch fill)Low while requests wait = misconfigurationV.03

Hardware — is the device healthy and working

MetricTrouble looks likeLesson
Power draw vs limit; energy counterA GPU far below its siblings; drawing idle power while allocatedV.02
Thermal / power throttling; clock frequencyAny throttling time; clocks below peersV.02
Xid eventsAny hardware Xid (48, 63/64, 74, 79, 94/95)V.02
ECC and row-remap countersRising on one deviceV.02
NVLink and PCIe error countersNon-zero growthV.02
Temperature (GPU and memory); coolant where liquid-cooledTrend upwardV.02
Tensor-pipe and memory-interface activityLow while “utilization” reads 100%V.02

Platform — is capacity arriving and being shared fairly

MetricTrouble looks likeLesson
Pods pending for a GPU; time in pendingMinutesV.05
Cold-start duration, by stageLonger than a burst your queue can absorbV.05
Desired vs ready replicasA persistent gapV.05
Slowest replica vs fleet (TTFT, tokens/s)Ratio well above 1V.05
Router decision latency; spill rateRisingV.05
Tokens by tenant vs quotaOne tenant crowding the restV.05

Cost and efficiency — review weekly

MetricTrouble looks likeLesson
$ per million output tokens per modelAbove target or risingV.06
Occupancy; allocation ratioLow occupancy with high allocation = paying for idleV.06
Tokens per joule; site power vs capFalling; power near the capV.06
Idle GPU-hoursAny large numberV.06
Telemetry volume and series count per teamA step upII.05, IV.04

Quality — gate rollouts on these

MetricTrouble looks likeLesson
Finish-reason mix; output length distributionA shift after any changeV.07
Structured-output and tool-call failuresRisingV.07
Corrupted / NaN outputsAnyV.07
Evaluation scores, canary vs baselineA significant dropV.07
Which model actually answered (fallback share)Rising silentlyV.07

Projects and standards: status on 3 October 2026#

ThingStatusWhy you careRe-check at
OpenTelemetry traces, metrics, logsStableThe default instrumentationopentelemetry.io/status
OTel declarative configurationStable (2026)One config file for SDKsopentelemetry.io/blog
OTel ProfilesPublic alpha (March 2026)The fourth signalopentelemetry.io/blog
OTel eBPF Instrumentation (OBI)Beta (April 2026); 1.0 targetedZero-code traces and metricsThe OBI repository’s releases
OTel GenAI semantic conventionsAll Development; own repository since June 2026Standard LLM, agent and tool spansopen-telemetry/semantic-conventions-genai
Prometheus3.x series; 3.15 released 24 September 2026; native histograms stable since 3.8; OTLP ingestion built inThe metrics baselineprometheus.io/blog, GitHub releases
vLLM metricsv0.30 (22 September 2026)Names change between releasesdocs.vllm.ai → Usage → Metrics
Gateway API Inference ExtensionInferencePool v1 stable; endpoint picker now developed in llm-dRouting on engine metricsThe project’s releases page
llm-dCNCF sandbox since March 2026Kubernetes-native distributed inference, with tracingllm-d.ai/blog
Kubernetes DRACore stable since 1.34; device taints stable in 1.37 (August 2026); partitionable devices still pre-GAHow GPUs are requested and countedkubernetes.io/blog release posts
NVIDIA DCGM / dcgm-exporterMature; field list grows with each GPU generationHardware telemetryThe exporter’s default-counters.csv
LLM observability toolsConsolidating into general platforms (Langfuse → ClickHouse, January 2026)Where prompts and evals liveEach project’s changelog
InferenceX; MLPerf InferenceContinuously re-run; MLPerf v6.0 results April 2026Reference curves to compare againstTheir sites

Feeds worth a regular look#

FeedFor
OpenTelemetry blog and the status pageSignal and convention stability
Prometheus blog and release notesQuery and storage changes
CNCF and Kubernetes blogs (release posts)DRA, Gateway API, scheduling
vLLM blog and release notes; SGLang releasesNew engine metrics, tracing, disaggregation
llm-d blogRouting, tracing and scheduling in practice
NVIDIA developer blog and GPU Operator / DCGM release notesNew hardware fields, health tooling
Grafana Labs and ClickHouse engineering blogsStorage and query engines for telemetry
KubeCon + CloudNativeCon and SREcon talksWhat operators actually run

Vendor blogs report their best case. Read them for mechanisms, and verify numbers yourself.

A quarterly review, in six questions#

  1. Did any metric or attribute we alert on get renamed upstream?
  2. Is anything we depend on still below “stable”? Did it move?
  3. Which dashboards did nobody open? Delete them.
  4. Which pages needed no action? Fix or delete them.
  5. What did the last incident lack? Was it added?
  6. What did telemetry cost, per team, against the infrastructure it watches?

Evaluating a new tool or signal#

Ask, in order: which job does it do (lesson 01)? Does it speak OTLP / PromQL, or its own format? What is its stability label? Can it be joined to what you have — same resource attributes? What does it cost at your volume? What is the exit?

Code#

Turn the first list into something executable: a health roll-up that names the first unhealthy layer, which is where an investigation should start.

Go
// watch.go — evaluate a snapshot of the key metrics and point at the first unhealthy layer.
package main

import "fmt"

type Check struct {
	Layer, Name string
	Value       float64
	Bad         func(float64) bool
	Unit        string
}

func main() {
	gt := func(limit float64) func(float64) bool { return func(v float64) bool { return v > limit } }
	lt := func(limit float64) func(float64) bool { return func(v float64) bool { return v < limit } }

	// A snapshot. In production each value is the result of one query.
	checks := []Check{
		{"1 user", "TTFT p95", 0.92, gt(0.5), "s"},
		{"1 user", "TPOT p95", 0.031, gt(0.040), "s"},
		{"1 user", "goodput", 0.91, lt(0.99), "ratio"},
		{"2 engine", "requests waiting (max replica)", 2, gt(8), ""},
		{"2 engine", "KV-cache usage", 0.71, gt(0.90), "ratio"},
		{"2 engine", "preemptions per min", 0, gt(0), ""},
		{"2 engine", "prefix-cache hit rate", 0.22, lt(0.50), "ratio"},
		{"3 hardware", "throttling seconds per min", 0, gt(0), "s"},
		{"3 hardware", "hardware Xid events (1 h)", 0, gt(0), ""},
		{"4 platform", "pods pending for a GPU", 0, gt(0), ""},
		{"4 platform", "slowest replica TTFT vs fleet", 1.2, gt(2), "x"},
		{"5 quality", "finish_reason=length share", 0.012, gt(0.03), "ratio"},
	}

	first := ""
	for _, c := range checks {
		status := "ok"
		if c.Bad(c.Value) {
			status = "BAD"
			if first == "" || (c.Layer > "1 user" && first == "1 user") {
				first = c.Layer
			}
		}
		fmt.Printf("%-11s %-34s %8.3f %-6s %s\n", c.Layer, c.Name, c.Value, c.Unit, status)
	}
	fmt.Printf("\nStart investigating at: %s\n", first)
	fmt.Println("Here: users see slow first tokens, the queue is short, the hardware is healthy,")
	fmt.Println("and the prefix cache stopped hitting — look for a prompt or routing change, not for capacity.")
}

Remember this#

  • About two dozen metrics cover an AI serving platform; read them top-down from user experience.
  • Each has a direction that means trouble; write it next to the panel.
  • Standards and tools carry stability labels. Re-check them quarterly, from primary sources.
  • Evaluate anything new by job, standard, stability, joinability, cost and exit.

Try it#

  1. Run watch.go. Change the snapshot so the cause is hardware throttling. Which values move together?
  2. Tick off the metric tables against your own dashboards. What is missing from each layer?
  3. Do the six-question quarterly review for a system you know.

Check yourself#

  1. In what order do you read the layers during an incident, and why?
  2. Which three items in the status table were not yet stable in October 2026?
  3. What six questions do you ask of a new tool?

↑↓ navigate↵ openesc close