The idea in one minute#
One engine’s metrics tell you about one replica. A platform has many replicas, several models, a router deciding who gets which request, an autoscaler deciding how many replicas exist, and tenants competing for all of it. Fleet observability is about the differences: between replicas, between what the scheduler was asked for and what it delivered, between tenants.
Three things deserve most of the attention: routing quality (did requests land where their cache was?), scaling signals (queue and KV pressure, not GPU utilization), and capacity that is paid for but not serving (pending pods, cold starts, idle GPUs).
An analogy#
Air traffic control. Each pilot watches their own instruments. The controller watches spacing, queues for each runway, which runway is closed, and whether another one should open. No single cockpit shows any of that.
A picture#
flowchart TB
CL["Clients, tenants"] --> GW["Gateway<br/>auth, quotas, per-tenant accounting"]
GW --> RT["Router / endpoint picker<br/>scores replicas by queue, KV usage, prefix match"]
RT --> R1["Replica A<br/>engine metrics"]
RT --> R2["Replica B"]
RT --> R3["Replica C"]
R1 --- G1[("GPU metrics")]
R2 --- G2[("GPU metrics")]
R3 --- G3[("GPU metrics")]
AS["Autoscaler"] -->|"reads waiting, KV usage"| R1
AS -->|"adds or removes replicas"| K8S["Kubernetes<br/>scheduler, DRA, node pools"]
K8S --> R3
class CL neutral
class GW,RT,AS queue
class R1,R2,R3 compute
class G1,G2,G3 memory
class K8S ioHow it really works#
What each component should tell you#
| Component | Metrics | Question answered |
|---|---|---|
| Gateway | Requests, tokens and errors by tenant, model, route; 429s; time spent in the gateway | Who is using what; who is being throttled |
| Router | Decision time; chosen replica; score inputs; prefix-match rate; retries and fallbacks | Is routing helping or scattering caches |
| Engines | Lesson 03, per replica | Which replica is the outlier |
| Autoscaler | Desired vs current replicas; scale events; the signal value it acted on | Is it reacting, and to what |
| Kubernetes | Pod phase, pending time, restarts, OOM kills, node conditions (kube-state-metrics) | Is capacity materializing |
| GPUs | Lesson 02, joined to pods | Is the hardware under each replica healthy |
| Model lifecycle | Load time, download time, readiness time per model version | How long is a cold start, really |
Outliers: compare replicas with each other#
Fleet averages hide the broken one. For every important engine metric, plot the spread:
# Slowest replica's TTFT p95 against the fleet's
max by (model_name) (
histogram_quantile(0.95, sum by (le, pod, model_name) (rate(vllm:time_to_first_token_seconds_bucket[5m]))))
/
histogram_quantile(0.95, sum by (le, model_name) (rate(vllm:time_to_first_token_seconds_bucket[5m])))
# Load imbalance: busiest replica's share of waiting requests
max by (model_name) (vllm:num_requests_waiting) / clamp_min(avg by (model_name) (vllm:num_requests_waiting), 1)A persistent outlier is usually hardware (lesson 02: throttling, a degraded link, a different GPU type), placement (a noisy neighbour, a wrong NUMA node), or routing (one replica attracts the long prompts).
Routing quality#
LLM routing is not round-robin: sending a conversation back to the replica that already holds
its prefix in the KV cache skips most of prefill. In Kubernetes this is the job of the
Gateway API Inference Extension — whose InferencePool API has been stable (v1) since its
1.0 release — and its endpoint picker, which as of 2026 is maintained in the llm-d
project (a CNCF sandbox project since March 2026). The picker scrapes each replica’s queue
depth and KV-cache usage and scores endpoints per request.
What to watch:
- Fleet prefix-cache hit rate, and its split by replica. Good routing raises it; a rollout or a scale-up temporarily lowers it.
- Router decision latency — it is on the path of every request.
- Staleness of the router’s view: it acts on scraped metrics a second or more old.
- Saturation-driven spills: how often the preferred (cache-holding) replica was too busy and the request went elsewhere.
Scaling signals#
| Signal | Use it? | Why |
|---|---|---|
| GPU utilization | No | Pinned near 100% under light load (lesson 02) |
| GPU memory used | No | Pre-allocated |
| Requests per second | Weak | Requests are not a unit of work |
| Requests waiting (queue depth) | Yes | Direct measure of unmet demand |
| KV-cache usage | Yes | The real memory pressure |
| Running requests ÷ configured concurrency | Yes | Headroom |
| TTFT p95 vs SLO | As a guard | A lagging symptom; scale before it moves |
| Tokens/s vs measured capacity | Yes, for planning | Needs a capacity figure from a load test |
In Kubernetes these reach the autoscaler through KEDA or the Prometheus adapter. Two timings decide whether autoscaling works at all, and both must be measured:
time to scale = detection delay + pod scheduling + node provisioning (if no spare GPU)
+ image pull + model download + weight load + warm-upOn a cold node this is commonly several minutes. Record each stage as a metric or a span. If the total exceeds how long your queue can absorb a burst, you need warm spare capacity, not a faster autoscaler.
Capacity that is not serving#
| Waste | How it shows up |
|---|---|
| Pods pending for a GPU | kube_pod_status_phase{phase="Pending"} with a GPU request; time-in-pending histogram |
| GPUs allocated to a pod that is idle | GPU power near idle with a pod label (lesson 02) |
| GPUs not allocated at all | Allocatable minus requested, per node pool |
| Replicas loading a model | Readiness lag after start |
| Fragmentation | Free GPUs spread so that no node can fit the next multi-GPU pod |
| Over-provisioned concurrency | KV usage and running count persistently far below limits |
Allocation ratio (GPUs requested ÷ allocatable) and effective use (tokens produced ÷ tokens the allocated GPUs could produce) are the two fleet efficiency numbers. The first is a scheduler metric; the second needs a capacity figure per model and GPU type.
With Dynamic Resource Allocation — core APIs stable since Kubernetes 1.34, device taints
stable in 1.37, partitionable devices still maturing — GPUs are requested by attribute through
ResourceClaim objects rather than as an opaque count. Your allocation dashboards must follow:
count claims and devices, not just nvidia.com/gpu requests.
Tenants#
Per-tenant accounting belongs at the gateway, where identity is known: tokens in, tokens out, cached tokens, requests, errors, throttles, TTFT. Keep cardinality in check (II.05): metrics for the top tenants and tiers, wide events for everyone. Watch for noisy neighbours — one tenant’s long prompts raising everyone’s ITL — by comparing a tenant’s token share with the fleet’s latency.
Rollouts#
A model or engine upgrade is the most common cause of incidents. Put model_revision,
engine_version and config_hash on every metric and span, and compare canary with baseline
on: TTFT, TPOT, error ratio, finish reasons, output length distribution, prefix hit rate — and
quality (lesson 07). A canary can pass every latency check and still produce worse answers.
Multi-node and disaggregated serving#
When one model spans several GPUs or nodes, or prefill and decode run on separate pools, add:
KV transfer time and bytes, transfer failures, the balance between the prefill and decode pools
(one of them is always the bottleneck), and link health (lesson 02). vLLM exports transfer
metrics for its NIXL connector (vllm:nixl_xfer_time_seconds, vllm:nixl_bytes_transferred,
failure counters); NVIDIA Dynamo and llm-d expose their own.
Code#
Why queue depth is a better scaling signal than a utilization gauge: simulate both autoscalers on the same traffic.
// autoscale.go — scale on GPU utilization vs on unmet demand, with a slow cold start.
package main
import (
"fmt"
"math"
)
func main() {
const (
perReplica = 50.0 // requests a replica can serve per tick within the SLO
coldStart = 6 // ticks until a new replica serves traffic
maxReplicas = 12
ticks = 60
)
demand := func(t int) float64 {
if t >= 15 && t < 40 {
return 260 // a burst
}
return 40
}
policies := []struct {
name string
want func(util, served, waiting float64, replicas int) int
}{
{"GPU utilization (up >80%, down <30%)", func(util, _, _ float64, r int) int {
switch {
case util > 0.8:
return r + 1
case util < 0.3:
return r - 1
}
return r
}},
{"served + waiting requests", func(_, served, waiting float64, _ int) int {
return int(math.Ceil((served + waiting) / perReplica))
}},
}
fmt.Println("policy replica-ticks paid max replicas ticks with a queue peak queue")
for _, p := range policies {
ready := 1
var starting []int // tick at which each new replica becomes ready
queue, peakQueue := 0.0, 0.0
lateTicks, paid, most := 0, 0, 1
for t := 0; t < ticks; t++ {
for len(starting) > 0 && starting[0] <= t {
ready++
starting = starting[1:]
}
queue += demand(t)
served := math.Min(float64(ready)*perReplica, queue)
queue -= served
// A GPU serving even one request reports ~100% "utilization".
util := 0.0
if served > 0 {
util = 1.0
}
target := p.want(util, served, queue, ready+len(starting))
target = max(1, min(target, maxReplicas))
for ready+len(starting) < target {
starting = append(starting, t+coldStart)
}
for ready+len(starting) > target {
if len(starting) > 0 {
starting = starting[:len(starting)-1]
} else {
ready--
}
}
if queue > 0 {
lateTicks++
}
peakQueue = math.Max(peakQueue, queue)
paid += ready + len(starting)
most = max(most, ready+len(starting))
}
fmt.Printf("%-38s %18d %12d %18d %10.0f\n", p.name, paid, most, lateTicks, peakQueue)
}
fmt.Println("\nUtilization reads 100% whenever anything is served, so that policy climbs to the")
fmt.Println("maximum and stays there: it is not autoscaling, it is paying for peak all day.")
fmt.Println("The demand-based policy follows the load; the queue it shows during the burst is")
fmt.Println("the cold start, and only warm capacity or a faster start removes it.")
}Remember this#
- Fleet observability is about differences: replica vs replica, desired vs actual, tenant vs tenant.
- Scale on waiting requests and KV-cache usage. Measure every stage of a cold start.
- Routing quality shows up as prefix-cache hit rate.
- Track capacity that is paid for and not serving: pending, idle, loading, fragmented.
- Version labels on everything; compare canary and baseline on latency and quality.
Try it#
- Run
autoscale.go. SetcoldStartto 1, then to 15. What changes for each policy, and what does that say about warm capacity? - Write PromQL for: pods pending with a GPU request for more than five minutes; GPUs allocated per node pool; prefix-cache hit rate per replica.
- List every stage of a cold start in a system you know and how you would time each.
Check yourself#
- Why are GPU utilization and GPU memory poor autoscaling signals?
- What does a falling fleet prefix-cache hit rate after a scale-up tell you?
- Name four kinds of capacity that is paid for but not serving.