1. What to measure#
The complete metric set for an LLM inference service, organized by what each answers.
GOLDEN SIGNALS (what users experience)
ttft_seconds histogram, labeled by prompt-length bucket
inter_token_latency_seconds histogram
e2e_latency_seconds histogram, labeled by output-length bucket
request_errors_total counter, by type
requests_total counter, by status
CAPACITY (what limits you)
kv_cache_usage_ratio gauge ← THE capacity signal
num_requests_running gauge
num_requests_waiting gauge
queue_wait_seconds histogram ← THE overload signal
preemptions_total counter
admission_rejections_total counter, by reason
EFFICIENCY (what you're paying for)
avg_running_batch_size gauge ← the most under-collected metric
prefill_tokens_total counter
decode_tokens_total counter
prefix_cache_hit_tokens_total counter
prefix_cache_query_tokens_total counter
gpu_dram_active gauge (from DCGM)
gpu_sm_active gauge (from DCGM)
tokens_generated_after_abort counter ← pure waste
BUSINESS
cost_per_million_tokens computed
tokens_by_tenant_total counter, by tenant and model
goodput_requests_total counter (met both SLOs)
HEALTH
model_load_duration_seconds histogram
gpu_memory_used_bytes gauge
gpu_power_watts gauge (throttling indicator)
gpu_clock_throttle_reasons gauge
nccl_errors_total counterThe three most under-collected and most valuable: avg_running_batch_size,
queue_wait_seconds, and tokens_generated_after_abort. Together they tell you whether you’re
traffic-limited, overloaded, or wasting capacity.
2. Labels — get these right#
GOOD LABELS (low cardinality, high value)
model "llama-3-70b"
model_version "v1.5.0-fp8" ← include the full serving config identity
tenant "team-search" (if you have < ~100 tenants)
prompt_bucket "0-512", "512-2k", "2k-8k", "8k-32k", "32k+"
status "success", "error", "cancelled"
error_type "oom", "timeout", "rejected", "upstream"
BAD LABELS (cardinality explosion)
request_id ✗ unbounded
user_id ✗ potentially millions
prompt_hash ✗ unbounded
exact_length ✗ use bucketsCardinality kills Prometheus. Each unique label combination is a separate time series;
model × version × tenant × bucket × status at 3×2×50×5×3 = 4,500 series is fine, adding
user_id makes it millions.
Put high-cardinality data in logs or traces, not metrics.
3. The dashboard hierarchy#
Dashboard 1 — Service health (the one on the wall)#
┌─────────────────────┬─────────────────────┬─────────────────────┐
│ TTFT p50/p95/p99 │ ITL p50/p95/p99 │ Error rate │
│ (by prompt bucket) │ │ (by type) │
├─────────────────────┼─────────────────────┼─────────────────────┤
│ Requests/sec │ Output tokens/sec │ Goodput % │
├─────────────────────┼─────────────────────┼─────────────────────┤
│ Queue wait p95 │ KV cache usage % │ Running batch size │
│ ↑ overload signal │ ↑ capacity signal │ ↑ efficiency signal │
└─────────────────────┴─────────────────────┴─────────────────────┘Those bottom three panels are the LLM-specific ones, and they’re what distinguish a useful dashboard from a generic web-service dashboard.
Dashboard 2 — Efficiency and cost#
cost per M output tokens (computed, trended)
average running batch / max_num_seqs
DCGM DRAM_ACTIVE
prefix cache hit rate
preemption rate
aborted token fraction
tokens by tenant (stacked)
prefill vs decode time splitDashboard 3 — Fleet#
per-replica: KV usage, batch size, queue depth (heatmap across replicas)
→ imbalance shows immediately as a heatmap with hot rows
GPU memory, power, temperature, throttle reasons
model load durations
restart countsThe per-replica heatmap is the fastest way to spot a routing problem. If one replica is at 95% KV and others at 20%, your load balancer is wrong (Section VIII.06).
4. The alerts that matter#
groups:
- name: llm-inference
rules:
# SLO violations — page
- alert: TTFTSLOViolation
expr: |
histogram_quantile(0.95,
sum(rate(ttft_seconds_bucket{prompt_bucket=~"0-512|512-2k"}[5m])) by (le, model)
) > 1.0
for: 10m
labels: {severity: page}
- alert: ITLSLOViolation
expr: |
histogram_quantile(0.95, sum(rate(itl_seconds_bucket[5m])) by (le, model)) > 0.08
for: 10m
labels: {severity: page}
# Capacity — page before it becomes an SLO violation
- alert: QueueWaitHigh
expr: |
histogram_quantile(0.95, sum(rate(queue_wait_seconds_bucket[5m])) by (le)) > 2
for: 5m
labels: {severity: page}
- alert: KVCacheNearExhaustion
expr: avg_over_time(kv_cache_usage_ratio[5m]) > 0.92
for: 10m
labels: {severity: warn}
# Efficiency — ticket, not page
- alert: LowBatchUtilization
expr: avg_over_time(avg_running_batch_size[30m]) / max_num_seqs < 0.25
for: 30m
labels: {severity: ticket}
annotations:
summary: "Running at {{ $value }} of capacity — consolidate or investigate routing"
- alert: PreemptionThrashing
expr: rate(preemptions_total[5m]) / rate(engine_steps_total[5m]) > 0.02
for: 10m
labels: {severity: warn}
- alert: HighAbortedTokenFraction
expr: |
rate(tokens_generated_after_abort[15m]) / rate(decode_tokens_total[15m]) > 0.15
for: 30m
labels: {severity: ticket}
# Health
- alert: GPUThrottling
# newer DCGM releases name this field DCGM_FI_DEV_CLOCKS_EVENT_REASONS, and it is not in
# dcgm-exporter's default counter set — check your exporter's metric list
expr: DCGM_FI_DEV_CLOCK_THROTTLE_REASONS != 0
for: 15m
labels: {severity: warn}
- alert: ModelLoadSlow
expr: histogram_quantile(0.9, rate(model_load_duration_seconds_bucket[1h])) > 300
labels: {severity: ticket}Note the alert on LowBatchUtilization. Most teams alert on things being slow; alerting on
being inefficient is what catches the “16 under-utilized replicas” problem from Section X.04.
5. Deriving cost per token#
# Cost per million output tokens
(
sum(gpu_count) * hourly_gpu_cost / 3600
)
/
(
sum(rate(decode_tokens_total[5m]))
) * 1e6With hourly_gpu_cost as a recording rule or a static config value. Track this as a first-class
metric — it’s the number that connects engineering work to the business.
Also useful:
# Prefix cache hit rate
sum(rate(prefix_cache_hit_tokens_total[5m]))
/ sum(rate(prefix_cache_query_tokens_total[5m]))
# Prefill share of GPU time
sum(rate(prefill_time_seconds_total[5m]))
/ (sum(rate(prefill_time_seconds_total[5m])) + sum(rate(decode_time_seconds_total[5m])))
# Effective utilization (Section X.04)
avg(gpu_busy_ratio)
* (avg(avg_running_batch_size) / max_num_seqs)
* (1 - rate(tokens_generated_after_abort[5m]) / rate(decode_tokens_total[5m]))6. Tracing#
Metrics tell you what; traces tell you which request and why.
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
async def handle_request(req):
with tracer.start_as_current_span("inference.request") as span:
span.set_attribute("gen_ai.request.model", req.model)
span.set_attribute("gen_ai.request.max_tokens", req.max_tokens)
span.set_attribute("gen_ai.usage.input_tokens", n_prompt)
with tracer.start_as_current_span("tokenize"): ...
with tracer.start_as_current_span("queue"): ... # ← the interesting span
with tracer.start_as_current_span("prefill"): ...
with tracer.start_as_current_span("decode"):
span.set_attribute("batch_size_at_start", bs)
span.set_attribute("gen_ai.usage.output_tokens", n_out)
span.set_attribute("gen_ai.response.finish_reason", reason)OpenTelemetry has semantic conventions for GenAI (gen_ai.* attributes) — use them, so
your traces are interpretable by standard tooling.
Sample intelligently: trace 100% of errors and slow requests, 0.1% of normal ones. Tracing every request at 1,000 req/s is expensive and unnecessary.
7. Structured logging#
{
"ts": "2026-03-14T10:23:45.123Z",
"level": "info",
"event": "request_complete",
"request_id": "req_abc123",
"model": "llama-3-70b",
"model_version": "v1.5.0-fp8",
"tenant": "team-search",
"prompt_tokens": 1843,
"output_tokens": 200,
"cached_prompt_tokens": 1536,
"ttft_ms": 234,
"queue_wait_ms": 12,
"prefill_ms": 187,
"itl_p50_ms": 21.4,
"itl_p99_ms": 48.2,
"e2e_ms": 4512,
"finish_reason": "stop",
"batch_size_at_admission": 47,
"preempted": false
}One structured record per request enables analyses metrics cannot:
-- Which tenant is driving the p99?
SELECT tenant, percentile_cont(0.99) WITHIN GROUP (ORDER BY ttft_ms)
FROM requests WHERE ts > now() - interval '1 hour' GROUP BY tenant;
-- Does prefix caching help this tenant?
SELECT tenant, avg(cached_prompt_tokens::float / prompt_tokens)
FROM requests GROUP BY tenant;
-- What's the actual output length distribution?
SELECT width_bucket(output_tokens, 0, 4096, 32), count(*)
FROM requests GROUP BY 1 ORDER BY 1;That last query is what you need for capacity planning (Section XI.02) and for building a realistic benchmark (Section X.07). You cannot get it from metrics.
8. Production implications#
- Emit all three: metrics, traces, logs. They answer different questions.
- Get the LLM-specific metrics right: KV usage, running batch, queue wait, prefix hit rate, aborted tokens. Generic APM tools won’t provide them.
- Deploy
dcgm-exporterfor the GPU-side metrics (Section X.04). - Keep cardinality bounded. No user IDs in labels.
- Bucket TTFT by prompt length. Unbucketed percentiles are meaningless.
- Alert on efficiency, not just latency.
- Retain structured logs long enough for capacity analysis — at least 30 days.
- Include
model_versioneverywhere, so you can correlate regressions with deployments.
9. Common mistakes#
Only golden signals. You need capacity and efficiency metrics too.
Unbucketed TTFT. Dominated by the length tail.
Cardinality explosion. From user IDs or exact lengths in labels.
Alerting on GPU utilization. It’s always 100% (Section X.04).
Alerting on queue depth instead of queue wait. Depth without cost is meaningless.
No efficiency alerts. You’ll never notice you’re running at 20% capacity.
Tracing everything. Expensive and unnecessary; sample.
Not logging per-request shape. You lose the ability to do capacity analysis.
10. Hands-on exercise#
A. Instrument fully. Add all the metrics from section 1 to a real server. Verify each appears in Prometheus with sensible values.
B. Build the dashboards. Create the three dashboards from section 3. Load-test the system and watch which panels move.
C. Test the alerts. Implement the alerts from section 4. Trigger each one deliberately (overload for queue wait, undersize for preemption, split traffic for low batch). Verify each fires.
D. Compute cost. Implement the cost-per-token query. Trend it over a day. Does it match your Section V.15 calculation?
E. Log analysis. Emit structured per-request logs. Write the three queries from section 7. Use the output-length distribution to build a realistic benchmark (Section X.07).
F. Find the imbalance. Build the per-replica heatmap. Deliberately misconfigure the load balancer and confirm the heatmap reveals it.
11. Interview questions#
- What metrics would you put on an LLM inference dashboard, and why those?
- Why must TTFT be bucketed by prompt length?
- What are the three most under-collected LLM serving metrics?
- Why can’t you alert on GPU utilization?
- How would you compute cost per million tokens in Prometheus?
- What goes in metrics vs traces vs logs?
- What per-request data do you need for capacity planning, and where do you store it?
12. Further reading#
- [REFERENCE] Prometheus best practices; histogram vs summary
- [REFERENCE] OpenTelemetry GenAI semantic conventions
- [REFERENCE] NVIDIA DCGM exporter
- [REFERENCE] vLLM’s
/metricsendpoint — a good reference metric set. Names change between releases (KV-cache usage isvllm:kv_cache_usage_percin current versions); read your version’s metrics page - Go deeper: the Observability Engineering path covers this subject as a full course — histograms and cardinality (II), SLO alerting (IV.03), GPU telemetry field by field (V.02) and engine metrics (V.03)
- [FUNDAMENTAL] Google SRE Book, “Monitoring Distributed Systems”
- Next: 11 — Case studies