PidokuInfra

Metrics, Prometheus, and Dashboards

Advanced Intermediate 1h 15m Difficulty 3/5 Topic 10 of 11

Prerequisites 04, VIII.03


1. What to measure#

The complete metric set for an LLM inference service, organized by what each answers.

GOLDEN SIGNALS (what users experience)
  ttft_seconds                     histogram, labeled by prompt-length bucket
  inter_token_latency_seconds      histogram
  e2e_latency_seconds              histogram, labeled by output-length bucket
  request_errors_total             counter, by type
  requests_total                   counter, by status

CAPACITY (what limits you)
  kv_cache_usage_ratio             gauge  ← THE capacity signal
  num_requests_running             gauge
  num_requests_waiting             gauge
  queue_wait_seconds               histogram  ← THE overload signal
  preemptions_total                counter
  admission_rejections_total       counter, by reason

EFFICIENCY (what you're paying for)
  avg_running_batch_size           gauge  ← the most under-collected metric
  prefill_tokens_total             counter
  decode_tokens_total              counter
  prefix_cache_hit_tokens_total    counter
  prefix_cache_query_tokens_total  counter
  gpu_dram_active                  gauge (from DCGM)
  gpu_sm_active                    gauge (from DCGM)
  tokens_generated_after_abort     counter  ← pure waste

BUSINESS
  cost_per_million_tokens          computed
  tokens_by_tenant_total           counter, by tenant and model
  goodput_requests_total           counter (met both SLOs)

HEALTH
  model_load_duration_seconds      histogram
  gpu_memory_used_bytes            gauge
  gpu_power_watts                  gauge (throttling indicator)
  gpu_clock_throttle_reasons       gauge
  nccl_errors_total                counter

The three most under-collected and most valuable: avg_running_batch_size, queue_wait_seconds, and tokens_generated_after_abort. Together they tell you whether you’re traffic-limited, overloaded, or wasting capacity.


2. Labels — get these right#

GOOD LABELS (low cardinality, high value)
  model            "llama-3-70b"
  model_version    "v1.5.0-fp8"     ← include the full serving config identity
  tenant           "team-search"     (if you have < ~100 tenants)
  prompt_bucket    "0-512", "512-2k", "2k-8k", "8k-32k", "32k+"
  status           "success", "error", "cancelled"
  error_type       "oom", "timeout", "rejected", "upstream"

BAD LABELS (cardinality explosion)
  request_id       ✗ unbounded
  user_id          ✗ potentially millions
  prompt_hash      ✗ unbounded
  exact_length     ✗ use buckets

Cardinality kills Prometheus. Each unique label combination is a separate time series; model × version × tenant × bucket × status at 3×2×50×5×3 = 4,500 series is fine, adding user_id makes it millions.

Put high-cardinality data in logs or traces, not metrics.


3. The dashboard hierarchy#

Dashboard 1 — Service health (the one on the wall)#

┌─────────────────────┬─────────────────────┬─────────────────────┐
│ TTFT p50/p95/p99    │ ITL p50/p95/p99     │ Error rate          │
│ (by prompt bucket)  │                     │ (by type)           │
├─────────────────────┼─────────────────────┼─────────────────────┤
│ Requests/sec        │ Output tokens/sec   │ Goodput %           │
├─────────────────────┼─────────────────────┼─────────────────────┤
│ Queue wait p95      │ KV cache usage %    │ Running batch size  │
│ ↑ overload signal   │ ↑ capacity signal   │ ↑ efficiency signal │
└─────────────────────┴─────────────────────┴─────────────────────┘

Those bottom three panels are the LLM-specific ones, and they’re what distinguish a useful dashboard from a generic web-service dashboard.

Dashboard 2 — Efficiency and cost#

  cost per M output tokens (computed, trended)
  average running batch / max_num_seqs
  DCGM DRAM_ACTIVE
  prefix cache hit rate
  preemption rate
  aborted token fraction
  tokens by tenant (stacked)
  prefill vs decode time split

Dashboard 3 — Fleet#

  per-replica: KV usage, batch size, queue depth (heatmap across replicas)
  → imbalance shows immediately as a heatmap with hot rows
  GPU memory, power, temperature, throttle reasons
  model load durations
  restart counts

The per-replica heatmap is the fastest way to spot a routing problem. If one replica is at 95% KV and others at 20%, your load balancer is wrong (Section VIII.06).


4. The alerts that matter#

YAML
groups:
- name: llm-inference
  rules:
  # SLO violations — page
  - alert: TTFTSLOViolation
    expr: |
      histogram_quantile(0.95,
        sum(rate(ttft_seconds_bucket{prompt_bucket=~"0-512|512-2k"}[5m])) by (le, model)
      ) > 1.0
    for: 10m
    labels: {severity: page}

  - alert: ITLSLOViolation
    expr: |
      histogram_quantile(0.95, sum(rate(itl_seconds_bucket[5m])) by (le, model)) > 0.08
    for: 10m
    labels: {severity: page}

  # Capacity — page before it becomes an SLO violation
  - alert: QueueWaitHigh
    expr: |
      histogram_quantile(0.95, sum(rate(queue_wait_seconds_bucket[5m])) by (le)) > 2
    for: 5m
    labels: {severity: page}

  - alert: KVCacheNearExhaustion
    expr: avg_over_time(kv_cache_usage_ratio[5m]) > 0.92
    for: 10m
    labels: {severity: warn}

  # Efficiency — ticket, not page
  - alert: LowBatchUtilization
    expr: avg_over_time(avg_running_batch_size[30m]) / max_num_seqs < 0.25
    for: 30m
    labels: {severity: ticket}
    annotations:
      summary: "Running at {{ $value }} of capacity — consolidate or investigate routing"

  - alert: PreemptionThrashing
    expr: rate(preemptions_total[5m]) / rate(engine_steps_total[5m]) > 0.02
    for: 10m
    labels: {severity: warn}

  - alert: HighAbortedTokenFraction
    expr: |
      rate(tokens_generated_after_abort[15m]) / rate(decode_tokens_total[15m]) > 0.15
    for: 30m
    labels: {severity: ticket}

  # Health
  - alert: GPUThrottling
    # newer DCGM releases name this field DCGM_FI_DEV_CLOCKS_EVENT_REASONS, and it is not in
    # dcgm-exporter's default counter set — check your exporter's metric list
    expr: DCGM_FI_DEV_CLOCK_THROTTLE_REASONS != 0
    for: 15m
    labels: {severity: warn}

  - alert: ModelLoadSlow
    expr: histogram_quantile(0.9, rate(model_load_duration_seconds_bucket[1h])) > 300
    labels: {severity: ticket}

Note the alert on LowBatchUtilization. Most teams alert on things being slow; alerting on being inefficient is what catches the “16 under-utilized replicas” problem from Section X.04.


5. Deriving cost per token#

PromQL
# Cost per million output tokens
(
  sum(gpu_count) * hourly_gpu_cost / 3600
)
/
(
  sum(rate(decode_tokens_total[5m]))
) * 1e6

With hourly_gpu_cost as a recording rule or a static config value. Track this as a first-class metric — it’s the number that connects engineering work to the business.

Also useful:

PromQL
# Prefix cache hit rate
sum(rate(prefix_cache_hit_tokens_total[5m]))
/ sum(rate(prefix_cache_query_tokens_total[5m]))

# Prefill share of GPU time
sum(rate(prefill_time_seconds_total[5m]))
/ (sum(rate(prefill_time_seconds_total[5m])) + sum(rate(decode_time_seconds_total[5m])))

# Effective utilization (Section X.04)
avg(gpu_busy_ratio)
* (avg(avg_running_batch_size) / max_num_seqs)
* (1 - rate(tokens_generated_after_abort[5m]) / rate(decode_tokens_total[5m]))

6. Tracing#

Metrics tell you what; traces tell you which request and why.

Python
from opentelemetry import trace
tracer = trace.get_tracer(__name__)

async def handle_request(req):
    with tracer.start_as_current_span("inference.request") as span:
        span.set_attribute("gen_ai.request.model", req.model)
        span.set_attribute("gen_ai.request.max_tokens", req.max_tokens)
        span.set_attribute("gen_ai.usage.input_tokens", n_prompt)
        with tracer.start_as_current_span("tokenize"): ...
        with tracer.start_as_current_span("queue"):    ...   # ← the interesting span
        with tracer.start_as_current_span("prefill"):  ...
        with tracer.start_as_current_span("decode"):
            span.set_attribute("batch_size_at_start", bs)
        span.set_attribute("gen_ai.usage.output_tokens", n_out)
        span.set_attribute("gen_ai.response.finish_reason", reason)

OpenTelemetry has semantic conventions for GenAI (gen_ai.* attributes) — use them, so your traces are interpretable by standard tooling.

Sample intelligently: trace 100% of errors and slow requests, 0.1% of normal ones. Tracing every request at 1,000 req/s is expensive and unnecessary.


7. Structured logging#

JSON
{
  "ts": "2026-03-14T10:23:45.123Z",
  "level": "info",
  "event": "request_complete",
  "request_id": "req_abc123",
  "model": "llama-3-70b",
  "model_version": "v1.5.0-fp8",
  "tenant": "team-search",
  "prompt_tokens": 1843,
  "output_tokens": 200,
  "cached_prompt_tokens": 1536,
  "ttft_ms": 234,
  "queue_wait_ms": 12,
  "prefill_ms": 187,
  "itl_p50_ms": 21.4,
  "itl_p99_ms": 48.2,
  "e2e_ms": 4512,
  "finish_reason": "stop",
  "batch_size_at_admission": 47,
  "preempted": false
}

One structured record per request enables analyses metrics cannot:

SQL
-- Which tenant is driving the p99?
SELECT tenant, percentile_cont(0.99) WITHIN GROUP (ORDER BY ttft_ms)
FROM requests WHERE ts > now() - interval '1 hour' GROUP BY tenant;

-- Does prefix caching help this tenant?
SELECT tenant, avg(cached_prompt_tokens::float / prompt_tokens)
FROM requests GROUP BY tenant;

-- What's the actual output length distribution?
SELECT width_bucket(output_tokens, 0, 4096, 32), count(*)
FROM requests GROUP BY 1 ORDER BY 1;

That last query is what you need for capacity planning (Section XI.02) and for building a realistic benchmark (Section X.07). You cannot get it from metrics.


8. Production implications#

  • Emit all three: metrics, traces, logs. They answer different questions.
  • Get the LLM-specific metrics right: KV usage, running batch, queue wait, prefix hit rate, aborted tokens. Generic APM tools won’t provide them.
  • Deploy dcgm-exporter for the GPU-side metrics (Section X.04).
  • Keep cardinality bounded. No user IDs in labels.
  • Bucket TTFT by prompt length. Unbucketed percentiles are meaningless.
  • Alert on efficiency, not just latency.
  • Retain structured logs long enough for capacity analysis — at least 30 days.
  • Include model_version everywhere, so you can correlate regressions with deployments.

9. Common mistakes#

Only golden signals. You need capacity and efficiency metrics too.

Unbucketed TTFT. Dominated by the length tail.

Cardinality explosion. From user IDs or exact lengths in labels.

Alerting on GPU utilization. It’s always 100% (Section X.04).

Alerting on queue depth instead of queue wait. Depth without cost is meaningless.

No efficiency alerts. You’ll never notice you’re running at 20% capacity.

Tracing everything. Expensive and unnecessary; sample.

Not logging per-request shape. You lose the ability to do capacity analysis.


10. Hands-on exercise#

A. Instrument fully. Add all the metrics from section 1 to a real server. Verify each appears in Prometheus with sensible values.

B. Build the dashboards. Create the three dashboards from section 3. Load-test the system and watch which panels move.

C. Test the alerts. Implement the alerts from section 4. Trigger each one deliberately (overload for queue wait, undersize for preemption, split traffic for low batch). Verify each fires.

D. Compute cost. Implement the cost-per-token query. Trend it over a day. Does it match your Section V.15 calculation?

E. Log analysis. Emit structured per-request logs. Write the three queries from section 7. Use the output-length distribution to build a realistic benchmark (Section X.07).

F. Find the imbalance. Build the per-replica heatmap. Deliberately misconfigure the load balancer and confirm the heatmap reveals it.


11. Interview questions#

  1. What metrics would you put on an LLM inference dashboard, and why those?
  2. Why must TTFT be bucketed by prompt length?
  3. What are the three most under-collected LLM serving metrics?
  4. Why can’t you alert on GPU utilization?
  5. How would you compute cost per million tokens in Prometheus?
  6. What goes in metrics vs traces vs logs?
  7. What per-request data do you need for capacity planning, and where do you store it?

12. Further reading#

  • [REFERENCE] Prometheus best practices; histogram vs summary
  • [REFERENCE] OpenTelemetry GenAI semantic conventions
  • [REFERENCE] NVIDIA DCGM exporter
  • [REFERENCE] vLLM’s /metrics endpoint — a good reference metric set. Names change between releases (KV-cache usage is vllm:kv_cache_usage_perc in current versions); read your version’s metrics page
  • Go deeper: the Observability Engineering path covers this subject as a full course — histograms and cardinality (II), SLO alerting (IV.03), GPU telemetry field by field (V.02) and engine metrics (V.03)
  • [FUNDAMENTAL] Google SRE Book, “Monitoring Distributed Systems”
  • Next: 11 — Case studies

↑↓ navigate↵ openesc close