Section X.10 covered the metrics catalogue and dashboards. This file covers what’s specific to LLM services and what a generic APM setup will miss.
1. What generic observability misses#
A standard APM setup gives you:
✓ request rate, error rate, latency percentiles
✓ CPU, memory, network
✓ traces across services
It does NOT give you:
✗ KV cache utilization ← your actual capacity limit
✗ running batch size ← your actual efficiency
✗ queue wait vs prefill time ← the TTFT decomposition
✗ prefix cache hit rate
✗ preemption rate
✗ tokens generated after abort
✗ per-request token counts and cost
✗ output quality signals
✗ GPU-specific health (Xid, ECC, throttling)Every item in the second list is essential and none of it comes for free. Instrumenting them is a deliberate project.
2. The three pillars, applied#
METRICS aggregate, cheap, alertable
→ "is the service healthy right now?"
→ the catalogue in Section X.10
TRACES per-request, sampled, detailed
→ "why was THIS request slow?"
→ span the gateway → scheduler → engine → response
LOGS per-request, structured, queryable
→ "what is the distribution of X across all requests?"
→ the source for capacity planning and benchmark designFor LLM services, logs do the heaviest lifting, because the questions you need to answer are distributional (length distributions, per-tenant patterns, cost attribution) and metrics can’t carry that cardinality.
3. The per-request log record#
{
"ts": "2026-03-14T10:23:45.123Z",
"event": "inference_complete",
"request_id": "req_abc123",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"tenant": "team-search",
"api_key_id": "key_xyz",
"model": "llama-3-70b",
"model_version": "v1.5.0-fp8-tp8",
"replica": "llama-70b-7f9d4-x2k1p",
"region": "us-east-1",
"prompt_tokens": 1843,
"cached_prompt_tokens": 1536,
"output_tokens": 200,
"max_tokens_requested": 2048,
"queue_wait_ms": 12.4,
"tokenize_ms": 1.2,
"prefill_ms": 187.3,
"ttft_ms": 234.1,
"itl_p50_ms": 21.4,
"itl_p95_ms": 38.2,
"itl_max_ms": 112.0,
"detokenize_ms": 8.1,
"e2e_ms": 4512.0,
"batch_size_at_admission": 47,
"kv_blocks_peak": 128,
"kv_block_seconds": 578.0,
"preempted": false,
"preemption_count": 0,
"finish_reason": "stop",
"temperature": 0.7,
"top_p": 0.9,
"guided_format": null,
"n": 1,
"status": 200,
"client_disconnected": false,
"tokens_after_disconnect": 0
}Fields most often missing and most valuable:
cached_prompt_tokens— measures prefix caching’s actual benefitkv_block_seconds— the missing term in cost attribution (Section XI.03)batch_size_at_admission— efficiency signal per requesttokens_after_disconnect— pure waste, per requestreplica— enables the per-replica imbalance analysis
4. Tracing an LLM request properly#
// go.opentelemetry.io/otel — one span per stage of the request.
var tracer = otel.Tracer("inference")
func handle(ctx context.Context, req *Request) error {
ctx, root := tracer.Start(ctx, "inference.request")
defer root.End()
// OpenTelemetry GenAI semantic conventions
root.SetAttributes(
attribute.String("gen_ai.system", "vllm"),
attribute.String("gen_ai.request.model", req.Model),
attribute.Int("gen_ai.request.max_tokens", req.MaxTokens),
attribute.Float64("gen_ai.request.temperature", req.Temperature),
)
// span runs one stage and records its duration and error.
span := func(name string, fn func(trace.Span) error) error {
_, s := tracer.Start(ctx, name)
defer s.End()
return fn(s)
}
span("auth", func(trace.Span) error { return authenticate(ctx, req) })
span("rate_limit", func(trace.Span) error { return checkQuota(ctx, req) })
span("tokenize", func(s trace.Span) error {
req.IDs = tokenize(req)
s.SetAttributes(attribute.Int("gen_ai.usage.input_tokens", len(req.IDs)))
return nil
})
span("route", func(s trace.Span) error {
req.Replica = router.Choose(req)
s.SetAttributes(attribute.String("inference.replica", req.Replica.ID),
attribute.Bool("inference.prefix_cache_hit", req.Replica.LikelyHit))
return nil
})
span("queue", func(s trace.Span) error { // ← the interesting span
<-req.Admitted
s.SetAttributes(attribute.Int("inference.batch_size", engine.RunningCount()),
attribute.Float64("inference.kv_usage", engine.KVUsage()))
return nil
})
span("prefill", func(s trace.Span) error {
s.SetAttributes(attribute.Int("inference.cached_tokens", req.CachedTokens))
return nil
})
return span("decode", func(s trace.Span) error {
s.SetAttributes(attribute.Int("gen_ai.usage.output_tokens", req.OutputTokens),
attribute.String("gen_ai.response.finish_reason", req.FinishReason))
return nil
})
}The queue span with the batch size and KV usage as attributes is what makes a trace
diagnostic rather than decorative: when you look at a slow request, you immediately see the
system state it encountered.
Sampling: 100% of errors, 100% of requests exceeding the SLO, 0.1-1% of the rest. Head-based sampling misses slow requests; use tail-based sampling if your tracing backend supports it.
5. Quality observability#
The dimension generic tooling has no concept of.
CHEAP, CONTINUOUS SIGNALS (compute for every request)
output token count
finish_reason distribution
structured output validity (if applicable)
presence of refusal patterns (regex on a sample)
repetition detection (n-gram repetition rate)
language of the response vs the request
SAMPLED SIGNALS (1% of traffic)
LLM-as-judge score against a rubric
embedding-based similarity to the reference model's output
toxicity / safety classifier score
USER SIGNALS (all traffic, sparse)
thumbs up/down
regeneration rate ← strong negative signal
copy/paste rate ← positive signal if measurable
conversation abandonment
turns to resolution
PERIODIC (daily)
full evaluation suite against production traffic samples
drift detection: has the input distribution changed?Regeneration rate is the best single user-side quality signal available in most products: it requires no explicit feedback, correlates well with dissatisfaction, and moves quickly.
6. The alerting philosophy#
PAGE (wake someone up)
SLO violation sustained > 10 minutes
error rate > 5%
a region or model completely unavailable
security event
TICKET (fix during business hours)
efficiency degradation (low batch, low DRAM_ACTIVE)
cost per token drift > 20%
preemption rate elevated
quality signals outside bounds
hardware degradation (ECC errors accumulating)
DASHBOARD ONLY (no alert)
normal variation
informational metrics
NEVER ALERT ON
GPU utilization
raw queue depth (use wait time)
individual request latency
anything that fires more than once a week without actionThe last line is the discipline that keeps alerting useful. An alert that fires weekly and is always acknowledged without action should be deleted or converted to a ticket.
7. The on-call runbook#
For each alert, the runbook entry should contain:
ALERT: TTFTSLOViolation
WHAT IT MEANS
p95 TTFT for prompts ≤ 2k tokens has exceeded 800 ms for 10 minutes.
FIRST CHECKS (in order)
1. Grafana → Service Health → is queue_wait_p95 elevated?
YES → capacity problem. Go to CAPACITY below.
NO → prefill got slower. Go to PREFILL below.
2. Is this one replica or all? (per-replica heatmap)
one → check that replica's health; consider removing it
3. Did anything deploy in the last hour? (deployment annotations)
yes → consider rollback
CAPACITY
- check avg_running_batch vs max_num_seqs
- check kv_cache_usage
- check arrival rate vs the last 24 hours
- if arrival rate is up: scale (link to the runbook)
- if arrival rate is normal but batch is low: routing problem
PREFILL
- check the prompt length distribution (has it shifted?)
- check prefix_cache_hit_rate (has it dropped? routing change?)
- check for GPU throttling
ESCALATION
if unresolved in 20 minutes, page <team>
RELATED
dashboards, past incidents, the capacity planA runbook that names specific dashboards and specific next queries is worth ten times one that says “investigate the cause.”
8. Production implications#
- Budget real engineering time for observability. The LLM-specific signals don’t come from any product.
- Structured per-request logs are the foundation. Retain 30+ days.
- Use OpenTelemetry GenAI semantic conventions so your traces are portable and tool-readable.
- Instrument quality, not just performance. It’s the failure mode that matters.
- Tail-based trace sampling so you capture the slow requests.
- Write runbooks with specific next steps, and update them after every incident.
- Prune alerts quarterly. Delete anything that fires without action.
- Emit
model_versionandreplicaeverywhere — correlation depends on it.
9. Common mistakes#
Relying on generic APM. Misses everything LLM-specific.
No quality observability. You learn about regressions from support tickets.
Head-based trace sampling. Systematically misses slow requests.
Metrics without per-request logs. Can’t do distributional analysis or capacity planning.
Alerting on GPU utilization.
Runbooks that say “investigate.”
Not emitting model_version. Can’t correlate regressions with deployments.
Alert fatigue. Alerts that fire without action train people to ignore them.
10. Hands-on exercise#
A. Emit the full record. Implement the per-request log from section 3, including the four “often missing” fields. Verify each is populated correctly.
B. Trace properly. Implement the tracing from section 4, with the queue span carrying system state. Trace a slow request and confirm you can see why it was slow from the trace alone.
C. Quality signals. Implement the cheap continuous quality signals from section 5. Establish baselines over a week. Then deploy a degraded model and see which signals move first and by how much.
D. Write a runbook. For your top three alerts, write the runbook entries in the format of section 7, with specific dashboards and queries.
E. Audit alerts. For an existing service, list every alert and when it last fired. How many fired without resulting in action? Delete or downgrade those.
F. Distributional analysis. Using your per-request logs, answer: what’s the output-length distribution? Which tenant has the highest cost per request? What’s the prefix cache benefit per tenant? These queries are impossible without the logs.
11. Interview questions#
- What does generic APM miss for an LLM service?
- What per-request fields would you log, and what does each enable?
- Why do logs matter more than metrics for LLM capacity planning?
- How would you observe output quality continuously?
- Why is tail-based trace sampling important here?
- What makes a good runbook entry?
- What would you never alert on, and why?
12. Further reading#
- [REFERENCE] OpenTelemetry GenAI semantic conventions — widely supported, and as of October
2026 still entirely at “Development” stability in their own repository
(
open-telemetry/semantic-conventions-genai): pin versions and expect renames - Go deeper: the Observability Engineering path — especially V.04 Tracing LLM requests, V.05 Platform and fleet, V.07 Quality and evals and VI.02 How it is changing
- [FUNDAMENTAL] Google SRE Book, “Monitoring Distributed Systems” and “Being On-Call”
- [REFERENCE] vLLM
/metricsand its Prometheus integration - Next: 09 — Security and isolation