The idea in one minute#
An inference platform can be fast, cheap and available while producing worse answers than last week. Nothing in the first five layers detects that. Quality has to be measured as its own signal: a score attached to a response, produced by a rule, another model, or a person, and then aggregated, graphed and alerted on like latency.
For an infrastructure engineer the point is narrow and important: infrastructure changes alter outputs. Quantization, a new engine version, a different sampling default, a truncated context, a cache bug — each can degrade answers silently. Quality telemetry is how you catch your own regressions.
An analogy#
A bakery can measure loaves per hour, oven temperature and delivery times and still be selling bread nobody wants. Someone has to taste it: a sample from each batch, against a standard, every day.
A picture#
flowchart TB REQ["Request and response<br/>with trace and response ID"] --> ON["Online checks<br/>cheap, every request"] REQ --> SAMP["Sample"] SAMP --> JUDGE["Evaluators<br/>rules, model judge, human review"] USER["User feedback<br/>thumbs, edits, retries"] --> SCORE ON --> SCORE["Scores<br/>joined to the trace by ID"] JUDGE --> SCORE SCORE --> AGG["Aggregate by model revision,<br/>engine version, quantization, tenant"] AGG --> ALERT["Regression alert,<br/>canary gate"] OFF["Offline eval set<br/>fixed prompts, known answers"] --> GATE["Release gate"] class REQ,USER neutral class ON,JUDGE,SAMP compute class SCORE,AGG memory class ALERT,GATE queue class OFF neutral
How it really works#
Signals that cost nothing#
Before any evaluator, the engine and gateway already report quality-adjacent facts:
| Signal | What a change means |
|---|---|
| Finish reason distribution | More length → truncated answers; fewer tool_calls → the model stopped using tools |
| Output length distribution | A sudden shift usually means a template, stop-token or sampling change |
| Empty or whitespace outputs | Chat-template or tokenizer mismatch |
| Structured-output parse failures | Guided decoding broke, or was turned off |
| Tool-call validity | Arguments that fail schema validation |
| Refusal rate | A safety-filter or system-prompt change |
| Retry / regenerate rate | Users voting with their feet |
| NaN or corrupted outputs | Numerical failure; vLLM counts these as vllm:corrupted_requests |
| Repetition (n-gram loops) | A sampling parameter or a quantization problem |
These are counters and histograms. Put them on the rollout dashboard next to TTFT.
Evaluators#
| Kind | Example | Cost | Strength | Weakness |
|---|---|---|---|---|
| Deterministic | Valid JSON? Contains a citation? Passes the unit tests? | Negligible | Exact, reproducible | Narrow |
| Reference-based | Compare with a known answer (exact match, similarity) | Low | Good for regression sets | Needs references |
| Model-graded | A judge model scores relevance, faithfulness, tone | A model call per item | Flexible | The judge has its own errors and drift; needs calibration against people |
| Human | Review a sample | High | The standard everything else is calibrated to | Slow, small samples |
| Behavioural | Thumbs, edits, task completion, abandonment | Free | Real users | Sparse, biased |
Offline and online#
- Offline: a fixed evaluation set run against every candidate — a new model revision, a quantized variant, an engine upgrade. It is a test suite, and it gates releases. For an infrastructure change the bar is simple: the outputs on the set should be equivalent within noise to the baseline. At temperature 0 you can even diff them.
- Online: score a sample of real traffic continuously. It catches what the fixed set does not cover, and drift in what users ask.
Quality as telemetry#
Treat a score as an event attached to the thing it scored:
- Key it by trace ID or
gen_ai.response.id(lesson 04), so a low score opens the exact request, and a trace shows its scores. - OpenTelemetry’s GenAI conventions define an evaluation-result event for this
(
gen_ai.evaluation.result, still in development), carrying the evaluator’s name, the score and an optional explanation. - Aggregate the same way as latency: by
model_revision,engine_version, quantization,tenant, prompt version. A quality regression isolated to one engine version is an infrastructure bug.
Scores are noisy. Report them with sample sizes and confidence intervals, and compare canary with baseline on the same prompts where you can.
What infrastructure changes do to quality#
| Change | Possible effect | How it is caught |
|---|---|---|
| Weight quantization (FP8, INT4, FP4) | Small accuracy loss, larger on some tasks | Offline set per precision; per-task breakdown |
| KV-cache quantization | Degradation growing with context length | Long-context eval cases |
| New engine or kernel version | Numerical differences; changed defaults | Output diff at temperature 0 |
| Chat-template or tokenizer change | Wrong formatting, empty or rambling output | Length and finish-reason distributions |
| Context truncation at the gateway | Missing information, confident wrong answers | Faithfulness evals; truncation counter |
| Speculative decoding, sampling changes | Should be distribution-preserving; bugs are not | Output diff; acceptance-rate metrics |
| Prefix-cache or KV-sharing bug | One user’s context leaking into another’s answer | A correctness and security incident: canary strings in tests |
| Routing to a fallback model | A weaker model answers under load | Log gen_ai.response.model; score by actual model |
The last row is common and easy to miss: a cascade or fallback that keeps availability at 100% by quietly answering with a smaller model. Always record which model actually answered.
Guardrails are telemetry too#
Input and output filters (safety, PII, prompt-injection detection) are components on the request path: measure their latency, their trigger rate by category and tenant, and their false-positive rate from review. A filter whose trigger rate triples overnight is either an attack or a broken deploy; either way you want the alert.
Privacy#
Evaluation needs content. Everything in lesson 04’s content policy applies: opt-in, redaction, sampling, short retention, restricted access. Where content cannot be kept, run evaluators inline and keep only the scores.
Code#
A canary comparison: the latency metrics look identical, the quality metrics do not.
// canary.go — a rollout that passes every latency check and fails on quality.
package main
import (
"fmt"
"math"
"math/rand"
)
type Stats struct {
n int
ttft, length, truncated, parseFail float64
scoreSum, scoreSq float64
}
func run(rng *rand.Rand, n int, quantized bool) Stats {
var s Stats
for i := 0; i < n; i++ {
s.n++
s.ttft += 0.18 + rng.ExpFloat64()*0.05
length := 220 + rng.NormFloat64()*60
score := 0.86 + rng.NormFloat64()*0.08
trunc, parse := rng.Float64() < 0.01, rng.Float64() < 0.004
if quantized { // the candidate: same speed, subtly worse outputs
length += 35
score -= 0.03
trunc = rng.Float64() < 0.035
parse = rng.Float64() < 0.02
}
if trunc {
s.truncated++
}
if parse {
s.parseFail++
}
s.length += length
s.scoreSum += score
s.scoreSq += score * score
}
return s
}
func main() {
rng := rand.New(rand.NewSource(13))
base, cand := run(rng, 4000, false), run(rng, 4000, true)
row := func(name string, b, c float64, unit string) {
fmt.Printf("%-28s %10.3f %10.3f %s\n", name, b, c, unit)
}
fmt.Printf("%-28s %10s %10s\n", "", "baseline", "canary")
row("TTFT mean", base.ttft/float64(base.n), cand.ttft/float64(cand.n), "s ← looks fine")
row("output length mean", base.length/float64(base.n), cand.length/float64(cand.n), "tokens ← shifted")
row("finish_reason=length", 100*base.truncated/float64(base.n), 100*cand.truncated/float64(cand.n), "% ← tripled")
row("JSON parse failures", 100*base.parseFail/float64(base.n), 100*cand.parseFail/float64(cand.n), "% ← 5x")
mean := func(s Stats) float64 { return s.scoreSum / float64(s.n) }
se := func(s Stats) float64 {
m := mean(s)
return math.Sqrt((s.scoreSq/float64(s.n) - m*m) / float64(s.n))
}
diff := mean(cand) - mean(base)
ci := 1.96 * math.Sqrt(se(base)*se(base)+se(cand)*se(cand))
row("judge score mean", mean(base), mean(cand), "")
fmt.Printf("\nscore difference %.3f ± %.3f (95%%) → ", diff, ci)
if math.Abs(diff) > ci {
fmt.Println("a real regression. Block the rollout.")
} else {
fmt.Println("within noise.")
}
}Remember this#
- Quality is a separate signal. No infrastructure metric detects worse answers.
- Free signals first: finish reasons, output length, parse failures, refusals, retries.
- Offline evals gate releases; online evals watch real traffic. Join scores to traces by ID.
- Infrastructure changes — quantization, engine upgrades, truncation, fallbacks — change outputs. Treat them as changes that need a quality check.
- Record which model actually answered.
Try it#
- Run
canary.gowith 200 samples instead of 4,000. Is the score difference still significant? What does that say about sample sizes for canaries? - List five free quality signals available in a system you know, and which of them you graph.
- Design the offline check for an engine upgrade that should change nothing.
Check yourself#
- Why can a rollout pass every latency and error check and still be a regression?
- What are the trade-offs between deterministic, model-graded and human evaluation?
- How do you connect a quality score back to the request that produced it?