PidokuInfra

Quality and Evals

Advanced 50 min Difficulty 3/5 Topic 07 of 07

Prerequisites 01, 04

The idea in one minute#

An inference platform can be fast, cheap and available while producing worse answers than last week. Nothing in the first five layers detects that. Quality has to be measured as its own signal: a score attached to a response, produced by a rule, another model, or a person, and then aggregated, graphed and alerted on like latency.

For an infrastructure engineer the point is narrow and important: infrastructure changes alter outputs. Quantization, a new engine version, a different sampling default, a truncated context, a cache bug — each can degrade answers silently. Quality telemetry is how you catch your own regressions.

An analogy#

A bakery can measure loaves per hour, oven temperature and delivery times and still be selling bread nobody wants. Someone has to taste it: a sample from each batch, against a standard, every day.

A picture#

flowchart TB
  REQ["Request and response<br/>with trace and response ID"] --> ON["Online checks<br/>cheap, every request"]
  REQ --> SAMP["Sample"]
  SAMP --> JUDGE["Evaluators<br/>rules, model judge, human review"]
  USER["User feedback<br/>thumbs, edits, retries"] --> SCORE
  ON --> SCORE["Scores<br/>joined to the trace by ID"]
  JUDGE --> SCORE
  SCORE --> AGG["Aggregate by model revision,<br/>engine version, quantization, tenant"]
  AGG --> ALERT["Regression alert,<br/>canary gate"]
  OFF["Offline eval set<br/>fixed prompts, known answers"] --> GATE["Release gate"]
  class REQ,USER neutral
  class ON,JUDGE,SAMP compute
  class SCORE,AGG memory
  class ALERT,GATE queue
  class OFF neutral

How it really works#

Signals that cost nothing#

Before any evaluator, the engine and gateway already report quality-adjacent facts:

SignalWhat a change means
Finish reason distributionMore length → truncated answers; fewer tool_calls → the model stopped using tools
Output length distributionA sudden shift usually means a template, stop-token or sampling change
Empty or whitespace outputsChat-template or tokenizer mismatch
Structured-output parse failuresGuided decoding broke, or was turned off
Tool-call validityArguments that fail schema validation
Refusal rateA safety-filter or system-prompt change
Retry / regenerate rateUsers voting with their feet
NaN or corrupted outputsNumerical failure; vLLM counts these as vllm:corrupted_requests
Repetition (n-gram loops)A sampling parameter or a quantization problem

These are counters and histograms. Put them on the rollout dashboard next to TTFT.

Evaluators#

KindExampleCostStrengthWeakness
DeterministicValid JSON? Contains a citation? Passes the unit tests?NegligibleExact, reproducibleNarrow
Reference-basedCompare with a known answer (exact match, similarity)LowGood for regression setsNeeds references
Model-gradedA judge model scores relevance, faithfulness, toneA model call per itemFlexibleThe judge has its own errors and drift; needs calibration against people
HumanReview a sampleHighThe standard everything else is calibrated toSlow, small samples
BehaviouralThumbs, edits, task completion, abandonmentFreeReal usersSparse, biased

Offline and online#

  • Offline: a fixed evaluation set run against every candidate — a new model revision, a quantized variant, an engine upgrade. It is a test suite, and it gates releases. For an infrastructure change the bar is simple: the outputs on the set should be equivalent within noise to the baseline. At temperature 0 you can even diff them.
  • Online: score a sample of real traffic continuously. It catches what the fixed set does not cover, and drift in what users ask.

Quality as telemetry#

Treat a score as an event attached to the thing it scored:

  • Key it by trace ID or gen_ai.response.id (lesson 04), so a low score opens the exact request, and a trace shows its scores.
  • OpenTelemetry’s GenAI conventions define an evaluation-result event for this (gen_ai.evaluation.result, still in development), carrying the evaluator’s name, the score and an optional explanation.
  • Aggregate the same way as latency: by model_revision, engine_version, quantization, tenant, prompt version. A quality regression isolated to one engine version is an infrastructure bug.

Scores are noisy. Report them with sample sizes and confidence intervals, and compare canary with baseline on the same prompts where you can.

What infrastructure changes do to quality#

ChangePossible effectHow it is caught
Weight quantization (FP8, INT4, FP4)Small accuracy loss, larger on some tasksOffline set per precision; per-task breakdown
KV-cache quantizationDegradation growing with context lengthLong-context eval cases
New engine or kernel versionNumerical differences; changed defaultsOutput diff at temperature 0
Chat-template or tokenizer changeWrong formatting, empty or rambling outputLength and finish-reason distributions
Context truncation at the gatewayMissing information, confident wrong answersFaithfulness evals; truncation counter
Speculative decoding, sampling changesShould be distribution-preserving; bugs are notOutput diff; acceptance-rate metrics
Prefix-cache or KV-sharing bugOne user’s context leaking into another’s answerA correctness and security incident: canary strings in tests
Routing to a fallback modelA weaker model answers under loadLog gen_ai.response.model; score by actual model

The last row is common and easy to miss: a cascade or fallback that keeps availability at 100% by quietly answering with a smaller model. Always record which model actually answered.

Guardrails are telemetry too#

Input and output filters (safety, PII, prompt-injection detection) are components on the request path: measure their latency, their trigger rate by category and tenant, and their false-positive rate from review. A filter whose trigger rate triples overnight is either an attack or a broken deploy; either way you want the alert.

Privacy#

Evaluation needs content. Everything in lesson 04’s content policy applies: opt-in, redaction, sampling, short retention, restricted access. Where content cannot be kept, run evaluators inline and keep only the scores.

Code#

A canary comparison: the latency metrics look identical, the quality metrics do not.

Go
// canary.go — a rollout that passes every latency check and fails on quality.
package main

import (
	"fmt"
	"math"
	"math/rand"
)

type Stats struct {
	n                                  int
	ttft, length, truncated, parseFail float64
	scoreSum, scoreSq                  float64
}

func run(rng *rand.Rand, n int, quantized bool) Stats {
	var s Stats
	for i := 0; i < n; i++ {
		s.n++
		s.ttft += 0.18 + rng.ExpFloat64()*0.05
		length := 220 + rng.NormFloat64()*60
		score := 0.86 + rng.NormFloat64()*0.08
		trunc, parse := rng.Float64() < 0.01, rng.Float64() < 0.004
		if quantized { // the candidate: same speed, subtly worse outputs
			length += 35
			score -= 0.03
			trunc = rng.Float64() < 0.035
			parse = rng.Float64() < 0.02
		}
		if trunc {
			s.truncated++
		}
		if parse {
			s.parseFail++
		}
		s.length += length
		s.scoreSum += score
		s.scoreSq += score * score
	}
	return s
}

func main() {
	rng := rand.New(rand.NewSource(13))
	base, cand := run(rng, 4000, false), run(rng, 4000, true)

	row := func(name string, b, c float64, unit string) {
		fmt.Printf("%-28s %10.3f %10.3f %s\n", name, b, c, unit)
	}
	fmt.Printf("%-28s %10s %10s\n", "", "baseline", "canary")
	row("TTFT mean", base.ttft/float64(base.n), cand.ttft/float64(cand.n), "s      ← looks fine")
	row("output length mean", base.length/float64(base.n), cand.length/float64(cand.n), "tokens ← shifted")
	row("finish_reason=length", 100*base.truncated/float64(base.n), 100*cand.truncated/float64(cand.n), "%      ← tripled")
	row("JSON parse failures", 100*base.parseFail/float64(base.n), 100*cand.parseFail/float64(cand.n), "%      ← 5x")

	mean := func(s Stats) float64 { return s.scoreSum / float64(s.n) }
	se := func(s Stats) float64 {
		m := mean(s)
		return math.Sqrt((s.scoreSq/float64(s.n) - m*m) / float64(s.n))
	}
	diff := mean(cand) - mean(base)
	ci := 1.96 * math.Sqrt(se(base)*se(base)+se(cand)*se(cand))
	row("judge score mean", mean(base), mean(cand), "")
	fmt.Printf("\nscore difference %.3f ± %.3f (95%%) → ", diff, ci)
	if math.Abs(diff) > ci {
		fmt.Println("a real regression. Block the rollout.")
	} else {
		fmt.Println("within noise.")
	}
}

Remember this#

  • Quality is a separate signal. No infrastructure metric detects worse answers.
  • Free signals first: finish reasons, output length, parse failures, refusals, retries.
  • Offline evals gate releases; online evals watch real traffic. Join scores to traces by ID.
  • Infrastructure changes — quantization, engine upgrades, truncation, fallbacks — change outputs. Treat them as changes that need a quality check.
  • Record which model actually answered.

Try it#

  1. Run canary.go with 200 samples instead of 4,000. Is the score difference still significant? What does that say about sample sizes for canaries?
  2. List five free quality signals available in a system you know, and which of them you graph.
  3. Design the offline check for an engine upgrade that should change nothing.

Check yourself#

  1. Why can a rollout pass every latency and error check and still be a regression?
  2. What are the trade-offs between deterministic, model-graded and human evaluation?
  3. How do you connect a quality score back to the request that produced it?

↑↓ navigate↵ openesc close