PidokuInfra

Versioning, Canary Deployments, and A/B Testing

Intermediate 1h Difficulty 3/5 Topic 10 of 14

Prerequisites 07, 09, IV.12


1. What is it?#

How you change what’s serving without breaking anything, and how you find out whether the change was good.

VERSIONING   identifying exactly what is deployed
CANARY       exposing a change to a small fraction first
A/B TESTING  measuring whether the change is better
ROLLBACK     undoing it quickly when it isn't

The LLM-specific difficulty: the failure mode is usually “slightly worse answers,” which no conventional monitoring detects.


2. Why it’s different for LLMs#

ORDINARY SERVICE                     LLM SERVICE
errors are 4xx/5xx                   errors include "the answer got worse"
canary detects errors in minutes     quality regressions take hours-days to detect
rollback is instant                  rollback requires reloading weights (minutes)
"the same version" is unambiguous    version = weights + quantization + engine +
                                       kernels + sampling defaults + template
deterministic outputs                nondeterministic (Section IV.12)

That fourth row is the one teams get wrong. “We deployed Llama-3-70B” is not a version. The serving configuration is part of the model’s identity.


3. What “version” must capture#

model_version = {
  weights:        checkpoint hash / URI
  quantization:   method, bits, group size, excluded layers, calibration set hash
  engine:         vLLM 0.x.y / TRT-LLM engine hash
  kernels:        attention backend, GEMM backend
  parallelism:    TP degree, PP degree
  precision:      dtype, KV cache dtype
  chat_template:  hash
  sampling:       default temperature/top_p if the server sets them
  tokenizer:      version/hash
  runtime:        CUDA version, driver version
}

Two deployments differing in any of these can produce different outputs. Record all of it, emit it as a label on every metric, and return it in an API response header (X-Model-Version) so you can correlate user reports with deployments.


4. Deployment strategies#

BLUE/GREEN
  Stand up the full new fleet, switch traffic, keep the old fleet warm.
  ✓ instant rollback (switch back)
  ✗ 2x GPU cost during the transition — expensive for LLMs
  → use for high-risk changes if you can afford it

ROLLING
  Replace replicas one at a time.
  ✓ no extra capacity
  ✗ slow (each replica takes minutes to start and drain)
  ✗ mixed versions serving simultaneously for a long window
  → the default; be aware of the mixed-version window

CANARY
  Route a small % of traffic to the new version, increase gradually.
  ✓ limits blast radius
  ✓ enables measurement before full rollout
  ✗ needs traffic-splitting infrastructure and enough traffic to be significant
  → the right default for model changes

SHADOW / MIRROR
  Send a copy of production traffic to the new version; discard its responses.
  ✓ zero user risk
  ✓ real traffic distribution
  ✗ 2x inference cost for the mirrored fraction
  ✗ can't measure user-facing outcomes (nobody sees the responses)
  → excellent for performance validation, limited for quality

For a quantization change or a model upgrade, use canary with quality measurement. For an engine version bump, shadow first (to catch crashes and performance regressions), then canary.


5. What to measure in a canary#

TIER 1 — must not regress (minutes to detect)
  error rate (5xx, 429, timeouts)
  p95 TTFT, p95 ITL
  throughput per GPU
  OOM / crash rate
  → automated rollback triggers

TIER 2 — quality proxies (minutes to hours)
  output length distribution        ← a good early signal!
  stop-reason distribution (eos vs length vs stop-string)
  refusal rate
  structured-output validity rate (if applicable)
  tool-call success rate (if applicable)
  → alert, investigate

TIER 3 — quality (hours to days)
  offline eval suite on the exact serving config
  LLM-as-judge on sampled production traffic
  human evaluation on a sample
  user-facing signals: thumbs up/down, regeneration rate, conversation length,
    task completion
  → the real decision

Tier 2’s “output length distribution” deserves emphasis. It is a surprisingly sensitive and cheap early indicator: a quantized model that starts rambling or truncating shows up in the length histogram within minutes, long before any eval completes.


6. Tiny example — a canary configuration#

YAML
# Traffic splitting at the gateway
routes:
  - model: llama-3-70b
    versions:
      - id: v1.4.0-fp16
        weight: 95
      - id: v1.5.0-fp8          # the canary
        weight: 5
    sticky_by: conversation_id   # keep a conversation on one version!

# Automated rollback
guards:
  - metric: error_rate
    version: v1.5.0-fp8
    threshold: "> 1.5x baseline"
    window: 5m
    action: rollback
  - metric: p95_itl_ms
    threshold: "> 1.2x baseline"
    window: 10m
    action: rollback
  - metric: mean_output_tokens
    threshold: "outside [0.8x, 1.25x] baseline"
    window: 30m
    action: alert

sticky_by: conversation_id matters. If turn 1 of a conversation is served by v1.4 and turn 2 by v1.5, the user sees an inconsistent voice and your prefix cache misses. Pin a conversation to a version.


7. A/B testing for quality#

DESIGN
  - Randomize by USER or CONVERSATION, not by request.
    (Per-request randomization means a user sees both versions mid-conversation.)
  - Run long enough for significance. LLM quality signals are noisy;
    you typically need thousands of conversations, not hundreds.
  - Pre-register the metric. Deciding what to measure after seeing the data
    is how you fool yourself.

METRICS THAT ACTUALLY WORK
  - Regeneration rate (user asks again) — strong negative signal
  - Conversation length / turns to resolution
  - Explicit feedback (thumbs), though sparse and biased
  - Task completion (if measurable in your product)
  - LLM-as-judge on paired outputs (cheap, correlates reasonably with humans,
    but has its own biases — position bias, length bias)

METRICS THAT MISLEAD
  - Perplexity (doesn't measure what users care about)
  - Public benchmarks (not your distribution)
  - Average response length alone (longer ≠ better)

Statistical caution: LLM quality differences from quantization are often 1-3%, and A/B tests on noisy human signals need large samples to detect that. Budget for a long test or accept that you’re measuring “no catastrophic regression” rather than “which is better.”


8. Rollback#

Fast rollback requires the old version to still be RUNNABLE:
  - keep the old checkpoint in the local node cache
  - keep the old engine artifacts (TRT engines, compile caches)
  - keep the old container image pulled
  - ideally keep old replicas warm during the canary

Rollback time:
  traffic shift only (blue/green or canary):     seconds
  redeploy old version (rolling):                minutes (full cold start)

Always be able to roll back in seconds for the first hours of a deployment. That means keeping the old fleet warm, which costs money — budget it as part of the deployment.


9. Common mistakes#

Treating the model checkpoint as the version. The serving config is part of it.

No quality measurement in the pipeline. Every conventional test passes while answers degrade.

Per-request A/B randomization. Users see inconsistent behavior mid-conversation.

Canary too small to be significant. 1% of traffic for 30 minutes tells you nothing about quality.

No automated rollback guards. Manual detection takes hours.

Rolling update with a short grace period. Truncates in-flight generations.

Not keeping the old version warm. Rollback takes as long as a deployment.

Mixed versions with prefix caching and no stickiness. Cache misses and inconsistent output.


10. Hands-on exercise#

A. Define your version. Write the full version record for a model you serve, capturing every field in section 3. Emit it as a metric label and an API header.

B. Build a canary. Set up two versions behind a router with configurable traffic weights and conversation stickiness. Verify a conversation stays on one version.

C. Length distribution as a signal. Deploy a quantized version alongside FP16. Compare the output length distributions on identical prompts. Is the difference detectable? How many samples do you need?

D. Automated guards. Implement the rollback guards from section 6. Deliberately deploy a broken version (e.g. wrong chat template) and verify the guard fires.

E. A/B power analysis. For a quality metric you can measure (e.g. regeneration rate), compute how many conversations you need to detect a 3% relative change at 95% confidence. Is that achievable with your traffic?


11. Interview questions#

  1. What constitutes a “model version” for an LLM service?
  2. Why does conventional canary monitoring fail to catch LLM quality regressions?
  3. What early signals indicate a quality problem within minutes?
  4. Why randomize A/B tests by user rather than by request?
  5. Compare blue/green, rolling, canary, and shadow for an LLM model change.
  6. How do you achieve seconds-fast rollback, and what does it cost?
  7. Why does version stickiness matter with prefix caching?

12. Further reading#

  • [FUNDAMENTAL] Google SRE Book, canary and release engineering chapters
  • [REFERENCE] Argo Rollouts / Flagger documentation (progressive delivery)
  • [ESTABLISHED] Kohavi et al., Trustworthy Online Controlled Experiments
  • Next: 11 — vLLM architecture

↑↓ navigate↵ openesc close