Section VIII.10 covered the mechanics of versioning and canary. This file covers the operational process: what a safe model release actually looks like end to end.
1. What makes model rollouts different from code rollouts#
CODE ROLLOUT MODEL ROLLOUT
failures are errors (5xx) failures are "worse answers"
detected in minutes detected in hours to days
rollback is instant rollback needs the old weights loaded
artifact is MB artifact is 10-800 GB
identical behavior guaranteed nondeterministic by nature
tests are deterministic evaluation is statisticalThe core problem: your existing deployment safety machinery cannot detect the failure mode that matters. Every metric is green while quality degrades.
Diagram — A gated rollout#
flowchart LR
B["Build<br/>pinned weights + engine + config"] --> EV["Offline eval<br/>quality and performance gates"]
EV --> SH["Shadow traffic"]
SH --> C1["Canary 1-5%"]
C1 --> G{"SLO and quality<br/>gates hold?"}
G -->|"yes"| C2["25%, then 100%"]
G -->|"no"| RB["Automatic rollback"]
C2 --> OLD["Keep the old version warm<br/>until confident"]
class B,EV neutral
class SH,C1 io
class G queue
class C2,OLD compute
class RB warn2. The release process#
STAGE 0 — PREPARE
□ artifact built and checksummed
□ quantization done offline, config recorded
□ engine artifacts (TRT engines / compile cache) built per GPU arch
□ distributed to all regional storage
□ node caches pre-populated
□ full version record created (Section VIII.10)
STAGE 1 — OFFLINE EVALUATION
□ numerical: logit diff, KL divergence, top-1 agreement vs reference
□ task: your benchmark suite, on YOUR domain
□ behavioral: 500+ token generations, instruction following,
structured output validity, refusal behavior
□ regression suite: known-tricky prompts from past incidents
→ GATE: all within tolerance, or explicitly waived with justification
STAGE 2 — SHADOW (optional but recommended for engine changes)
□ mirror 1-5% of production traffic to the new version
□ discard responses; compare latency, memory, errors, crash rate
□ run for 2-24 hours
→ GATE: no crashes, latency within 10%, no OOM
STAGE 3 — CANARY
□ 1% of traffic, one region, non-premium tiers
□ automated guards active (Section VIII.10)
□ monitor: error rate, TTFT, ITL, output length distribution,
stop-reason distribution, structured validity, regeneration rate
□ 2-4 hours minimum, or until statistical significance
→ GATE: all Tier 1 and Tier 2 signals within bounds
STAGE 4 — PROGRESSIVE ROLLOUT
1% → 5% → 25% → 50% → 100%
hold at each step for at least 1 hour (longer for the early steps)
□ automated rollback guards remain active throughout
□ human check at 25% and 50%
→ GATE at each step
STAGE 5 — SOAK
□ 100% for 48-72 hours with the old version still deployable
□ monitor the slow signals: user feedback, conversation length,
task completion, support tickets
→ GATE: no degradation → decommission the old version
STAGE 6 — CLEANUP
□ remove old version's replicas
□ keep the old artifact for 30 days minimum
□ update the registry, document what changedStage 5 is the one that gets cut under schedule pressure, and it’s where the real quality signals appear. Defend it.
3. The signals, by detection speed#
MINUTES (automated guards, auto-rollback)
error rate, 5xx, timeouts
p95 TTFT, p95 ITL
OOM / crash rate
throughput per GPU
MINUTES-HOURS (automated alerts, human judgment)
output length distribution ← the most sensitive cheap signal
stop-reason distribution (eos vs length vs stop-string)
structured output validity rate
refusal rate
tool-call success rate
prefix cache hit rate (a template change would move this)
HOURS-DAYS (human review)
regeneration rate
conversation length / turns to resolution
explicit user feedback
support ticket volume and content
LLM-as-judge on sampled paired outputs
human evaluation on a sampleThe middle tier is what most teams lack, and it’s where quality regressions become visible first. Output length distribution in particular is nearly free to compute and moves detectably when a model starts rambling or truncating.
4. Automated rollback guards#
guards:
# Tier 1 — automatic rollback
- metric: error_rate_5xx
condition: "> 1.5 × baseline for 5m"
action: rollback
- metric: p95_ttft_ms
condition: "> 1.3 × baseline for 10m"
action: rollback
- metric: p95_itl_ms
condition: "> 1.25 × baseline for 10m"
action: rollback
- metric: oom_count
condition: "> 0 in 15m"
action: rollback
# Tier 2 — halt the rollout, alert a human
- metric: mean_output_tokens
condition: "outside [0.80, 1.25] × baseline for 30m"
action: halt_and_alert
- metric: structured_output_validity
condition: "< 0.98 × baseline for 30m"
action: halt_and_alert
- metric: refusal_rate
condition: "outside [0.7, 1.5] × baseline for 60m"
action: halt_and_alert
- metric: finish_reason_length_fraction
condition: "> 1.4 × baseline for 30m"
action: halt_and_alert“Halt and alert” is the right action for Tier 2, not automatic rollback — these signals are noisier and a human should look before reverting.
5. Rollback readiness#
FAST ROLLBACK (seconds) requires:
□ the old version still RUNNING (traffic shift only)
□ or: the old artifact in every node cache and engine artifacts built
SLOW ROLLBACK (minutes to an hour):
□ old artifact in regional storage
□ redeploy, reload, warm up
DURING A ROLLOUT: keep the old version warm.
Cost: 2x replicas for the rollout window.
Value: seconds vs an hour of rollback time.
→ for a 4-hour progressive rollout, that's 4 GPU-hours × replicas.
Cheap insurance.Decide explicitly whether you’re paying for fast rollback, and for how long. “We’ll roll back if needed” is not a plan unless the old version is deployable in the time your SLO allows.
6. What to do when you find a regression post-rollout#
1. ROLL BACK FIRST, INVESTIGATE SECOND.
Do not debug in production while users are affected.
2. PRESERVE EVIDENCE
Sample requests and responses from the bad version (with consent /
per your privacy policy). You need them to reproduce.
3. REPRODUCE OFFLINE
With the exact serving configuration, not the reference implementation.
4. ROOT CAUSE
Common causes, in order:
- quantization degraded a capability the eval suite didn't test
- chat template changed subtly
- sampling defaults changed
- a kernel/backend change altered numerics beyond tolerance
- the checkpoint itself is different from what was evaluated
5. ADD THE TEST
Whatever you missed goes into the evaluation suite. Permanently.
6. WRITE THE POSTMORTEM
Including: why the existing gates didn't catch it.Step 5 is the compounding one. An evaluation suite that grows with every incident becomes genuinely protective within a year.
7. Special cases#
ENGINE VERSION UPGRADE (not a model change)
→ shadow first (crashes and performance are the risk)
→ but ALSO validate numerics: kernel changes alter outputs
→ the model version record must include the engine version
QUANTIZATION CHANGE
→ the highest-risk change type
→ full Stage 1-5, no shortcuts
→ soak for at least 72 hours
CHAT TEMPLATE CHANGE
→ subtle and dangerous; the model is out of distribution if wrong
→ test that the rendered prompt is byte-identical for a corpus
of conversations before and after
→ also: it invalidates the prefix cache
SAMPLING DEFAULT CHANGE
→ affects everyone who doesn't specify parameters
→ treat as a model change; canary it
EMERGENCY SECURITY PATCH
→ skip stages 2-3, go to 5% canary for 15 minutes, then full
→ accept the risk explicitly, document it, soak afterwards8. Production implications#
- The version record includes the whole serving configuration, not just the checkpoint.
- Emit
model_versionon every metric and every response header. Correlation with deployments is otherwise impossible. - Build the middle-tier signals (output length, stop reasons, structured validity). They’re cheap and they’re what catches quality regressions.
- Keep the old version warm during rollout. Seconds vs an hour of rollback.
- Soak for 48-72 hours. The slow signals need time.
- Grow the evaluation suite from incidents.
- Never roll out a model change during a low-staffing window.
9. Common mistakes#
Treating a model rollout like a code rollout. Different failure mode, different timescale.
Stopping evaluation at perplexity and MMLU. They miss what users notice.
No middle-tier signals. You detect regressions from support tickets.
Cutting the soak period.
Rolling back slowly because the old version wasn’t kept deployable.
Not recording the full version. You can’t reproduce or correlate.
Debugging in production instead of rolling back first.
Not adding a test after an incident.
10. Hands-on exercise#
A. Build the version record. Implement the full version identity from Section VIII.10 for a service. Emit it as a metric label and a response header.
B. Implement the middle-tier signals. Compute and emit: output length distribution, stop reason distribution, structured output validity. Establish baselines.
C. Run a canary. Deploy a deliberately degraded version (e.g. INT4 where you normally run FP8) at 5%. Do your signals detect it? How long did it take?
D. Test the guards. Implement the automated guards from section 4. Trigger each deliberately and verify the action.
E. Rollback drill. Time a full rollback from a 100% deployment. Then time it with the old version kept warm. Quantify the difference and the cost.
F. Grow the suite. Take a past quality incident (yours or a documented one) and write the evaluation test that would have caught it.
11. Interview questions#
- Why can’t standard deployment safety machinery catch LLM quality regressions?
- What are the middle-tier signals and why do they matter?
- Walk me through a safe model rollout, stage by stage.
- What would you put in an automated rollback guard versus a halt-and-alert?
- Why keep the old version warm during a rollout, and what does it cost?
- A chat template change — why is it dangerous and how do you validate it?
- You find a quality regression after full rollout. What’s your sequence of actions?
12. Further reading#
- [FUNDAMENTAL] Google SRE Book, release engineering and canarying chapters
- [REFERENCE] Argo Rollouts / Flagger for progressive delivery
- [ESTABLISHED] Kohavi et al., Trustworthy Online Controlled Experiments
- Next: 08 — Observability