PidokuInfra

Model Rollouts

Advanced Intermediate 1h Difficulty 3/5 Topic 07 of 11

Prerequisites VIII.10, IV.12

Section VIII.10 covered the mechanics of versioning and canary. This file covers the operational process: what a safe model release actually looks like end to end.


1. What makes model rollouts different from code rollouts#

CODE ROLLOUT                        MODEL ROLLOUT
failures are errors (5xx)           failures are "worse answers"
detected in minutes                 detected in hours to days
rollback is instant                 rollback needs the old weights loaded
artifact is MB                      artifact is 10-800 GB
identical behavior guaranteed       nondeterministic by nature
tests are deterministic             evaluation is statistical

The core problem: your existing deployment safety machinery cannot detect the failure mode that matters. Every metric is green while quality degrades.

Diagram — A gated rollout#

flowchart LR
  B["Build<br/>pinned weights + engine + config"] --> EV["Offline eval<br/>quality and performance gates"]
  EV --> SH["Shadow traffic"]
  SH --> C1["Canary 1-5%"]
  C1 --> G{"SLO and quality<br/>gates hold?"}
  G -->|"yes"| C2["25%, then 100%"]
  G -->|"no"| RB["Automatic rollback"]
  C2 --> OLD["Keep the old version warm<br/>until confident"]

  class B,EV neutral
  class SH,C1 io
  class G queue
  class C2,OLD compute
  class RB warn

2. The release process#

STAGE 0 — PREPARE
  □ artifact built and checksummed
  □ quantization done offline, config recorded
  □ engine artifacts (TRT engines / compile cache) built per GPU arch
  □ distributed to all regional storage
  □ node caches pre-populated
  □ full version record created (Section VIII.10)

STAGE 1 — OFFLINE EVALUATION
  □ numerical: logit diff, KL divergence, top-1 agreement vs reference
  □ task: your benchmark suite, on YOUR domain
  □ behavioral: 500+ token generations, instruction following,
    structured output validity, refusal behavior
  □ regression suite: known-tricky prompts from past incidents
  → GATE: all within tolerance, or explicitly waived with justification

STAGE 2 — SHADOW (optional but recommended for engine changes)
  □ mirror 1-5% of production traffic to the new version
  □ discard responses; compare latency, memory, errors, crash rate
  □ run for 2-24 hours
  → GATE: no crashes, latency within 10%, no OOM

STAGE 3 — CANARY
  □ 1% of traffic, one region, non-premium tiers
  □ automated guards active (Section VIII.10)
  □ monitor: error rate, TTFT, ITL, output length distribution,
    stop-reason distribution, structured validity, regeneration rate
  □ 2-4 hours minimum, or until statistical significance
  → GATE: all Tier 1 and Tier 2 signals within bounds

STAGE 4 — PROGRESSIVE ROLLOUT
  1% → 5% → 25% → 50% → 100%
  hold at each step for at least 1 hour (longer for the early steps)
  □ automated rollback guards remain active throughout
  □ human check at 25% and 50%
  → GATE at each step

STAGE 5 — SOAK
  □ 100% for 48-72 hours with the old version still deployable
  □ monitor the slow signals: user feedback, conversation length,
    task completion, support tickets
  → GATE: no degradation → decommission the old version

STAGE 6 — CLEANUP
  □ remove old version's replicas
  □ keep the old artifact for 30 days minimum
  □ update the registry, document what changed

Stage 5 is the one that gets cut under schedule pressure, and it’s where the real quality signals appear. Defend it.


3. The signals, by detection speed#

MINUTES (automated guards, auto-rollback)
  error rate, 5xx, timeouts
  p95 TTFT, p95 ITL
  OOM / crash rate
  throughput per GPU

MINUTES-HOURS (automated alerts, human judgment)
  output length distribution        ← the most sensitive cheap signal
  stop-reason distribution (eos vs length vs stop-string)
  structured output validity rate
  refusal rate
  tool-call success rate
  prefix cache hit rate (a template change would move this)

HOURS-DAYS (human review)
  regeneration rate
  conversation length / turns to resolution
  explicit user feedback
  support ticket volume and content
  LLM-as-judge on sampled paired outputs
  human evaluation on a sample

The middle tier is what most teams lack, and it’s where quality regressions become visible first. Output length distribution in particular is nearly free to compute and moves detectably when a model starts rambling or truncating.


4. Automated rollback guards#

YAML
guards:
  # Tier 1 — automatic rollback
  - metric: error_rate_5xx
    condition: "> 1.5 × baseline for 5m"
    action: rollback
  - metric: p95_ttft_ms
    condition: "> 1.3 × baseline for 10m"
    action: rollback
  - metric: p95_itl_ms
    condition: "> 1.25 × baseline for 10m"
    action: rollback
  - metric: oom_count
    condition: "> 0 in 15m"
    action: rollback

  # Tier 2 — halt the rollout, alert a human
  - metric: mean_output_tokens
    condition: "outside [0.80, 1.25] × baseline for 30m"
    action: halt_and_alert
  - metric: structured_output_validity
    condition: "< 0.98 × baseline for 30m"
    action: halt_and_alert
  - metric: refusal_rate
    condition: "outside [0.7, 1.5] × baseline for 60m"
    action: halt_and_alert
  - metric: finish_reason_length_fraction
    condition: "> 1.4 × baseline for 30m"
    action: halt_and_alert

“Halt and alert” is the right action for Tier 2, not automatic rollback — these signals are noisier and a human should look before reverting.


5. Rollback readiness#

FAST ROLLBACK (seconds) requires:
  □ the old version still RUNNING (traffic shift only)
  □ or: the old artifact in every node cache and engine artifacts built

SLOW ROLLBACK (minutes to an hour):
  □ old artifact in regional storage
  □ redeploy, reload, warm up

DURING A ROLLOUT: keep the old version warm.
  Cost: 2x replicas for the rollout window.
  Value: seconds vs an hour of rollback time.
  → for a 4-hour progressive rollout, that's 4 GPU-hours × replicas.
    Cheap insurance.

Decide explicitly whether you’re paying for fast rollback, and for how long. “We’ll roll back if needed” is not a plan unless the old version is deployable in the time your SLO allows.


6. What to do when you find a regression post-rollout#

1. ROLL BACK FIRST, INVESTIGATE SECOND.
   Do not debug in production while users are affected.

2. PRESERVE EVIDENCE
   Sample requests and responses from the bad version (with consent /
   per your privacy policy). You need them to reproduce.

3. REPRODUCE OFFLINE
   With the exact serving configuration, not the reference implementation.

4. ROOT CAUSE
   Common causes, in order:
     - quantization degraded a capability the eval suite didn't test
     - chat template changed subtly
     - sampling defaults changed
     - a kernel/backend change altered numerics beyond tolerance
     - the checkpoint itself is different from what was evaluated

5. ADD THE TEST
   Whatever you missed goes into the evaluation suite. Permanently.

6. WRITE THE POSTMORTEM
   Including: why the existing gates didn't catch it.

Step 5 is the compounding one. An evaluation suite that grows with every incident becomes genuinely protective within a year.


7. Special cases#

ENGINE VERSION UPGRADE (not a model change)
  → shadow first (crashes and performance are the risk)
  → but ALSO validate numerics: kernel changes alter outputs
  → the model version record must include the engine version

QUANTIZATION CHANGE
  → the highest-risk change type
  → full Stage 1-5, no shortcuts
  → soak for at least 72 hours

CHAT TEMPLATE CHANGE
  → subtle and dangerous; the model is out of distribution if wrong
  → test that the rendered prompt is byte-identical for a corpus
    of conversations before and after
  → also: it invalidates the prefix cache

SAMPLING DEFAULT CHANGE
  → affects everyone who doesn't specify parameters
  → treat as a model change; canary it

EMERGENCY SECURITY PATCH
  → skip stages 2-3, go to 5% canary for 15 minutes, then full
  → accept the risk explicitly, document it, soak afterwards

8. Production implications#

  • The version record includes the whole serving configuration, not just the checkpoint.
  • Emit model_version on every metric and every response header. Correlation with deployments is otherwise impossible.
  • Build the middle-tier signals (output length, stop reasons, structured validity). They’re cheap and they’re what catches quality regressions.
  • Keep the old version warm during rollout. Seconds vs an hour of rollback.
  • Soak for 48-72 hours. The slow signals need time.
  • Grow the evaluation suite from incidents.
  • Never roll out a model change during a low-staffing window.

9. Common mistakes#

Treating a model rollout like a code rollout. Different failure mode, different timescale.

Stopping evaluation at perplexity and MMLU. They miss what users notice.

No middle-tier signals. You detect regressions from support tickets.

Cutting the soak period.

Rolling back slowly because the old version wasn’t kept deployable.

Not recording the full version. You can’t reproduce or correlate.

Debugging in production instead of rolling back first.

Not adding a test after an incident.


10. Hands-on exercise#

A. Build the version record. Implement the full version identity from Section VIII.10 for a service. Emit it as a metric label and a response header.

B. Implement the middle-tier signals. Compute and emit: output length distribution, stop reason distribution, structured output validity. Establish baselines.

C. Run a canary. Deploy a deliberately degraded version (e.g. INT4 where you normally run FP8) at 5%. Do your signals detect it? How long did it take?

D. Test the guards. Implement the automated guards from section 4. Trigger each deliberately and verify the action.

E. Rollback drill. Time a full rollback from a 100% deployment. Then time it with the old version kept warm. Quantify the difference and the cost.

F. Grow the suite. Take a past quality incident (yours or a documented one) and write the evaluation test that would have caught it.


11. Interview questions#

  1. Why can’t standard deployment safety machinery catch LLM quality regressions?
  2. What are the middle-tier signals and why do they matter?
  3. Walk me through a safe model rollout, stage by stage.
  4. What would you put in an automated rollback guard versus a halt-and-alert?
  5. Why keep the old version warm during a rollout, and what does it cost?
  6. A chat template change — why is it dangerous and how do you validate it?
  7. You find a quality regression after full rollout. What’s your sequence of actions?

12. Further reading#

  • [FUNDAMENTAL] Google SRE Book, release engineering and canarying chapters
  • [REFERENCE] Argo Rollouts / Flagger for progressive delivery
  • [ESTABLISHED] Kohavi et al., Trustworthy Online Controlled Experiments
  • Next: 08 — Observability

↑↓ navigate↵ openesc close