PidokuInfra

Multi-Region and Disaster Recovery

Advanced 1h Difficulty 3/5 Topic 06 of 11

Prerequisites 02, 05


1. Why multi-region, for LLM services specifically#

Three distinct reasons, with different implications:

1. LATENCY
   The speed of light. A user in Singapore talking to a US-East endpoint
   pays 200+ ms of RTT before any inference happens.
   → for a 800 ms TTFT SLO, that's 25% of the budget, spent on physics.

2. AVAILABILITY
   Survive the loss of a region or an availability zone.

3. DATA RESIDENCY
   Legal or contractual requirements that data not leave a jurisdiction.
   → often the actual driver, and it's non-negotiable.

For LLM services, reason 1 is stronger than for most services, because TTFT is the metric users feel and network RTT is a fixed tax on it.


2. What makes it harder than for stateless services#

STATELESS WEB SERVICE            LLM SERVICE
replicas are cheap and small     each replica is 8 GPUs and 140 GB
scale up in seconds              minutes
capacity is fungible             GPUs are scarce and region-constrained
state is in a shared database    KV cache is per-replica and ephemeral
failover = DNS change            failover = DNS change + capacity in the
                                   other region, which you must PAY FOR

The cost of multi-region is the duplicated GPU capacity, and GPUs are the expensive part. A 2-region active-active deployment with full failover capacity costs roughly 2x.


3. The deployment patterns#

PATTERN                    COST      FAILOVER    LATENCY     WHEN
Single region              1.0x      none        poor for    small scale,
                                                 distant     one market
                                                 users
Active-passive             1.5-2x    minutes     poor        compliance-driven
  (warm standby)                     to hours    (until fo)  DR requirement
Active-active,             2.0x      seconds     good        the usual answer
  each sized for full load
Active-active,             1.3-1.5x  degraded    good        cost-conscious;
  each sized for 65%                                          accept degradation
                                                              during failover
Multi-region with          varies    seconds     good        large scale
  overflow routing

“Active-active, each sized for 65%” is the pragmatic middle ground: both regions serve traffic normally at 65% utilization; if one fails, the other runs at 130% of its comfortable load — degraded (higher latency, some shedding) but not down.

Document that degradation explicitly as part of your DR plan.


4. Routing#

LATENCY-BASED DNS / ANYCAST
  Route users to the nearest healthy region.
  ✓ simple, works with standard infrastructure
  ✗ DNS TTL means failover takes minutes
  ✗ doesn't account for regional capacity

GLOBAL LOAD BALANCER (application layer)
  ✓ health-aware, capacity-aware, instant failover
  ✓ can implement overflow ("if EU is at 90%, send 10% to US")
  ✗ another component to run

HYBRID (the usual answer)
  latency-based DNS for the primary route
  + application-layer overflow between regions
  + health checks that remove a region quickly

LLM-specific consideration: session affinity. A multi-turn conversation should stay in one region, both for prefix cache locality (Section V.11) and for consistency. Route by conversation ID within a region, and only move sessions on failover.


5. What has to be replicated#

MUST BE IN EVERY REGION
  model weights (in regional object storage + node caches)
  the serving stack (container images)
  configuration
  the model registry (or a regional read replica)

SHOULD BE REGIONAL, NOT REPLICATED
  KV cache (ephemeral; never replicate)
  in-flight request state
  local caches

MUST BE GLOBALLY CONSISTENT (or carefully partitioned)
  rate limit / quota state    ← the hard one
  billing records
  auth tokens
  audit logs

Quota state across regions is genuinely hard. Options:

1. Partition quota by region (user's quota is 1/N per region)
   ✓ simple, no coordination
   ✗ a user hitting one region gets 1/N of their quota

2. Global quota store with eventual consistency
   ✓ correct enough; a user can briefly exceed by the sync window
   ✗ requires a cross-region store; adds latency

3. Local enforcement with periodic reconciliation
   ✓ fast, no cross-region latency in the request path
   ✗ over-consumption is detected after the fact
   → THE USUAL ANSWER: enforce locally with a generous local budget,
     reconcile centrally, and correct on the next window.

6. Model weight distribution#

Getting 140 GB of weights into every region:

  1. Push to regional object storage (S3 cross-region replication,
     or an explicit pipeline)
  2. Pre-populate node-local caches (DaemonSet, or bake into an AMI/image)
  3. Verify checksums per region

TIMING
  A model release must complete in all regions before you shift traffic.
  For a 140 GB model across 4 regions: hours, mostly transfer time.
  → build this into your release timeline.

COST
  Cross-region egress is expensive. Push once to each region's storage
  and pull locally, rather than pulling cross-region per node.

A common and expensive mistake: nodes in region B pulling weights from region A’s object storage. Egress charges for 140 GB × 50 nodes × every deployment adds up quickly.


7. The DR plan#

Write it down, with numbers:

SCENARIO: loss of region us-east-1

  DETECTION
    health checks fail for > 60 s, or error rate > 50% for > 30 s
    → automated: global LB removes the region

  IMPACT
    45% of traffic must move to us-west-2 and eu-west-1
    us-west-2 goes from 65% → 105% utilization → degraded
    eu-west-1 goes from 60% → 88% utilization → acceptable

  DEGRADATION (expected and accepted)
    p95 TTFT: 780 ms → 1,900 ms for affected users
    free-tier traffic shed: ~15%
    premium tier: unaffected (reserved capacity)

  RECOVERY
    RTO (time to serve degraded): 90 seconds (DNS + LB health check)
    RTO (time to serve normally): 45-90 minutes (spin up additional
        capacity in the surviving regions, IF quota is available)
    RPO: not applicable — inference is stateless; in-flight requests
         are lost

  ACTIONS
    1. automated: LB removes the region
    2. automated: degradation ladder engages (Section XI.04)
    3. manual: request additional capacity in surviving regions
    4. manual: communicate degraded status
    5. on recovery: gradual traffic return (10% steps) to avoid
       thundering herd on cold replicas

Step 5 matters. Returning 45% of traffic instantly to a region whose replicas just started produces a cold-start stampede. Ramp it.


8. Cost management#

Multi-region is expensive. Ways to reduce the cost:

1. ASYMMETRIC SIZING
   Size regions by their traffic, not equally.
   
2. SHARED OVERFLOW CAPACITY
   One region has extra capacity that any region can overflow into,
   rather than every region having its own N+1.

3. SPOT / PREEMPTIBLE for the overflow tier
   Cheaper, and you only need it during failover.
   
4. TIER-AWARE FAILOVER
   Only premium traffic fails over; free tier is shed.
   → dramatically reduces the capacity you must hold.

5. EXTERNAL PROVIDER AS THE DR TARGET
   Instead of holding idle GPUs, contract with an API provider for
   burst capacity. Expensive per token, cheap as insurance.

Option 4 is the most under-used. Holding full failover capacity for free-tier traffic is usually not worth it; shedding it during a regional failure is an acceptable degradation.


9. Production implications#

  • Decide the driver first: latency, availability, or compliance. They imply different architectures.
  • Size for degraded operation, not full failover, unless you can afford 2x.
  • Document the expected degradation numerically as part of the DR plan.
  • Keep sessions regional for prefix cache locality.
  • Enforce quota locally with central reconciliation.
  • Distribute weights regionally; never pull cross-region per node.
  • Test failover. A DR plan that has never been executed is a document, not a capability.
  • Ramp traffic back gradually after recovery.

10. Common mistakes#

Multi-region for availability when latency is the actual need (or vice versa) — leading to the wrong architecture.

Sizing every region for 100% of total load. 3x cost for a 3-region deployment.

Cross-region weight pulls. Large egress bills.

Global quota checks in the request path. Adds cross-region latency to every request.

Sessions moving between regions. Cold prefix cache, inconsistent behavior.

Instant traffic return after recovery. Cold-start stampede.

Never testing failover.

Not documenting the expected degradation, so an incident becomes a surprise.


11. Hands-on exercise#

A. Compute the latency benefit. For your user distribution, compute the p50 and p95 network RTT to a single region versus to the nearest of N regions. What fraction of your TTFT budget does each save?

B. Size the regions. For a given traffic distribution and a “survive one region loss with degradation” requirement, compute the capacity needed in each region. Compare to full-failover sizing.

C. Write the DR plan. Produce the document from section 7 for a service you know, with real numbers.

D. Quota design. Design cross-region quota enforcement. Implement local enforcement with periodic reconciliation and measure the over-consumption window.

E. Test it. In a test environment, simulate a regional failure. Measure: detection time, traffic shift time, degradation experienced, and recovery time. Compare to your plan’s RTO.


12. Interview questions#

  1. What are the three reasons for multi-region, and how do they differ in architecture?
  2. Why is multi-region more expensive for LLM services than for stateless ones?
  3. How would you size regions to survive one region’s loss without 2x cost?
  4. How do you handle rate limit quota across regions?
  5. Why should sessions stay in one region?
  6. What’s your RTO and RPO for an inference service, and why is RPO unusual here?
  7. Why ramp traffic back gradually after a region recovers?

13. Further reading#

  • [FUNDAMENTAL] Google SRE Book, chapters on reliability and disaster recovery
  • [REFERENCE] Cloud provider multi-region architecture guidance
  • [FUNDAMENTAL] Grigorik, High Performance Browser Networking — the latency argument
  • Next: 07 — Model rollouts

↑↓ navigate↵ openesc close