PidokuInfra

Fault Tolerance and Failure Modes

Advanced 1h 15m Difficulty 4/5 Topic 05 of 11

Prerequisites VIII.05, IX.10


1. The failure modes, catalogued#

FAILURE                        FREQUENCY   BLAST RADIUS      DETECTION
GPU Xid error / fault          weeks       1 GPU → 1 TP group  dmesg, DCGM
GPU falls off the bus          months      1 node              nvidia-smi fails
ECC uncorrectable error        months      1 GPU               DCGM, dmesg
Driver hang                    months      1 node              all CUDA calls hang
NCCL hang (one slow rank)      weeks       1 TP/PP group       timeout
OOM                            days-weeks  1 replica           exception or SIGKILL
Model load failure             per deploy  1 replica           readiness never passes
Network partition (multi-node) months      1 instance          NCCL error
Node preemption (spot)         hours       1 node              cloud signal
Deployment error               per deploy  everything          canary should catch
Quality regression             per deploy  everything          hard (Section VIII.10)
Upstream dependency failure    varies      varies              circuit breaker
Traffic spike beyond capacity  weekly      everything          queue wait

The two that are LLM-specific and most damaging: NCCL hangs and quality regressions. Both are hard to detect and both hold resources while broken.


2. The design principles#

1. FAIL FAST, NOT SLOW.
   A hung process holding 8 GPUs is worse than a crashed one.
   → timeouts everywhere, especially NCCL.

2. LIMIT THE BLAST RADIUS.
   DP replicas fail independently; TP groups fail together.
   → prefer more, smaller instances where latency permits.

3. DEGRADE, DON'T COLLAPSE.
   Shed load explicitly rather than queueing until failure.

4. MAKE FAILURE VISIBLE.
   A silently degraded replica serving slowly is worse than a dead one,
   because the load balancer keeps sending it traffic.

5. RECOVER AUTOMATICALLY, BUT BOUNDED.
   Restart, but with backoff and a crashloop limit.

6. PRESERVE WHAT YOU CAN.
   In-flight requests should fail cleanly with a clear error,
   not hang until the client times out.

3. GPU failures#

DETECTION
  dmesg -T | grep -i xid
  nvidia-smi --query-gpu=ecc.errors.uncorrected.aggregate.total --format=csv
  DCGM_FI_DEV_XID_ERRORS

COMMON XID CODES
  13, 31    illegal memory access (usually a software bug, not hardware)
  43        stopped processing (software)
  48        double-bit ECC error (HARDWARE — replace)
  63, 64    ECC page retirement (hardware degrading)
  74        NVLink error (hardware or cable)
  79        GPU has fallen off the bus (HARDWARE — node is unusable)
  94, 95    contained/uncontained ECC error

RESPONSE
  software Xids (13, 31, 43)  → restart the process; investigate the bug
  hardware Xids (48, 79, 94)  → CORDON THE NODE, drain, replace the GPU
  degrading (63, 64)          → schedule replacement; monitor

Automate the cordon. A node with an uncorrectable ECC error will keep failing; leaving it in the pool produces repeated incidents.

YAML
# NVIDIA GPU Operator's node-problem-detector can taint nodes automatically
# Or: a DaemonSet that watches DCGM_FI_DEV_XID_ERRORS and applies a taint

4. NCCL hangs — the LLM-specific one#

SYMPTOM
  A TP or PP group stops making progress. No error. No crash.
  All ranks are in a collective, waiting for one that never arrives.
  8 or 16 GPUs are held indefinitely.

CAUSES
  one rank crashed (the others wait forever)
  one rank is descheduled or throttled (Section II.12)
  a network partition
  a driver issue on one node
  a deadlock from mismatched collective ordering (a code bug)

DETECTION
  NCCL_TIMEOUT (set it! default is very long)
  TORCH_NCCL_ASYNC_ERROR_HANDLING=1
  a watchdog: if no step completed in N seconds, kill the process
  externally: queue depth growing while throughput is zero

RESPONSE
  1. Kill ALL ranks of the group (NCCL has no partial recovery)
  2. Restart the group
  3. If it recurs on the same node, cordon that node
Go
// StepWatchdog is an application-level watchdog — necessary because NCCL timeouts
// don't always fire, and a hang is much worse than a crash.
type StepWatchdog struct {
	last    atomic.Int64 // unix nanoseconds of the last engine step
	timeout time.Duration
}

func NewStepWatchdog(timeout time.Duration) *StepWatchdog {
	w := &StepWatchdog{timeout: timeout}
	w.Beat()
	go func() {
		for range time.Tick(10 * time.Second) {
			if idle := time.Since(time.Unix(0, w.last.Load())); idle > w.timeout {
				slog.Error("watchdog: no engine step, aborting", "idle", idle)
				os.Exit(1) // hard exit; let the supervisor restart us
			}
		}
	}()
	return w
}

// Beat is called by the engine loop after every step.
func (w *StepWatchdog) Beat() { w.last.Store(time.Now().UnixNano()) }

os._exit(1) rather than a graceful shutdown is correct here: a hung process may not be able to shut down gracefully, and the supervisor’s restart is the recovery path.


5. Handling in-flight requests during failure#

WHEN A REPLICA DIES
  its in-flight requests are lost. The clients see:
    - a connection reset (non-streaming)
    - a truncated stream (streaming) ← worse; the client got partial output

MITIGATIONS
  1. Fail cleanly: send an in-band error before dying, if possible
       data: {"error": {"message": "...", "type": "server_error"}}
  2. Client-side: detect truncation (no [DONE] received) and surface it
  3. Do NOT auto-retry a streamed request — the user has partial output
  4. For non-streaming, a retry is safe IF the original is confirmed dead

There is no way to migrate an in-flight generation to another replica without transferring its KV cache — which is possible in principle (Section IX.11) but not implemented in production systems for failure recovery. Accept the loss and handle it cleanly.


6. Circuit breakers and dependency failures#

An inference gateway typically depends on:
  auth service
  rate limit store (Redis)
  model registry
  the engines themselves
  (sometimes) a safety classifier, a retrieval service

FOR EACH: what happens if it's down?

  auth down        → fail closed (reject) or fail open (allow)?
                     SECURITY DECISION. Usually fail closed, with a
                     short-lived cache to ride out brief outages.
  rate limiter     → fail OPEN (allow), with a conservative local limit.
                     Rejecting all traffic because Redis is down is worse
                     than briefly allowing over-limit traffic.
  registry         → cache aggressively; it changes rarely.
  safety filter    → fail closed for high-risk content, or degrade to a
                     cheaper local check. A policy decision.
  an engine        → circuit-break it, route elsewhere.
Go
var ErrCircuitOpen = errors.New("circuit open")

type CircuitBreaker struct {
	mu        sync.Mutex
	failures  int
	threshold int           // consecutive failures before opening, e.g. 5
	cooldown  time.Duration // how long to stay open, e.g. 30s
	openedAt  time.Time
}

func (cb *CircuitBreaker) Call(fn func() error) error {
	cb.mu.Lock()
	if !cb.openedAt.IsZero() && time.Since(cb.openedAt) < cb.cooldown {
		cb.mu.Unlock()
		return ErrCircuitOpen // fail fast: do not even try the unhealthy replica
	}
	cb.mu.Unlock()

	err := fn()

	cb.mu.Lock()
	defer cb.mu.Unlock()
	if err == nil {
		cb.failures, cb.openedAt = 0, time.Time{}
		return nil
	}
	if cb.failures++; cb.failures >= cb.threshold {
		cb.openedAt = time.Now()
	}
	return err
}

Write down the fail-open/fail-closed decision for every dependency. It is a decision, and if you don’t make it deliberately the code makes it accidentally.


7. Testing failure#

GAME DAYS — practice these deliberately:

1. Kill one replica under load.
   Expect: LB removes it, in-flight requests fail cleanly, no cascade.

2. Kill one rank of a TP group.
   Expect: the group hangs, the watchdog fires within N seconds,
           the supervisor restarts all ranks, readiness gates traffic.

3. Fill the KV cache (send max-length requests).
   Expect: preemption, then admission rejection with 503. No OOM.

4. Saturate the network between nodes.
   Expect: NCCL slows, the watchdog eventually fires, or throughput
           degrades gracefully.

5. Stop the rate-limit store.
   Expect: fail-open with a local fallback limit, alarm fires.

6. Deploy a deliberately broken model version.
   Expect: readiness never passes, the rollout halts, no traffic shifted.

7. Exceed capacity by 3x.
   Expect: 503s with Retry-After, served requests keep their SLO,
           recovery when load drops (no metastable state).

Most teams have never tested #2 or #7, and both are the ones that cause real incidents.


8. Production implications#

  • Set NCCL_TIMEOUT and an application watchdog. Hangs are worse than crashes.
  • Automate node cordoning on hardware Xid errors.
  • Prefer more, smaller instances where latency permits — smaller blast radius.
  • Write down fail-open/fail-closed for every dependency.
  • Handle in-flight request loss cleanly. Don’t leave clients hanging.
  • Run game days. Quarterly, on the seven scenarios above.
  • Track MTTR per failure mode, not just availability.
  • Crashloop limits: a replica that has restarted 5 times in 10 minutes should stay down and alert, not keep cycling.

9. Common mistakes#

No NCCL timeout. A hang holds 8-16 GPUs indefinitely.

Restarting one rank of a TP group. The others still hang.

No watchdog. Relying on NCCL’s timeout alone is insufficient.

Leaving failed nodes in the pool. Repeated incidents from the same hardware.

Auto-retrying streamed requests. Duplicate partial output.

Fail-closed on the rate limiter. Total outage because Redis blipped.

Never testing failure. The first time you exercise the path is during an incident.

Unbounded crashloops. A replica cycling forever, consuming GPUs and generating noise.


10. Hands-on exercise#

A. Implement the watchdog. Add the step watchdog from section 4 to an engine. Test it by suspending a rank (kill -STOP) and verifying the watchdog fires.

B. The Xid response. Write a DaemonSet or script that watches for hardware Xid errors and cordons the node. Test with a simulated error.

C. Dependency matrix. For a service you know, list every dependency and write the fail-open/fail-closed decision with justification. Implement circuit breakers.

D. Game day. Run all seven scenarios from section 7 on a test environment. Document what actually happened versus what you expected. Fix the gaps.

E. Clean failure. Implement in-band error reporting for a streaming endpoint. Kill the engine mid-generation and verify the client receives a structured error rather than a hang.

F. Blast radius. Compare the impact of a single GPU failure in a TP=8 deployment versus 8 independent replicas. Quantify in requests affected.


11. Interview questions#

  1. What happens when one rank of a TP group fails? How do you handle it?
  2. Why is a hang worse than a crash for a GPU workload?
  3. What is a hardware Xid error and how should you respond?
  4. Should the rate limiter fail open or closed? Justify.
  5. Can you migrate an in-flight generation to another replica? Why or why not?
  6. Design a game day for an LLM inference service.
  7. How does parallelism strategy affect blast radius?

12. Further reading#

  • [REFERENCE] NVIDIA Xid error documentation
  • [REFERENCE] NVIDIA GPU Operator and node-problem-detector
  • [FUNDAMENTAL] Google SRE Book, “Addressing Cascading Failures”
  • [FUNDAMENTAL] Netflix’s chaos engineering principles
  • Next: 06 — Multi-region and disaster recovery

↑↓ navigate↵ openesc close