PidokuInfra

Cascades and Fallbacks

Expert Advanced 1h Difficulty 3/5 Topic 09 of 11

Prerequisites 06, XI.03


CASCADE     try a cheap model first; escalate to an expensive one only
            when needed.
            → a COST optimization

FALLBACK    when the primary path fails or is overloaded, use an
            alternative.
            → an AVAILABILITY mechanism

They use similar machinery and are often confused. Build both; know which you’re building.


2. Cascades — the cost argument#

OBSERVATION: most queries are easy.

  "What's the capital of France?"        → an 8B model answers correctly
  "Prove that √2 is irrational and
   explain each step"                    → needs the 70B

If 70% of queries can be handled by a model 10x cheaper:
  cost = 0.70 × C_small + 0.30 × C_large + escalation_overhead
       = 0.70 × 0.1 + 0.30 × 1.0 + overhead
       = 0.07 + 0.30 + overhead
       = ~0.40 × the cost of using the large model for everything
       
  → 60% cost reduction, IF the routing decision is good.

The whole thing hinges on the routing decision. A cascade that escalates 80% of queries saves nothing and adds latency.


3. How to decide whether to escalate#

STRATEGY 1 — ROUTE UP FRONT (a classifier decides before generating)
  a small classifier predicts which model is needed
  ✓ no wasted generation
  ✓ no added latency for escalated queries
  ✗ needs training data (query → which model sufficed)
  ✗ the classifier can be wrong in both directions

STRATEGY 2 — GENERATE THEN CHECK (the small model answers; a check
  decides whether to escalate)
  ✓ the check can use the actual answer
  ✗ escalated queries pay BOTH costs and BOTH latencies
  ✗ the check is itself a hard problem

STRATEGY 3 — CONFIDENCE-BASED
  use the small model's own signal: mean token logprob, entropy,
  or an explicit "I'm not sure" detection
  ✓ free (you have the logprobs)
  ✗ LLM confidence is poorly calibrated — a confidently wrong answer
    looks confident
  → use with caution; validate the correlation on your task

STRATEGY 4 — HEURISTIC
  prompt length, presence of code, task type from the system prompt,
  explicit user tier
  ✓ trivial, transparent, debuggable
  ✗ coarse
  → START HERE. It captures a surprising amount of the benefit.

Strategy 4 first, then 1. Heuristics are transparent and you can measure their accuracy before investing in a classifier.

Go
func chooseModel(r *Request) string {
	// Strategy 4: heuristics, in order of confidence
	switch {
	case r.Tenant.Tier == "premium":
		return "large"
	case r.Requires("code_execution"):
		return "large"
	case len(r.TokenIDs) > 8000:
		return "large" // long context → likely complex
	case knownSimpleTasks[r.SystemPromptHash]:
		return "small"
	case slices.Contains([]string{"classification", "extraction", "simple_qa", "translation"}, detectTaskType(r)):
		return "small"
	}
	return "large" // default to quality
}

Default to the large model. A cascade that defaults to small and escalates on failure degrades quality; one that defaults to large and routes down on confidence preserves it.


4. Measuring a cascade#

YOU MUST MEASURE
  escalation rate           what fraction goes to the large model?
  cost per request          the actual saving
  quality delta             on the queries handled by the small model
  latency delta             especially for escalated queries (strategy 2)
  routing accuracy          how often did the small model suffice
                            when we sent it to large, and vice versa?

THE VALIDATION
  Sample N queries. Run BOTH models on all of them.
  Have a judge (human or LLM) determine which answers are acceptable.
  → gives you the ceiling (what fraction COULD go small)
  → and the routing accuracy against that ceiling

Without this measurement you cannot tell a good cascade from a quality regression that happens to be cheaper. It’s the same discipline as validating quantization (Section IV.12).


5. Fallbacks — the availability mechanism#

FALLBACK LADDER (in order)

1. RETRY ON ANOTHER REPLICA
   the primary replica failed or is saturated
   ✓ same model, same quality
   ✗ only helps if another replica has capacity

2. FALL BACK TO ANOTHER REGION
   the local region is degraded
   ✗ higher latency
   → Section XI.06

3. FALL BACK TO A SMALLER MODEL
   the large model's pool is saturated
   ✗ quality degradation, but a response
   → often better than an error

4. FALL BACK TO AN EXTERNAL PROVIDER
   ✗ expensive, data leaves your infrastructure
   ✓ far cheaper than an outage
   → requires a contract and a data-handling decision IN ADVANCE

5. FAIL WITH A USEFUL ERROR
   503 with Retry-After
   → honest, and better than a timeout
Go
func generateWithFallback(ctx context.Context, r *Request) (*Result, error) {
	for _, attempt := range fallbackChain {
		if !attempt.Enabled || !attempt.AllowedFor(r.Tenant) {
			continue
		}
		actx, cancel := context.WithTimeout(ctx, attempt.Timeout)
		result, err := attempt.Execute(actx, r)
		cancel()

		switch {
		case err == nil:
			if attempt.Degraded {
				fallbackUsed.WithLabelValues(attempt.Name).Inc()
				result.Metadata["degraded"] = attempt.Name
			}
			return result, nil
		case errors.Is(err, ErrInvalidRequest):
			return nil, err // don't fall back on client errors!
		case errors.Is(err, ErrSaturated), errors.Is(err, ErrUnavailable), errors.Is(err, context.DeadlineExceeded):
			fallbackAttempted.WithLabelValues(attempt.Name, err.Error()).Inc()
			continue // try the next level
		default:
			return nil, err
		}
	}
	return nil, &ServiceUnavailable{RetryAfter: 30 * time.Second}
}

raise on client errors is important. A malformed request should not cascade through your entire fallback chain, consuming capacity at each level.


6. Telling the client#

Should the client know it got a degraded response?

ARGUMENTS FOR
  ✓ they can decide whether to accept it or retry later
  ✓ honesty; they may be paying for the premium model
  ✓ debuggability

ARGUMENTS AGAINST
  ✗ leaks internal state
  ✗ clients may implement their own retry logic that makes things worse

PRACTICAL ANSWER
  include it in the response metadata, not as an error:
    X-Model-Served: llama-3-8b        (they asked for chat-large)
    X-Degraded: capacity
  → the client CAN see it; most won't look; those who care can act.

If a tenant is paying for a specific model, silently serving a different one is a contractual problem, not just a technical one. Decide the policy explicitly and document it.


7. The interaction with quality SLOs#

A cascade or fallback CHANGES THE MODEL. That means:
  - the quality SLO applies to the composite, not to one model
  - your evaluation must cover the cascade, not just each model
  - a change in escalation rate changes aggregate quality

MEASURE
  aggregate quality (weighted by the actual routing distribution)
  quality per path
  escalation rate over time (drift!)

Escalation rate drift is a real failure mode: a product change shifts the query distribution, the escalation rate falls from 30% to 10%, cost drops, and quality quietly degrades. Alert on escalation rate as a quality signal.


8. Production implications#

  • Start with heuristic routing. Transparent, debuggable, captures much of the benefit.
  • Default to the larger model and route down, not up.
  • Measure the ceiling by running both models on a sample. You need to know what’s achievable.
  • Alert on escalation rate drift — it’s a quality signal.
  • Fallbacks are for availability; cascades are for cost. Build both, keep them distinct.
  • Don’t fall back on client errors.
  • Expose the served model in the response metadata.
  • Decide the external-provider fallback policy in advance, including data handling.
  • Evaluate the cascade as a system, not the models individually.

9. Common mistakes#

Cascading without measuring quality. A cost saving that’s actually a quality regression.

Defaulting to the small model. Quality degrades on the queries you routed wrong.

Using LLM confidence as the escalation signal without validating it. Poorly calibrated.

Strategy 2 (generate then check) without accounting for the double cost on escalated queries.

Falling back on client errors. Wastes capacity across the whole chain.

Silent model substitution for a paying customer.

Not alerting on escalation rate. Drift goes unnoticed.

One evaluation per model, none for the cascade.


10. Hands-on exercise#

A. Measure the ceiling. Take 500 real queries. Run an 8B and a 70B model on each. Have a judge determine for which queries the 8B answer is acceptable. What fraction? That’s your ceiling.

B. Build the heuristic router. Implement strategy 4 for your task. Measure its accuracy against the ceiling from A: how often does it route correctly in each direction?

C. Compute the saving. With your measured escalation rate and per-model costs, compute the actual cost reduction. Compare to the theoretical maximum from A.

D. Test the confidence signal. For the queries where the small model was wrong, was its mean logprob lower? Plot the distributions. Is confidence a usable signal for your task?

E. Build the fallback chain. Implement the chain from section 5. Test each level by simulating the corresponding failure. Verify client errors don’t cascade.

F. Drift detection. Simulate a query distribution shift and verify your escalation-rate alert fires.


11. Interview questions#

  1. What’s the difference between a cascade and a fallback?
  2. How would you decide whether to escalate a query?
  3. Why default to the large model rather than the small one?
  4. Why is LLM confidence a poor escalation signal?
  5. How would you measure whether a cascade is working?
  6. Why must you not fall back on client errors?
  7. Should the client know it got a degraded response? Justify.

12. Further reading#

  • [ESTABLISHED] Chen et al., “FrugalGPT” (2023) — LLM cascades
  • [EMERGING] Ong et al., “RouteLLM” (2024) — learned routing between models
  • [FUNDAMENTAL] Circuit breaker and fallback patterns (Nygard, Release It!)
  • Next: 10 — Cluster and capacity management

↑↓ navigate↵ openesc close