PidokuInfra

Routing Strategies

Expert Advanced 1h 15m Difficulty 4/5 Topic 06 of 11

Prerequisites VIII.06, 05

Section VIII.06 covered load balancing between replicas of one model. This file covers the platform’s routing layer, which also decides which model and which pool.


1. The routing decisions#

A platform router makes several decisions per request, in order:

1. WHICH MODEL?          alias resolution, cascades, fallbacks
2. WHICH POOL?           general / long-context / premium / batch
3. WHICH REPLICA?        load-aware, prefix-aware, session-affine
4. ADMIT OR REJECT?      quota, capacity, priority

Each has different inputs and different failure modes.


2. Decision 1 — which model#

ALIAS RESOLUTION
  client asks for "chat-large" → registry says llama-3-70b-instruct v1.5.0
  → decouples client-facing names from deployments (Section XII.02)

VERSION SELECTION
  canary: 5% → v1.6.0, 95% → v1.5.0, sticky by conversation_id
  (Section VIII.10)

COST-AWARE / CASCADE ROUTING
  route by request characteristics:
    short, simple query        → the 8B model
    complex reasoning          → the 70B model
    code                       → the code-specialized model
  → Section XII.09

TENANT-SPECIFIC
  tenant has a fine-tuned adapter → route to a replica with that LoRA
  tenant is restricted to certain models → enforce here

3. Decision 2 — which pool#

POOL SEGREGATION (Section VIII.06 — one of the highest-value decisions)

  general        the bulk of traffic; tuned for typical request shapes
  long-context   requests > 16k tokens; different max_model_len,
                 different max_num_seqs, isolated so they don't
                 poison the general pool's batch
  premium        reserved capacity, lower batch size, better ITL
  batch          no latency SLO; huge batch; may use spot instances
  canary         the version under test

ROUTING RULES
  if request.prompt_tokens > 16000:       → long-context pool
  elif tenant.tier == "premium":          → premium pool
  elif request.priority == "batch":       → batch pool
  elif canary_selected(request):          → canary pool
  else:                                   → general pool

Segregating long-context requests is worth calling out again. A single 128k request in a pool tuned for 4k consumes 32 sequences’ worth of KV and blocks prefill for everyone. It’s a one-line routing rule that protects your p99.


4. Decision 3 — which replica#

Covered in Section VIII.06; the platform version adds the multi-model dimension:

def choose_replica(request, model, pool):
    candidates = routing_table[model][pool]
    candidates = [r for r in candidates if r.ready and not r.draining]

    # 1. Session affinity (for prefix cache locality)
    if request.conversation_id:
        preferred = consistent_hash(request.conversation_id, candidates)
        if preferred.kv_usage < AFFINITY_RELEASE_THRESHOLD:   # e.g. 0.85
            return preferred

    # 2. Prefix affinity (for shared system prompts)
    prefix_key = hash(request.token_ids[:PREFIX_WINDOW])
    warm = [r for r in candidates if prefix_key in r.known_prefixes
                                  and r.kv_usage < AFFINITY_RELEASE_THRESHOLD]
    if warm:
        return min(warm, key=lambda r: r.kv_usage)

    # 3. Load-aware: power of two choices
    a, b = random.sample(candidates, min(2, len(candidates)))
    chosen = a if a.load_score() < b.load_score() else b
    remember_prefix(prefix_key, chosen)
    return chosen

def load_score(r):
    # combine the signals; KV usage dominates
    return (0.6 * r.kv_usage
          + 0.3 * (r.queue_depth / r.max_queue)
          + 0.1 * (r.running_batch / r.max_batch))

The AFFINITY_RELEASE_THRESHOLD is the crucial tuning knob. Too high and a popular prefix overloads one replica; too low and you lose cache locality. 0.80-0.90 is a reasonable range; tune by measuring end-to-end TTFT.


5. Decision 4 — admit or reject#

The router's admission decision is DIFFERENT from the engine's:

ROUTER-LEVEL
  is the tenant over quota?              → 429
  is EVERY replica of this model at capacity? → 503
  is the estimated wait > the client's deadline? → 503
  is the request malformed / over limits? → 400

ENGINE-LEVEL (Section VIII.03)
  is there KV space right now?
  does this fit in the token budget for this step?

The router should reject early (cheap) but the engine has the authoritative state. Both are needed.

The router's view of capacity is STALE (metrics scraped every 1-5 s).
→ it should be conservative, and the engine must still be able to reject.
→ never let the router promise what the engine can't deliver.

6. The routing table#

STRUCTURE
  model_alias → model_id, version(s), weights
  model_id + version → pools
  pool → [replica endpoints, capabilities, current load]

DISTRIBUTION (Section XII.01)
  control plane computes it, pushes to routers
  routers cache it with a TTL and last-known-good behavior
  load metrics scraped separately, more frequently

FRESHNESS REQUIREMENTS
  routing table (which replicas exist):  seconds to minutes — slow is fine
  load metrics (how busy):               1-5 seconds — matters
  prefix location hints:                 seconds; approximate is fine

Load metrics need to be fresher than the routing table, and they’re the thing that must keep flowing. A router with a 5-minute-old routing table but current load data works fine; the reverse does not.


7. Routing for multi-LoRA#

With multi-LoRA (Section VIII.09), routing gains a dimension:

  request specifies adapter "team-x-v3"
  → prefer a replica that already has it loaded (avoids a load)
  → and concentrate requests for the same adapter on the same replica
    (fewer distinct adapters per batch → lower grouped-GEMM overhead)

def choose_replica_lora(request, candidates):
    adapter = request.lora_id
    loaded = [r for r in candidates if adapter in r.loaded_adapters]
    if loaded:
        return min(loaded, key=load_score)
    # not loaded anywhere: pick the least loaded replica with room
    # for another adapter
    room = [r for r in candidates if len(r.loaded_adapters) < r.max_loras]
    return min(room or candidates, key=load_score)

Adapter affinity matters more than you’d expect: the grouped-GEMM overhead grows with the number of distinct adapters in a batch, so concentrating them is worth some load imbalance.


8. Failure handling in the router#

REPLICA UNHEALTHY
  remove from candidates. But: distinguish
    - not ready (loading, warming) → exclude
    - saturated (KV at 100%)        → DEPRIORITIZE, don't exclude
                                      (excluding cascades — Section VIII.06)
    - failing health checks         → exclude, alert
    - draining                      → exclude for new requests

ALL REPLICAS SATURATED
  → 503 with Retry-After, don't queue at the router
    (queueing at the router hides the problem from the engine's
     scheduler, which has better information)

MODEL NOT DEPLOYED
  → if it's a cold-tier model: trigger a load, queue the request,
    tell the client to expect latency (or 503 with a long Retry-After)
  → else: 404

CONTROL PLANE UNREACHABLE
  → keep using the cached routing table (Section XII.01)
  → alert, but keep serving

“Deprioritize, don’t exclude” for saturated replicas is the rule that prevents cascading failure: excluding them concentrates load on the remaining replicas, which then saturate.


9. Production implications#

  • The router is where LLM-specific logic lives. Build it; don’t use a generic LB (Section XII.01).
  • Segregate long-context requests into their own pool. Single highest-value routing rule.
  • Session affinity by conversation ID, with a load release valve.
  • Load metrics must be fresh (1-5 s). Stale load data makes routing worse than random.
  • Never exclude saturated replicas; deprioritize them.
  • Don’t queue at the router. Let the engine’s scheduler decide.
  • Adapter affinity for multi-LoRA.
  • Instrument routing decisions: which replica, why, and what the alternatives were. When routing goes wrong you need to see the decision.

10. Common mistakes#

Generic load balancer with round-robin. Wrong on every axis (Section VIII.06).

Excluding saturated replicas. Cascading failure.

Queueing at the router. Hides state from the scheduler and adds a second queue.

Stale load metrics. Routing on 60-second-old data is worse than random.

No session affinity with prefix caching enabled. 1/N hit rate.

Affinity without a release valve. One popular conversation overloads a replica.

No long-context segregation. One request poisons the pool.

Global least-loaded across router instances. Herding (Section VIII.06).


11. Hands-on exercise#

A. Build the router. Implement the four decisions from section 1. Support: aliases, pool selection, session affinity, prefix affinity, load-aware selection with power-of-two, and admission. This is a core piece of Project 15.

B. Measure the affinity threshold. Sweep AFFINITY_RELEASE_THRESHOLD from 0.5 to 1.0 in a simulation with a popular shared prefix. Plot prefix hit rate and p99 TTFT. Find the optimum.

C. Demonstrate the cascade. Simulate excluding saturated replicas versus deprioritizing them under increasing load. Show the cascade in the first case.

D. Metric freshness. Simulate routing with load metrics that are 1 s, 5 s, 30 s, and 300 s stale. At what staleness does load-aware routing become worse than random?

E. Long-context segregation. Simulate a pool with 2% of requests at 64k context, with and without segregation. Measure p99 TTFT for the 98%.

F. LoRA affinity. Simulate 20 adapters across 4 replicas with and without adapter affinity. Measure the average number of distinct adapters per batch.


12. Interview questions#

  1. What decisions does a platform router make, and in what order?
  2. Why segregate long-context requests, and what does it protect?
  3. Why deprioritize rather than exclude a saturated replica?
  4. Why shouldn’t the router queue requests?
  5. How fresh must load metrics be, and what happens when they’re stale?
  6. How does multi-LoRA change replica selection?
  7. What’s the tension between cache affinity and load balance, and how do you resolve it?

13. Further reading#

  • [FUNDAMENTAL] Mitzenmacher, “The Power of Two Choices”
  • [ESTABLISHED] Kubernetes Gateway API Inference Extension — InferencePool is a stable (v1) API; the endpoint picker now lives in llm-d
  • [ESTABLISHED] SGLang’s router implementation
  • [ESTABLISHED] Envoy load balancing documentation for the general patterns
  • Next: 07 — Prefix-aware routing

↑↓ navigate↵ openesc close