1. What is it?#
The three mechanisms that determine whether your service degrades gracefully or collapses.
BACKPRESSURE signalling upstream that you're full, so it slows down
TIMEOUTS giving up on work that has taken too long
RETRIES trying again after a failure — and the discipline not toFor LLM serving, all three need different settings and different reasoning than for ordinary services, because requests are long-lived, expensive, and non-idempotent.
2. Why it exists#
Because the default behavior of a naive system under overload is not “slower” — it’s collapse, and often a collapse it cannot recover from.
Overload → queue grows → latency grows → clients time out → clients retry
→ load increases → queue grows more → ...
And critically, in LLM serving:
the original request is STILL RUNNING, holding a KV slot, generating tokens
nobody will read.The system is now doing 3x the work for 0x the value. Arrival rate returning to normal doesn’t help, because the system is busy with abandoned work. This is a metastable failure.
3. Simple analogy#
A ticket office with a queue out the door.
Without backpressure: everyone joins the queue. After an hour, people at the back give up and rejoin at the front (retry), making the queue longer. The clerks serve people who left twenty minutes ago.
With backpressure: a sign at the door says “queue full, back at 3pm.” The queue stays short, served people are happy, and turned-away people know to come back rather than waiting three hours for nothing.
4. Tiny example — the retry storm#
// retries.go — an overloaded server, with and without retries and cancellation.
package main
import (
"container/heap"
"fmt"
"math"
"math/rand"
)
type event struct {
t float64
attempt int
}
type events []event
func (h events) Len() int { return len(h) }
func (h events) Less(i, j int) bool { return h[i].t < h[j].t }
func (h events) Swap(i, j int) { h[i], h[j] = h[j], h[i] }
func (h *events) Push(x any) { *h = append(*h, x.(event)) }
func (h *events) Pop() any { old := *h; x := old[len(old)-1]; *h = old[:len(old)-1]; return x }
type floats []float64
func (h floats) Len() int { return len(h) }
func (h floats) Less(i, j int) bool { return h[i] < h[j] }
func (h floats) Swap(i, j int) { h[i], h[j] = h[j], h[i] }
func (h *floats) Push(x any) { *h = append(*h, x.(float64)) }
func (h *floats) Pop() any { old := *h; x := old[len(old)-1]; *h = old[:len(old)-1]; return x }
func simulate(arrivalRate float64, retries, cancellation bool, simTime float64) (completed int, wasted float64) {
rng := rand.New(rand.NewSource(1))
slots := make(floats, 32) // when each of 32 slots becomes free
var q events
for t := rng.ExpFloat64() / arrivalRate; t < simTime; t += rng.ExpFloat64() / arrivalRate {
heap.Push(&q, event{t, 0}) // the original requests
}
for q.Len() > 0 {
ev := heap.Pop(&q).(event)
if ev.t >= simTime {
break
}
deadline := ev.t + 30 // the client gives up after 30 s
start := math.Max(ev.t, slots[0])
service := math.Exp(1.0 + 0.8*rng.NormFloat64())
var busyUntil float64
switch {
case start+service <= deadline: // finished while the client was still listening
busyUntil = start + service
completed++
case cancellation: // server notices the client left and aborts early
busyUntil = math.Max(start, deadline)
wasted += busyUntil - start
default: // server runs to completion for nobody
busyUntil = start + service
wasted += service
}
slots[0] = busyUntil
heap.Fix(&slots, 0)
if start+service > deadline && retries && ev.attempt < 2 {
heap.Push(&q, event{deadline, ev.attempt + 1}) // the client tries again
}
}
return completed, wasted
}
func main() {
for _, c := range []struct {
label string
retries, cancel bool
}{{"no retries, no cancel", false, false}, {"retries, no cancel", true, false}, {"retries + cancel", true, true}} {
done, wasted := simulate(12, c.retries, c.cancel, 300)
fmt.Printf("%-25s completed=%5d wasted_gpu_sec=%8.0f\n", c.label, done, wasted)
}
}The output:
no retries, no cancel completed= 645 wasted_gpu_sec= 11599
retries, no cancel completed= 639 wasted_gpu_sec= 30218 ← WORSE
retries + cancel completed= 1906 wasted_gpu_sec= 6684 ← bestRetries without cancellation make things worse. Retries with cancellation help. The combination is what matters, and most systems implement only the first half.
5. Technical explanation#
Backpressure — the mechanisms#
1. REJECT (503 + Retry-After)
The primary mechanism. Fast, explicit, honest.
2. BOUNDED QUEUES at every layer
TCP backlog, API queue, engine waiting queue. Each bounded.
3. FLOW CONTROL on streams
If a client isn't reading, stop generating. TCP backpressure propagates
naturally IF you check it.
4. RATE LIMITING (token bucket, per tenant)
Prevents one client from consuming the fleet.
5. CONCURRENCY LIMITS per tenant
Simpler than rate limiting and directly bounds slot occupancy.Mechanism 3 is LLM-specific and often missed. If a client stops reading the stream, the
socket buffer fills, writes block, and — if you’re not checking — you either block the event
loop or buffer unboundedly in userspace. Check is_disconnected() and handle write backpressure.
Timeouts — you need several#
TIMEOUT TYPICAL VALUE PURPOSE
connect 1-5 s network reachability
TLS handshake 2-5 s
queue wait 5-30 s don't queue forever
time to first token 30-60 s catch prefill problems
inter-token stall 10-30 s catch engine hangs ← LLM-specific
total generation 5-15 min absolute bound
idle connection > max generation don't kill live streamsThe inter-token stall timeout is the important one. A total-request timeout of 10 minutes is useless for detecting a wedged engine; a “no token in 30 seconds” timeout catches it in 30 seconds.
var ErrEngineStalled = errors.New("engine stalled")
// streamWithStallTimeout forwards tokens, but fails if the engine goes quiet.
// The timeout is per TOKEN, not per request: a long answer is fine, a silent engine is not.
func streamWithStallTimeout(ctx context.Context, in <-chan string, out chan<- string, stall time.Duration) error {
timer := time.NewTimer(stall)
defer timer.Stop()
for {
select {
case tok, ok := <-in:
if !ok {
return nil // generation finished normally
}
out <- tok
timer.Reset(stall)
case <-timer.C:
return fmt.Errorf("%w: no token in %v", ErrEngineStalled, stall)
case <-ctx.Done():
return ctx.Err() // client went away
}
}
}Retries — the discipline#
RULE 1: Never retry a request that may still be running.
Cancel first, confirm cancellation, then retry.
RULE 2: Never retry after streaming has started.
The client has partial output. A retry duplicates it.
RULE 3: Retry only idempotent-ish failures:
connection refused, 503 with Retry-After, 429.
NOT: timeouts (the work may be in flight), 500 (may have partial effect).
RULE 4: Budget retries.
Cap retries at ~10% of total requests, globally. A retry budget
prevents a partial outage from becoming a total one.
RULE 5: Exponential backoff with jitter. Always jitter.
Without jitter, retries synchronize into waves.
RULE 6: Respect Retry-After.Retry budget implementation:
// RetryBudget allows retries only while they are a small fraction of total traffic.
type RetryBudget struct {
mu sync.Mutex
ratio float64 // e.g. 0.1: retries may add at most 10% load
window time.Duration // e.g. one minute
requests, retries []time.Time
}
func (b *RetryBudget) RecordRequest(now time.Time) {
b.mu.Lock()
defer b.mu.Unlock()
b.requests = append(b.requests, now)
}
func (b *RetryBudget) AllowRetry(now time.Time) bool {
b.mu.Lock()
defer b.mu.Unlock()
cutoff := now.Add(-b.window)
expire := func(ts []time.Time) []time.Time {
for len(ts) > 0 && ts[0].Before(cutoff) {
ts = ts[1:]
}
return ts
}
b.requests, b.retries = expire(b.requests), expire(b.retries)
if float64(len(b.retries)) >= b.ratio*float64(max(len(b.requests), 10)) {
return false // budget spent: fail now rather than add load
}
b.retries = append(b.retries, now)
return true
}This one class prevents most retry storms. It is 15 lines and almost nobody implements it.
Cancellation propagation#
Client disconnects
→ the ASGI server sets the disconnect flag
→ the streaming handler notices (checks between tokens)
→ calls engine.abort(request_id)
→ the scheduler removes the request from `running`
→ the KV manager frees its blocks
→ the slot is immediately availableEvery link in that chain must exist. If any is missing, capacity leaks. Test it explicitly: start a generation, kill the client, and verify the KV blocks are freed within one step.
6-9. Under the hood, performance, production, mistakes#
Under the hood — measuring abandonment:
Instrument:
requests_aborted_total{reason="client_disconnect"}
tokens_generated_after_disconnect_total ← pure waste; should be ~0
abort_latency_seconds ← time from disconnect to slot freedIn a chat product, client_disconnect is typically 10-30% of requests. If
tokens_generated_after_disconnect is large, cancellation isn’t working.
Performance: the capacity impact of cancellation:
20% of requests abandoned at 50% of their generation
→ 10% of all generated tokens are wasted
→ fixing it is a 10-11% capacity increase, for a day of workProduction:
- Implement cancellation first. It’s the highest return-on-effort item in this file.
- Set a stall timeout, not just a total timeout.
- Implement a retry budget.
- Return
Retry-Afteron 429 and 503, and honor it in your own clients. - Bound every queue.
- Test the overload path. Load-test past capacity and verify you reject cleanly rather than collapsing. Most teams never do this.
- Alert on rejection rate, not just on latency — rejections are the healthy failure mode.
Mistakes:
- Retries without cancellation. Actively harmful.
- Retrying on timeout. The original may still be running.
- Retrying after streaming started.
- No jitter. Synchronized retry waves.
- No retry budget. Partial outage → total outage.
- Idle timeout shorter than generation time. Streams die at 60 seconds.
- No stall timeout. A wedged engine holds slots forever.
- Never testing the overload path.
10. Hands-on exercise#
A. Reproduce the storm. Run the simulation in section 4. Vary the arrival rate to find where retries-without-cancellation cause collapse. Add cancellation and re-run.
B. Implement cancellation end to end. In a real server, wire up disconnect detection → abort → KV free. Verify by killing a client mid-generation and checking that the KV block count recovers within one step.
C. Stall timeout. Implement the stall-timeout wrapper. Simulate an engine hang and verify it fires in 30 seconds rather than at the total timeout.
D. Retry budget. Implement the RetryBudget class. Simulate a partial outage (30% of
requests failing) with and without the budget. Compare total load on the backend.
E. Overload test. Load-test your service at 2x capacity. Does it reject cleanly? What’s the p99 latency of served requests? Does it recover when load drops?
11. Interview questions#
- What is a metastable failure and how do retries cause one?
- Why are retries harmful without cancellation in LLM serving?
- What timeouts does an LLM service need? Why is a stall timeout necessary?
- What is a retry budget and why does it matter?
- Why should you never retry after streaming has started?
- How does cancellation propagate from a disconnected client to freed KV blocks?
- What is the healthy failure mode under overload, and how do you verify you have it?
12. Further reading#
- [FUNDAMENTAL] Google SRE Book, “Handling Overload,” “Addressing Cascading Failures”
- [FUNDAMENTAL] Bronson et al., “Metastable Failures in Distributed Systems” (HotOS 2021)
- [ESTABLISHED] AWS Builders’ Library: “Timeouts, retries, and backoff with jitter”
- Next: 06 — Load balancing and routing