1. What is it?#
Every instinct a good backend engineer has — stateless services, horizontal scaling, uniform request cost, retries are cheap, autoscale on CPU — is either wrong or dangerously incomplete for LLM inference.
This file is a systematic list of the differences, so that when your intuition fires, you know whether to trust it.
2. Why does it exist?#
Because the failure mode is subtle. A team builds an LLM service the way they build every other service. It works in staging. It works at launch. Then traffic grows 3x and it collapses in a way no runbook covers: latency climbs to 30 seconds, autoscaling adds replicas that take 8 minutes to become useful, retries multiply the load, and the KV cache thrashes.
Every one of those symptoms traces back to an assumption that holds for stateless CRUD services and does not hold here.
3. Simple analogy#
A normal web service is a supermarket checkout. LLM inference is an operating theatre.
At a checkout: customers take roughly similar time, you open another lane when the queue grows, a lane can be opened in seconds, and if a transaction fails you just do it again.
In an operating theatre: procedures take between 20 minutes and 9 hours with no way to know in advance, you cannot “open another theatre” in under a day, the theatre is occupied for the whole procedure, resources (blood, ICU beds) are consumed progressively and can run out mid-procedure, and you absolutely cannot “just retry.”
Every operational difference below is a version of that.
4. Tiny example#
The retry storm — the shortest illustration of why intuition fails.
// Perfectly reasonable backend code
func callModel(ctx context.Context, prompt string) (resp *http.Response, err error) {
for attempt := 0; attempt < 3; attempt++ {
ctx, cancel := context.WithTimeout(ctx, 30*time.Second)
defer cancel()
req, _ := http.NewRequestWithContext(ctx, "POST", modelURL, strings.NewReader(prompt))
if resp, err = http.DefaultClient.Do(req); err == nil {
return resp, nil
}
time.Sleep(time.Second << attempt) // exponential backoff
}
return nil, err
}On a normal service under load: a request times out, you retry, it succeeds, everyone is happy. Cost of a retry ≈ cost of the original request ≈ 10 ms of CPU.
On an LLM service under load:
t=0 request arrives, gets a KV slot, starts generating
t=30s client times out, retries
...but the original request is STILL RUNNING, still holding its KV slot,
still consuming GPU, still generating tokens nobody will read
t=30s retry #1 enters the queue behind everything else
t=60s retry #1 times out; original may still be running; retry #2 entersYou have tripled the load on a system that was already overloaded, and none of the work is discardable because the engine does not know the client left. This is a metastable failure: the system cannot recover even after arrival rate returns to normal, because it is now busy serving abandoned work.
The fixes are inference-specific: cancellation propagation (abort the sequence when the client disconnects), retry budgets, admission control that rejects rather than queues, and never retrying a request that may still be running.
5. Technical explanation#
The eleven differences#
1. Requests are not fungible — cost varies by 1000x.
Request A: 20-token prompt, 10-token answer → ~30 token-units
Request B: 100k-token prompt, 4000-token answer → ~104,000 token-unitsConsequences: round-robin load balancing is wrong; per-request rate limits are wrong (use token buckets); “requests in flight” is a poor load signal (use token throughput or KV occupancy); p99 latency without normalizing for length is meaningless.
2. Requests are long-lived and stateful.
A single request occupies a KV cache slot for its entire generation — seconds to minutes. The server holds per-request state that cannot be moved to another replica. This makes inference replicas stateful in practice, which breaks graceful-drain assumptions, blue/green deployments, and any pattern that assumes a request can be re-routed mid-flight.
3. Capacity is memory-shaped and cliff-like.
Normal services degrade gracefully: more load → more latency. LLM servers degrade gracefully until KV memory runs out, then either reject, preempt/swap sequences (adding huge latency spikes), or OOM-crash. The transition is abrupt.
latency
^ ┃ ← KV exhaustion
| ┃
| ______/┃
| _________________/ (preemption / rejection / crash)
|_________/
+--------------------------------------> concurrent sequences4. Scaling out takes minutes, not seconds.
Container pull (multi-GB image) 30-120 s
Model weights load (140 GB) 60-600 s ← dominated by storage bandwidth
CUDA init, graph capture, warmup 20-60 s
──────────────────────────────────────────────
Total cold start 2-12 minutesYour traffic spike is over before the new replica is ready. This forces pre-provisioned headroom rather than reactive autoscaling, which changes the cost model fundamentally. (Sections VIII.07 and XI.04 cover the mitigations: warm pools, fast weight loading, snapshotting, model streaming.)
5. The unit of work is a token, not a request.
Metrics, rate limits, quotas, billing, and load balancing should all be token-denominated. Teams that build request-denominated systems rebuild them within a year.
6. Cost per unit of work is 100-1000x higher.
A typical web request costs microdollars. An LLM request can cost cents. This changes what is worth engineering: caching a response that costs $0.00001 is not worth the complexity; caching a prefix that costs $0.02 absolutely is. Efficiency work has direct, visible P&L impact — which is why inference engineers are paid what they are paid.
7. Output is nondeterministic by default.
Same input, different output — from sampling, from batch-size-dependent kernel selection, from floating-point non-associativity in reduction order. Consequences: cache keys must include sampling parameters; “reproduce the bug” is hard; A/B tests need larger samples; contract tests on exact output strings will flake.
8. Correctness is a distribution, not a boolean.
There is no 200-vs-500 for “the answer is wrong.” Quality regressions from a quantization change or a kernel swap are invisible to conventional monitoring. You need evaluation suites in CI and in production (Section XI.08).
9. Streaming changes the whole request lifecycle.
Once you have emitted the first token you cannot return an error status; you can only stop mid- stream. Load balancers must not buffer. Timeouts must be per-token, not per-request. Connections are long-lived, which stresses connection-limited proxies.
10. Throughput and latency are controlled by the same knob.
In a normal service, adding capacity improves both. Here, batch size improves throughput and worsens latency simultaneously, and it is the scheduler — not a config file — that decides.
11. The hardware is scarce and expensive.
You cannot simply add 200 more GPUs next Tuesday. Capacity planning has procurement lead times measured in quarters. This makes efficiency work strategically important, not merely nice.
The translation table#
| Backend instinct | LLM reality |
|---|---|
| Stateless services | Stateful for the request’s lifetime (KV cache) |
| Autoscale on CPU% | Autoscale on KV occupancy / queue depth / token rate |
| Scale out in seconds | Minutes; pre-provision instead |
| Retry on timeout | Cancel first; retries can cause metastable collapse |
| Round-robin LB | Load-aware and prefix-aware routing |
| Requests/sec capacity | Tokens/sec, plus a concurrency limit from memory |
| p99 latency SLO | Separate TTFT and ITL SLOs, bucketed by length |
| Rate limit by RPS | Rate limit by tokens (input and output separately) |
| Graceful drain | Must wait for in-flight generations or cancel them |
| Cache the response | Cache the prefix’s KV, which is far more valuable |
| Errors are 4xx/5xx | Errors include “the answer is subtly worse” |
6. Under the hood#
Why is a replica effectively stateful? Because the KV cache for sequence s lives in that GPU’s HBM, in blocks referenced by a block table. Moving the request elsewhere means transferring hundreds of MB to GB over the network, or recomputing the prefill. Both are expensive enough that no production system does it casually — though KV transfer is exactly what disaggregated serving does deliberately, with fast interconnects (Section XIII.06).
Why is cold start so slow? Reading 140 GB from object storage at 1 GB/s takes 140 seconds; from
a local NVMe at 5 GB/s, 28 seconds; and then it must be copied to HBM. Mitigations — local
caching, parallel shard loading, safetensors mmap + direct-to-GPU, pre-baked images, or
keeping warm standby replicas — all trade money for time.
7. Performance implications#
- Queueing dominates TTFT under load, so admission control has more leverage than kernel tuning for tail latency.
- Cancellation is a performance feature. In a chat product, 10-30% of generations are abandoned. Freeing those slots immediately is equivalent to a 10-30% capacity increase — one of the highest-return changes available, and frequently missing.
- Head-of-line blocking is real. One 100k-token prefill can stall every decode in flight for a second or more. Chunked prefill (Section XIII.05) is the fix.
8. Production implications#
A checklist you can apply to any LLM service:
[ ] Client disconnect propagates to the engine and aborts the sequence
[ ] Timeouts are per-token (stall detection), not just per-request
[ ] Retries have a budget and never duplicate in-flight generations
[ ] Rate limits are token-based, separately for input and output
[ ] Max input length and max output length are enforced at the gateway
[ ] Autoscaling signal is KV occupancy or queue wait, not CPU
[ ] Warm replicas / pre-provisioned headroom sized for spike + cold-start time
[ ] Load balancer does not buffer streamed responses
[ ] Separate TTFT and ITL SLOs, reported bucketed by prompt length
[ ] Admission control sheds load explicitly instead of queueing unboundedly
[ ] Quality evaluation runs in CI on the exact serving configuration
[ ] Graceful shutdown drains in-flight generations with a deadlineMost incidents in LLM serving trace to a missing item on that list rather than to a bad kernel.
9. Common mistakes#
Autoscaling on GPU utilization. It is ~100% whenever anything is running, so it is a useless signal. Use KV cache occupancy, queue depth, or token throughput.
Treating the engine as a black box behind a normal LB. Without load-awareness, one replica gets three 50k-token prefills and dies while others idle.
Unbounded queues. “We’ll just queue and drain later” produces requests that complete after the client gave up — pure waste. Bound queues and reject early.
Copying a microservices deployment strategy. Rolling updates that assume 10-second drains will kill in-flight generations.
Ignoring quality in the deployment pipeline. A quantization change that improves latency 2x and degrades reasoning quality by 5% will pass every conventional test you have.
Assuming idempotency. LLM calls are not idempotent (sampling) and are not free to repeat.
10. Hands-on exercise#
A. Simulate a retry storm. Write a simulator: Poisson arrivals, service times drawn from a lognormal distribution, a fixed concurrency limit, and clients that retry after a 30 s timeout without cancellation. Sweep arrival rate. Find the rate at which the system becomes metastable (does not recover after the spike ends). Then add cancellation and re-run. Quantify the difference.
B. Measure cold start. Time each phase of starting a real model server: image pull, weight load, CUDA init, first request, steady state. Where does the time actually go? What would you optimize first?
C. Cost of a bad load balancer. Simulate 4 replicas, round-robin vs least-outstanding-tokens routing, with a realistic (heavy-tailed) length distribution. Compare p99 TTFT.
D. Audit a real service. Take any LLM service you have access to (or vLLM locally) and go through the checklist in section 8. How many items pass?
11. Interview questions#
- Give five ways LLM serving violates standard backend assumptions.
- Why is autoscaling on GPU utilization wrong? What would you use instead?
- Explain a metastable failure in an LLM service and how to prevent it.
- Why are LLM inference replicas effectively stateful? What breaks as a result?
- A client disconnects mid-generation. What should happen, and what usually does?
- How would you rate-limit an LLM API fairly across tenants?
- Your deployment pipeline is green but users say answers got worse after a release. What was missing from the pipeline?
12. Further reading#
- [FUNDAMENTAL] Google SRE Book, “Handling Overload” and “Addressing Cascading Failures”
- [FUNDAMENTAL] Bronson et al., “Metastable Failures in Distributed Systems” (HotOS 2021)
- [ESTABLISHED] vLLM’s request abort / preemption implementation
- Next: 11 — Why inference gets expensive at scale