The idea in one minute#
Latency is not one number; it is a distribution, and it is almost always lopsided: most requests are quick and a few are very slow. The average blends them into a figure that describes nobody. Percentiles describe the distribution directly: the p99 is the value that 99% of requests were faster than.
Two traps follow. You cannot average percentiles from different machines or time windows. And a user who makes many requests meets the slow tail far more often than “1%” suggests.
An analogy#
Nine people in a room earn a normal salary and one is a billionaire. The average income in the room is enormous and tells you nothing about any of the ten. The median tells you about the nine; the maximum tells you about the one. You need both.
A picture#
flowchart LR
subgraph DIST["1,000 request latencies, sorted"]
direction LR
A["fastest"] --> B["p50<br/>500th value"] --> C["p90<br/>900th"] --> D["p99<br/>990th"] --> E["max"]
end
AVG["average<br/>pulled up by the tail"] -.->|"sits somewhere between<br/>p50 and p99"| C
class A,B neutral
class C queue
class D,E warn
class AVG memoryHow it really works#
Definitions#
Sort N measurements. The p-th percentile is the value at position p/100 × N.
| Percentile | Meaning | Use |
|---|---|---|
| p50 (median) | Half the requests were faster | The typical experience |
| p90 / p95 | One in ten / twenty was slower | Early warning |
| p99 | One in a hundred was slower | The usual SLO target |
| p99.9 | One in a thousand | Large fan-out systems |
| max | The worst | Debugging, never alerting |
Why the tail matters more than its size suggests#
If a page makes 20 backend calls, the chance that at least one of them lands in the slowest 1%
is 1 − 0.99^20 ≈ 18%. With 100 calls it is 63%. In a system with fan-out — or an AI agent
that calls a model thirty times to complete one task — the tail of the component becomes the
median of the whole.
You cannot average percentiles#
Two servers each report p99 = 100 ms and p99 = 900 ms. The fleet p99 is not 500 ms. It depends on how many requests each server handled and on the shapes of both distributions; the two numbers alone are not enough to compute it.
The fix is to ship something that can be merged: a histogram. Each server counts how many requests fell into each latency bucket. Bucket counts add up across servers and across time, and the percentile is computed afterwards from the summed buckets. That is exactly how Prometheus works (II.04), and it is the single most important reason histograms exist.
Where tails come from#
| Cause | Why it is intermittent |
|---|---|
| Queueing | A burst arrives while the server is busy |
| Garbage collection, compaction | Periodic pauses |
| Cold caches, cold starts | First request after idle |
| Noisy neighbours | Shared CPU, disk or GPU |
| Retries and timeouts | A slow dependency is retried |
| Uneven work | In LLM serving, a 30,000-token prompt next to a 30-token one |
The last row is why LLM latency is normalized per token (TPOT) and measured per phase (V.03): otherwise the “tail” is just the long requests.
Coordinated omission#
A load generator that waits for each response before sending the next request stops measuring exactly when the system is slow. It reports a healthy p99 for a system that stalled for ten seconds. Use an open-loop generator, which sends on a schedule regardless of responses, and measure from the intended send time.
Code#
// tails.go — the average hides the tail, and percentiles cannot be averaged.
package main
import (
"fmt"
"math"
"math/rand"
"sort"
)
func pct(sorted []float64, p float64) float64 {
if len(sorted) == 0 {
return math.NaN()
}
i := int(math.Ceil(p/100*float64(len(sorted)))) - 1
if i < 0 {
i = 0
}
return sorted[i]
}
func sample(rng *rand.Rand, n int, slowFraction float64) []float64 {
out := make([]float64, n)
for i := range out {
out[i] = 40 + rng.ExpFloat64()*15 // typical: ~55 ms
if rng.Float64() < slowFraction {
out[i] += 800 + rng.Float64()*1200 // a slow path: queueing, GC, cold cache
}
}
sort.Float64s(out)
return out
}
func mean(xs []float64) float64 {
s := 0.0
for _, x := range xs {
s += x
}
return s / float64(len(xs))
}
func main() {
rng := rand.New(rand.NewSource(7))
a := sample(rng, 90000, 0.002) // healthy server, most of the traffic
b := sample(rng, 10000, 0.08) // unhealthy server, little traffic
fmt.Println("server n mean p50 p90 p99")
for _, s := range []struct {
name string
xs []float64
}{{"A", a}, {"B", b}} {
fmt.Printf("%-6s %6d %6.0f %5.0f %5.0f %6.0f ms\n",
s.name, len(s.xs), mean(s.xs), pct(s.xs, 50), pct(s.xs, 90), pct(s.xs, 99))
}
all := append(append([]float64{}, a...), b...)
sort.Float64s(all)
fmt.Printf("\nfleet p99, computed correctly from all requests: %6.0f ms\n", pct(all, 99))
fmt.Printf("average of the two servers' p99s (WRONG): %6.0f ms\n", (pct(a, 99)+pct(b, 99))/2)
fmt.Println("\nChance a user meets the slowest 1% at least once:")
for _, calls := range []int{1, 5, 20, 100} {
fmt.Printf(" %3d calls per task → %4.1f%%\n", calls, 100*(1-math.Pow(0.99, float64(calls))))
}
}Remember this#
- Latency is a distribution. Report percentiles, never only the average.
- Percentiles cannot be averaged or summed. Histograms can be merged; compute the percentile after merging.
- Fan-out turns a component’s tail into the whole system’s typical case.
- Measure with an open-loop load generator, or you will not see your own stalls.
Try it#
- Run
tails.go. Change server B’s share of traffic from 10% to 50%. How does the correct fleet p99 move compared with the wrong one? - Add a p99.9 column. How many samples do you need before p99.9 is stable from run to run?
- An agent makes 30 model calls per task. What per-call percentile must meet your latency target for 95% of tasks to be fast throughout?
Check yourself#
- Why does the average mislead for latency?
- Why can histograms be combined across servers when percentiles cannot?
- What is coordinated omission?