PidokuInfra

Inference Engine Metrics

Advanced 1h 10m Difficulty 4/5 Topic 03 of 07

Prerequisites 01, II.04; helpful: Inference Engineering V.03, V.05, V.09

The idea in one minute#

An LLM server is a queue in front of a batch loop. A request waits, is prefilled (the prompt is read, producing the first token), then decodes one token per step alongside everyone else in the batch. Each phase has its own latency metric and its own cause of slowness, so the engine must report them separately.

Five families of metric tell you nearly everything: latency by phase (queue, TTFT, inter-token), throughput in tokens, saturation (waiting requests, KV-cache usage, preemptions), cache effectiveness (prefix cache hits), and errors and finish reasons.

An analogy#

A lift in a tall building. You wait for it (queue). It takes a moment to load everyone (prefill). Then it stops at every floor for every passenger (decode steps shared by the batch). “How long did my trip take” is the sum, but to fix a slow building you need to know which of the three is the problem — and how full the lift was.

A picture#

flowchart TB
  ARR["Request arrives"] --> W["WAITING<br/>queue time"]
  W --> PF["PREFILL<br/>read the prompt, cached prefix skipped"]
  PF --> FT["First token sent<br/>TTFT = queue + prefill"]
  FT --> DEC["DECODE<br/>one token per step, shared with the batch<br/>inter-token latency"]
  DEC --> END["Finished<br/>stop, length, error"]
  DEC -->|"KV cache full"| PRE["Preempted<br/>back to waiting"]
  PRE --> W
  KV[("KV cache blocks")] --- DEC
  class ARR,END,FT neutral
  class W queue
  class PF,DEC compute
  class KV memory
  class PRE warn

How it really works#

The latency metrics#

MetricDefinitionDominated byUsers feel it as
Queue timeArrival → scheduledLoad vs capacityPart of the initial wait
Prefill timeScheduled → first token computedUncached prompt length; compute-boundPart of the initial wait
TTFT (time to first token)Arrival → first token sentQueue + prefill (+ network)“Is it responding?”
ITL (inter-token latency)Gap between consecutive tokensBatch size, model size, memory bandwidthReading speed, stutter
TPOT (time per output token)(end − first token) ÷ (output tokens − 1)Same as ITL, averaged per requestOverall streaming speed
End-to-endArrival → last tokenTTFT + TPOT × output tokensTotal wait for non-streaming use

end-to-end ≈ TTFT + TPOT × (output_tokens − 1). Because output length varies so much, end-to-end latency is mostly a measure of how much was asked for. Set SLOs on TTFT and TPOT, and track end-to-end only per workload class.

ITL and TPOT are not the same thing: TPOT averages a request’s gaps, hiding a stall; ITL’s distribution shows it. A stream that pauses for two seconds mid-answer has a fine TPOT and a terrible ITL p99.

Throughput#

  • Output tokens/s and prompt tokens/s, per replica and per model — the load figures.
  • Tokens per engine step — how full the batch actually is.
  • Requests running — the current batch size.

Saturation#

SignalMeaningWhy it matters
Requests waitingQueue depthThe earliest, clearest overload signal; the best autoscaling input
KV-cache usageFraction of KV blocks allocatedThe true memory pressure (GPU “memory used” is not)
PreemptionsA running request evicted to free KV blocksThe engine is thrashing: lots of wasted work, sharp latency tail
Waiting by reasonWhy requests cannot start (capacity, KV space, …)Tells you whether to add replicas or change limits

Cache effectiveness#

Prefix cache hit rate = cached prompt tokens ÷ prompt tokens. For chat and agents with long repeated prefixes this is the single largest cost lever: a hit turns thousands of prefill tokens into a lookup. A drop in hit rate after a deploy (someone put a timestamp at the top of the system prompt) can double GPU cost with no other symptom than higher TTFT.

vLLM’s metrics, by name#

vLLM serves Prometheus metrics at /metrics, prefixed vllm:. Names below are from the metrics reference for v0.30 (September 2026). Names have changed between releases — for example KV-cache usage was once exported as vllm:gpu_cache_usage_perc — so check your version’s page.

FamilyMetricType
Latencyvllm:time_to_first_token_secondsHistogram
vllm:inter_token_latency_secondsHistogram
vllm:request_time_per_output_token_secondsHistogram
vllm:e2e_request_latency_secondsHistogram
vllm:request_queue_time_seconds, vllm:request_prefill_time_seconds, vllm:request_decode_time_seconds, vllm:request_inference_time_secondsHistograms
Throughputvllm:prompt_tokens, vllm:generation_tokensCounters
vllm:iteration_tokens_totalHistogram (tokens per engine step)
vllm:request_prompt_tokens, vllm:request_generation_tokensHistograms (size per request)
Saturationvllm:num_requests_running, vllm:num_requests_waitingGauges
vllm:num_requests_waiting_by_reasonGauge
vllm:kv_cache_usage_percGauge, 0–1
vllm:num_preemptions, vllm:request_num_preemptionsCounter, histogram
Cachevllm:prefix_cache_queries, vllm:prefix_cache_hitsCounters (tokens)
vllm:prompt_tokens_cached, vllm:prompt_tokens_by_sourceCounters
vllm:kv_block_lifetime_seconds, vllm:kv_block_idle_before_evict_secondsHistograms
Outcomevllm:request_success (labelled by finish reason)Counter
vllm:corrupted_requestsCounter (NaNs in the logits)
Efficiencyvllm:estimated_flops_per_gpu_total, vllm:estimated_read_bytes_per_gpu_totalCounters
Speculative decodingvllm:spec_decode_num_accepted_tokens_per_posCounter

Other engines expose the same ideas under other names: SGLang (sglang: prefix), NVIDIA’s Dynamo and Triton, TensorRT-LLM, llama.cpp’s server, and hosted APIs through their usage fields. Hugging Face’s TGI — archived in March 2026 — used tgi_ names you will still meet in older dashboards. Learn the families, then map the names.

The queries to keep#

PromQL
# TTFT p95 per model
histogram_quantile(0.95, sum by (le, model_name) (rate(vllm:time_to_first_token_seconds_bucket[5m])))

# Output tokens per second, per replica
sum by (pod) (rate(vllm:generation_tokens_total[1m]))

# Prefix cache hit rate
sum(rate(vllm:prefix_cache_hits_total[5m])) / sum(rate(vllm:prefix_cache_queries_total[5m]))

# Queue pressure per replica, and preemptions per minute
max by (pod) (vllm:num_requests_waiting)
sum by (pod) (rate(vllm:num_preemptions_total[5m])) * 60

# Where does request time go? Mean share spent queued
sum(rate(vllm:request_queue_time_seconds_sum[5m])) / sum(rate(vllm:e2e_request_latency_seconds_sum[5m]))

Prometheus appends _total to counters and _bucket / _sum / _count to histograms on exposition; confirm the exact series names by reading your own /metrics.

Reading the combinations#

SymptomWithLikely causeAction
TTFT upWaiting up, KV usage moderateNot enough compute for arrivalsScale out
TTFT upWaiting low, prefill time upLonger or less-cached promptsCheck prefix hit rate, prompt changes
ITL upRunning count upBigger batch: the designed trade-offLower the concurrency limit, or scale
ITL p99 spikesA long-prompt request arrivedPrefill is stalling decode stepsChunked prefill; separate prefill and decode
Preemptions above zeroKV usage near 1.0KV cache exhaustedFewer concurrent sequences, shorter contexts, more memory
Throughput downClock or power down on the GPUThermal or power throttlingLesson 02
Tokens/s fine, users unhappyFinish reason length risingTruncated outputsA quality problem (lesson 07)

Goodput#

Throughput counts every token. Goodput counts only requests that met the SLO: TTFT under its threshold and TPOT under its threshold and completed. As load rises, throughput keeps climbing while goodput peaks and collapses — the peak is your real capacity. Compute it per request (a gateway or client-side event is the natural place) and make it the number you size, autoscale and benchmark against.

Measure from outside as well#

Server-side TTFT stops at the engine’s socket. Users also pay for the gateway, the network and client buffering. Record TTFT and ITL in the client or with a synthetic prober, and compare: the difference is everything in front of the engine.

Code#

Compute the streaming metrics and goodput from token timestamps, as a load tester or gateway would.

Go
// streammetrics.go — TTFT, TPOT, ITL p99 and goodput from per-token timestamps.
package main

import (
	"fmt"
	"math/rand"
	"sort"
)

type Request struct {
	Arrival float64   // seconds
	Tokens  []float64 // time each output token was received
}

func simulate(rng *rand.Rand, n int, load float64) []Request {
	reqs := make([]Request, n)
	for i := range reqs {
		queue := rng.ExpFloat64() * 0.05 * load * load // queueing grows sharply with load
		prefill := 0.08 + rng.Float64()*0.25
		itl := 0.018 * (1 + 0.6*load) // bigger batches → slower steps
		t := queue + prefill
		toks := make([]float64, 40+rng.Intn(400))
		for j := range toks {
			toks[j] = t
			gap := itl * (0.9 + 0.2*rng.Float64())
			if rng.Float64() < 0.002*load { // another request's long prefill stalls the batch
				gap += 0.8
			}
			t += gap
		}
		reqs[i] = Request{0, toks}
	}
	return reqs
}

func p(xs []float64, q float64) float64 {
	s := append([]float64{}, xs...)
	sort.Float64s(s)
	return s[int(q*float64(len(s)-1))]
}

func main() {
	const ttftSLO, tpotSLO = 0.5, 0.040
	rng := rand.New(rand.NewSource(4))
	fmt.Println("load  TTFT p95  TPOT p95  ITL p99   tokens/s/req  good requests")
	for _, load := range []float64{0.5, 1, 2, 3, 4} {
		reqs := simulate(rng, 2000, load)
		var ttft, tpot, itl []float64
		good := 0
		for _, r := range reqs {
			first, last := r.Tokens[0], r.Tokens[len(r.Tokens)-1]
			tt := first - r.Arrival
			tp := (last - first) / float64(len(r.Tokens)-1)
			ttft, tpot = append(ttft, tt), append(tpot, tp)
			for i := 1; i < len(r.Tokens); i++ {
				itl = append(itl, r.Tokens[i]-r.Tokens[i-1])
			}
			if tt <= ttftSLO && tp <= tpotSLO {
				good++
			}
		}
		fmt.Printf("%4.1f  %6.0f ms  %5.1f ms  %6.0f ms  %10.1f  %10.1f%%\n", load,
			p(ttft, .95)*1000, p(tpot, .95)*1000, p(itl, .99)*1000, 1/p(tpot, .5), 100*float64(good)/float64(len(reqs)))
	}
	fmt.Printf("\nSLO: TTFT <= %.0f ms and TPOT <= %.0f ms. Capacity is the load where 'good' is still ~99%%.\n",
		ttftSLO*1000, tpotSLO*1000)
}

Remember this#

  • Three phases — wait, prefill, decode — each with its own metric and its own cause.
  • SLOs on TTFT and TPOT; watch the ITL distribution for stalls.
  • Saturation is requests waiting, KV-cache usage and preemptions — not GPU memory or utilization.
  • Prefix cache hit rate is a cost metric. Goodput, not throughput, is capacity.
  • Metric names differ by engine and version; the families do not.

Try it#

  1. Run streammetrics.go. At which load does goodput fall below 99%? Which SLO breaks first?
  2. Start any engine with a small model and read /metrics. Map every family in this lesson to a series name. No GPU: do the same against the vLLM metrics reference page.
  3. Build a six-panel dashboard from the queries above, in the order: traffic, errors, TTFT, ITL, saturation, cache.

Check yourself#

  1. Which two phases make up TTFT, and what drives each?
  2. How can TPOT look healthy while users see the stream freeze?
  3. What do preemptions tell you, and what is the usual fix?

↑↓ navigate↵ openesc close