PidokuInfra

Anatomy of an Inference Server

Intermediate 1h 15m Difficulty 3/5 Topic 01 of 14

Prerequisites V.09, V.10


1. What is it?#

The components every LLM inference server has, and how they fit together.

                    ┌──────────────────────────────────────┐
   HTTP/gRPC  ────► │  API LAYER                           │
                    │  parse, validate, auth, rate limit   │
                    │  chat template, tokenize             │
                    └──────────────┬───────────────────────┘
                                   │ Request objects
                    ┌──────────────▼───────────────────────┐
                    │  SCHEDULER                           │
                    │  waiting queue / running set         │
                    │  admission, preemption, priorities   │
                    │  prefix cache lookup                 │
                    └──────────────┬───────────────────────┘
                                   │ batch metadata
                    ┌──────────────▼───────────────────────┐
                    │  KV CACHE MANAGER                    │
                    │  block pool, block tables, refcounts │
                    │  prefix hash table, eviction         │
                    └──────────────┬───────────────────────┘
                                   │
                    ┌──────────────▼───────────────────────┐
                    │  MODEL EXECUTOR                      │
                    │  forward pass, CUDA graphs           │
                    │  distributed communication (TP/PP)   │
                    └──────────────┬───────────────────────┘
                                   │ logits
                    ┌──────────────▼───────────────────────┐
                    │  SAMPLER                             │
                    │  penalties, temperature, top-k/p     │
                    │  stop conditions, constrained decode │
                    └──────────────┬───────────────────────┘
                                   │ token ids
                    ┌──────────────▼───────────────────────┐
                    │  DETOKENIZER + STREAM                │
                    │  incremental decode, SSE framing     │
                    └──────────────────────────────────────┘

Every production engine has exactly these six components. The differences are in the policies inside each.

Diagram — The request path through a serving system#

flowchart LR
  C["Client"] --> GW["Gateway<br/>auth, limits, routing"]
  GW --> API["API server<br/>validate, tokenize, chat template"]
  API --> Q["Queue<br/>admission control"]
  Q --> SCH["Scheduler<br/>continuous batching"]
  SCH --> EX["Model executor<br/>forward pass on GPU"]
  EX <--> KV[("KV cache manager")]
  EX --> SMP["Sampler"]
  SMP --> DET["Detokenize + stream"]
  DET -->|"SSE"| C
  SMP -.->|"next step"| SCH

  class GW,API,DET io
  class Q,SCH queue
  class EX,SMP compute
  class KV memory
  class C neutral

2. Why it’s structured this way#

Because of the constraints established in Section V:

CONSTRAINT                              → COMPONENT
Requests have unknown, variable cost    → scheduler with per-step decisions
Memory is the binding capacity limit    → KV cache manager owns admission
Decode is memory-bound                  → batching is the executor's job
Tokens must be streamed                 → detokenizer is incremental and stateful
CPU work competes with GPU driving      → API layer separated from the engine loop
Requests are long-lived and stateful    → per-request state lives with the scheduler

The architecture is a direct consequence of the physics. That’s why every engine converged on roughly the same shape.


3. Simple analogy#

A hospital emergency department.

  • API layer — reception: check in, verify insurance, take vitals.
  • Scheduler — triage: who goes next, who waits, who is turned away.
  • KV cache manager — bed management: how many beds, who occupies which, when they free.
  • Executor — the treatment rooms: where work actually happens.
  • Sampler — the decision at each step: what to do next for this patient.
  • Detokenizer/stream — communicating with the patient and family as things progress.

The bottleneck is beds, not doctors. Triage exists because beds are scarce. That’s exactly the KV-cache-bound situation of an LLM server.


4. Tiny example — the main loop#

Every engine’s heart:

Go
type Engine struct {
	model            Model
	kv               *BlockManager
	waiting, running []*Request
	maxBatch         int
	maxBatchedTokens int
}

func (e *Engine) Add(r *Request) {
	r.Arrival = time.Now()
	e.waiting = append(e.waiting, r)
}

func (e *Engine) Step() (finished []*Request) {
	// --- 1. SCHEDULE: decide what runs this iteration ---
	prefill, decode := e.schedule()
	if len(prefill)+len(decode) == 0 {
		return nil
	}

	// --- 2. PREPARE: build the flat tensors and metadata ---
	batch := e.prepareInputs(prefill, decode)

	// --- 3. EXECUTE: one forward pass ---
	logits := e.model.Forward(batch) // a CUDA graph replay if decode-only

	// --- 4. SAMPLE ---
	tokens := e.sample(logits, batch.Sampling)

	// --- 5. UPDATE: append tokens, grow KV, detect completion ---
	finished = e.update(append(prefill, decode...), tokens)

	// --- 6. FREE ---
	for _, r := range finished {
		e.kv.Free(r.BlockTable)
		e.running = slices.DeleteFunc(e.running, func(x *Request) bool { return x == r })
	}
	return finished
}

func (e *Engine) schedule() (prefill, decode []*Request) {
	// admit new requests while memory and the token budget allow
	decode = slices.Clone(e.running)
	budget := e.maxBatchedTokens - len(decode) // decode uses 1 token each
	for len(e.waiting) > 0 && len(e.running) < e.maxBatch {
		r := e.waiting[0]
		need := min(r.RemainingPromptTokens(), budget)
		if need <= 0 || !e.kv.CanAllocate(e.kv.BlocksNeeded(need)) {
			break
		}
		e.waiting = e.waiting[1:]
		r.BlockTable = e.kv.AllocateWithPrefixCache(r)
		prefill, e.running = append(prefill, r), append(e.running, r)
		budget -= need
	}
	return prefill, decode
}

That’s a working continuous-batching engine. Real ones add: chunked prefill (splitting a prefill across steps), preemption, priorities, multi-LoRA, speculative decoding, and careful optimization of prepare_inputs — but the loop is this.


5. Technical explanation of each component#

API layer#

Responsibilities:
  - Protocol handling (HTTP/1.1 + SSE, HTTP/2, gRPC)
  - Auth, rate limiting, quota
  - Request validation (max lengths, parameter ranges)
  - Chat template application
  - Tokenization (thread pool, releases the GIL)
  - Response streaming and detokenization
  - Metrics emission

Design notes:
  - Runs async (asyncio/tokio) — thousands of long-lived connections
  - SEPARATE PROCESS from the engine, usually. Why? The GIL: HTTP parsing
    and tokenization must not block the engine's kernel-launch loop.
  - Communicates with the engine over ZeroMQ, shared memory, or an async queue

Scheduler#

State:
  waiting:  deque of admitted-but-not-started requests
  running:  list of requests with allocated KV
  swapped:  preempted requests (if using swap-based preemption)

Per-step decisions:
  1. Which waiting requests to admit
  2. How many prefill tokens to process (chunked prefill budget)
  3. Whether to preempt anyone
  4. Batch composition

Policies (this is where engines differ):
  FCFS, priority, fair-share, shortest-first
  prefill-priority vs decode-priority
  chunked vs whole prefill

KV cache manager#

  block pool:        preallocated GPU memory, divided into fixed blocks
  free list:         available block ids
  block tables:      per-sequence logical→physical mapping
  refcounts:         for copy-on-write sharing
  prefix hash table: content hash → block id, for prefix caching
  eviction:          LRU over refcount-0 cached blocks

This component owns capacity. Admission is really “does the KV manager have room?”

Model executor#

  - Owns the weights and the KV cache tensors
  - Builds input tensors from the scheduler's metadata
  - Runs the forward pass (eager or CUDA graph replay)
  - Handles distributed execution (TP ranks, PP stages)
  - Returns logits

In a TP setup: one executor process per rank, all running in lockstep,
communicating via NCCL. The scheduler runs on rank 0 and broadcasts decisions.

Sampler#

  Fully on GPU (Section V.13). Handles per-request parameters via
  vectorized per-row operations.
  Also: stop-string detection, constrained decoding masks, logprob extraction.

Detokenizer#

  Stateful per request (Section V.02). Must handle multi-byte UTF-8
  spanning tokens, and hold back partial stop sequences.
  Usually runs in the API process, not the engine process.

6. Under the hood — the process topology#

A typical vLLM deployment with TP=4:

Process: API server (Python, asyncio)
   ├─ uvicorn worker(s)
   ├─ tokenizer thread pool
   └─ ZeroMQ ──┐
               │
Process: Engine core (rank 0)          ← scheduler lives here
   ├─ scheduler
   ├─ KV manager
   ├─ model executor (rank 0)
   └─ NCCL ────┬──────┬──────┐
               │      │      │
Process: rank 1   rank 2   rank 3     ← executors only, driven by rank 0

Why separate processes rather than threads?

  1. The GIL. The engine loop must not be blocked by HTTP work.
  2. TP ranks need separate CUDA contexts and separate NCCL ranks.
  3. Fault isolation — an API-layer crash shouldn’t take down the model.

Cost: IPC serialization on every request and response. Engines optimize this heavily (shared memory for large payloads, msgpack rather than pickle).


7-9. Performance, production, mistakes#

Performance — where CPU time goes in the engine process, at batch 128:

prepare_inputs (building tensors, block tables)   30-45%
scheduler bookkeeping                              15-25%
sampling metadata construction                     10-15%
kernel launches (with CUDA graphs)                  5-10%
output processing                                  10-15%

Note that almost none of it is kernel launching once graphs are on. The bottleneck shifts to Python-side data structure manipulation, which is why engines have moved scheduler internals to C++/Rust or heavily optimized the Python (flat arrays, cached tensors, incremental updates).

Production:

  • Monitor each component separately. Queue depth (scheduler), KV utilization (cache manager), step time (executor), CPU per process.
  • The API process and engine process have different scaling characteristics. You may need several API workers per engine.
  • prepare_inputs cost scales with batch size — at very large batch it can dominate. Profile it.
  • Fault isolation matters: an engine crash should be detected and the process restarted, with in-flight requests failed cleanly rather than hanging.

Mistakes:

  • Running the API and engine in one Python process. GIL contention destroys ITL.
  • Not bounding the waiting queue. Unbounded memory growth and requests that complete after the client left.
  • Blocking the async event loop in the API layer.
  • Treating the engine as a black box. When it’s slow, you need to know which component.

10. Hands-on exercise#

A. Build it. Implement the six components as separate classes, with the main loop from section 4. Use a small model. Support: streaming, per-request sampling params, and stop conditions. This is Project 06 and the foundation for 07-09.

B. Map a real engine. Read vLLM’s source and locate each of the six components. Draw the call graph for one request from HTTP arrival to first token.

C. Measure the components. Instrument a real engine to report per-component time per step. Where does CPU time actually go at batch 8 vs batch 256?

D. Process topology. For a running vLLM server with TP>1, list the processes and threads (ps -eLf, py-spy dump on each). Draw the topology. Which process would you profile for an ITL problem? For a TTFT problem?


11. Interview questions#

  1. Name the six components of an inference server and what each owns.
  2. Why are the API layer and the engine separate processes?
  3. Which component owns capacity, and why?
  4. Where does CPU time go in the engine process at high batch?
  5. What is prepare_inputs and why does it matter?
  6. In a TP=4 deployment, where does the scheduler run and how do the other ranks know what to do?
  7. What breaks if you run the API and engine in one Python process?

12. Further reading#

  • [REFERENCE] vLLM source: entrypoints/openai/api_server.py, engine/llm_engine.py, core/scheduler.py, worker/model_runner.py
  • [ESTABLISHED] Yu et al., “Orca” (OSDI 2022) — the architecture that started this
  • [ESTABLISHED] Kwon et al., PagedAttention (SOSP 2023) §5 — vLLM’s system design
  • Next: 02 — APIs

↑↓ navigate↵ openesc close