1. What is it?#
- Process — an isolated address space with its own memory, file descriptors, and at least one thread. Isolation is strong; communication is expensive.
- Thread — an execution context (registers, stack, program counter) sharing the process’s address space. Communication is free; isolation is nonexistent.
- Context switch — the kernel saving one thread’s state and restoring another’s. Costs 1-10 µs directly, plus cache/TLB pollution that can cost much more.
For inference: the engine is typically one process per GPU (or per TP rank), with threads for HTTP handling, tokenization, and the engine loop — and Python’s GIL complicating all of it.
2. Why does it exist?#
Because you have more work than cores, and because work has different isolation and communication needs. Processes for fault isolation and for escaping the GIL; threads for cheap sharing of the big data structures (like a KV cache manager) that you cannot afford to copy.
3. Simple analogy#
A workshop. A process is a separate workshop with its own tools and locked door — safe, but lending a tool means walking it across town. A thread is a colleague in your workshop — instant sharing, and also instant ability to knock over your work.
A context switch is being interrupted mid-task: you must write down where you were, and when you come back your workbench has been rearranged (cold caches).
4. Tiny example#
Two threads of CPU-bound work, in Go:
// threads.go — CPU-bound work on one goroutine, then on two.
package main
import (
"fmt"
"sync"
"time"
)
//go:noinline
func cpuWork(n int) (x int) {
for i := 0; i < n; i++ {
x += i & 3
}
return x
}
func main() {
const n = 2_000_000_000
// Serial
t0 := time.Now()
cpuWork(n)
cpuWork(n)
serial := time.Since(t0)
// Parallel: goroutines are scheduled onto real OS threads, one per core
t0 = time.Now()
var wg sync.WaitGroup
for i := 0; i < 2; i++ {
wg.Add(1)
go func() { defer wg.Done(); cpuWork(n) }()
}
wg.Wait()
parallel := time.Since(t0)
fmt.Printf("serial %v 2 goroutines %v speedup %.2fx\n",
serial.Round(time.Millisecond), parallel.Round(time.Millisecond), float64(serial)/float64(parallel))
}In Go the speedup is ≈ 2.0x: the runtime schedules goroutines onto real OS threads, one per core, and they genuinely run at the same time.
Write the same program with Python’s threading module and, on CPython ≤3.12, the speedup is
≈ 1.0x (sometimes worse). Python threads do not run Python bytecode in parallel. Only one
thread holds the Global Interpreter Lock (GIL) at a time. This matters to you even as a Go
engineer, because most inference engines (vLLM, TGI, SGLang) are Python processes, and the GIL
shapes their architecture.
Inside such an engine, this still runs in parallel:
import torch
# GIL is RELEASED during this call — the C++ code runs in parallel
y = big_tensor @ big_matrixRule: the GIL is released around C extension calls, I/O, and CUDA launches. So Python threads are fine for an inference server’s I/O and for overlapping GPU work, and useless for CPU-bound Python logic. That single fact explains most of the threading design in vLLM, TGI, and friends.
5. Technical explanation#
Costs#
Thread creation ~10-50 µs
Process creation(fork) ~100 µs - 1 ms
Context switch 1-10 µs direct
+ cache pollution up to 100s of µs indirect (refilling L1/L2, TLB)
Mutex uncontended ~20 ns
Mutex contended ~1-10 µs (may involve a syscall + sleep)
Atomic increment ~5-20 ns (worse under contention)
Pipe/socket IPC ~5-20 µs round trip
Shared memory ~free after setupThe thread model of an inference server#
A typical vLLM-style server:
Process: API server (Python, asyncio)
├─ event loop thread HTTP, SSE streaming
├─ tokenizer thread pool releases GIL (Rust tokenizers)
└─ IPC to engine process
Process: Engine (one per TP rank)
├─ main loop thread scheduler + kernel launches
├─ CUDA driver threads internal
└─ NCCL threads collective communicationWhy separate processes? Two reasons: to escape the GIL (the engine loop must not be blocked by HTTP parsing), and because tensor-parallel ranks must be separate processes anyway (each needs its own CUDA context and its own NCCL rank).
Oversubscription — the classic self-inflicted wound#
32-core machine, running:
PyTorch intra-op threads: 32
OpenMP (via MKL): 32
Tokenizer pool: 16
HTTP workers: 8
────────────────────────────────
88 runnable threads on 32 cores → constant context switchingSymptom: high system CPU time, high cs in vmstat, throughput lower than with fewer
threads. Fix: set them explicitly.
export OMP_NUM_THREADS=8
export MKL_NUM_THREADS=8
export TOKENIZERS_PARALLELISM=false # when you already parallelize at the request level
python -c "import torch; torch.set_num_threads(8)"Async vs threads for the API layer#
An inference server’s HTTP layer is I/O-bound and long-lived (streaming responses last
seconds-to-minutes). Thread-per-connection would need thousands of threads. asyncio (or Rust’s
tokio, or Go goroutines) handles this with one or a few threads. This is why almost every
modern inference server’s front end is async.
6. Under the hood#
Watch it happen:
# Context switches per second (cs column)
vmstat 1
# Per-process voluntary/involuntary switches
pidstat -w -p $(pgrep -f vllm) 1
# Thread list with CPU usage
top -H -p $(pgrep -f vllm)
# What are threads waiting on?
cat /proc/<pid>/task/*/stack 2>/dev/null # kernel stacks
py-spy dump --pid <pid> # Python stacks of ALL threadspy-spy dump is the single most useful command for diagnosing a stuck or slow Python inference
server. Learn it now; you will use it in Section X.
Interpreting: voluntary switches mean the thread blocked (I/O, lock, sleep) — usually fine. Involuntary switches mean the scheduler preempted it — high counts mean oversubscription.
7. Performance implications#
- The engine loop is latency-critical. If it gets descheduled for 5 ms, every request’s ITL
suffers by 5 ms. Give it a dedicated core; consider
SCHED_FIFOin extreme cases (file 10). - GIL contention adds jitter. A CPU-heavy Python callback (e.g. a custom stopping criterion) holds the GIL and delays the engine loop.
- Lock contention on the scheduler’s data structures shows up at large batch. Keep critical sections short.
- Thread pools for tokenization should be sized to the CPU budget, not to
os.cpu_count()(which lies inside containers — see file 12).
8. Production implications#
- One engine process per GPU (or per TP group). Do not try to serve two models from one Python process expecting parallelism.
- Set all the thread-count environment variables explicitly in your container image. Relying on defaults inside containers is a reliable way to oversubscribe.
- Pin the engine thread to a NUMA-local core (file 03) on multi-socket hosts.
- Use
py-spyin production. It attaches without restarting and without instrumenting. Include it in your image. - Watch involuntary context switches as a health metric; a spike means CPU contention that will show up as ITL jitter.
9. Common mistakes#
Using Python threads for CPU-bound work. The GIL makes it pointless. Use processes, or push the work into C/Rust.
Leaving OMP_NUM_THREADS unset in a container. OpenMP sees the host’s core count, spawns
128 threads inside a 4-CPU cgroup, and you get pathological throttling.
Blocking the asyncio event loop. One synchronous time.sleep() or a heavy JSON dump in a
coroutine stalls every streaming connection on that loop.
Creating threads per request. At 1,000 req/s that’s 1,000 thread creations/sec plus scheduler pressure. Use pools.
Assuming os.cpu_count() reflects your quota. It doesn’t in containers. Read the cgroup
(file 12).
10. Hands-on exercise#
A. Scaling, and where it stops. Run the example in section 4 with 1, 2, 4, … goroutines up
to twice your core count. Plot the speedup. Where does it flatten, and why? Then set
GOMAXPROCS=1 and run it again: you have just reproduced what the GIL does to a Python engine.
B. Oversubscription. Run a CPU-heavy workload with OMP_NUM_THREADS = 1, 4, 8, 16, 32, 64
on an N-core machine. Plot throughput. Find the peak. Watch vmstat 1’s cs column at each
setting.
C. Profile a real server’s threads. Start any Python inference server under load. Run
top -H, pidstat -w, and py-spy dump. Identify: which thread is the engine loop, which is
handling HTTP, and where CPU is actually going.
D. Measure context-switch cost. Write a ping-pong benchmark between two threads using a condition variable, 1M round trips. Compute µs per switch.
11. Interview questions#
- What is the GIL and when does it not block parallelism?
- Why do inference servers use separate processes for the API layer and the engine?
- What is oversubscription and how does it manifest?
- What is the difference between voluntary and involuntary context switches, diagnostically?
- Why is the API layer of an inference server usually async rather than thread-per-request?
- How would you debug a Python inference server that has become unresponsive?
12. Further reading#
- [FUNDAMENTAL] Gregg, Systems Performance, ch. 6
- [REFERENCE]
py-spydocumentation — https://github.com/benfred/py-spy - [REFERENCE] Python
asynciodocs; CPython GIL design notes (and PEP 703 on removing it) - Next: 06 — Memory allocation and virtual memory