1. What is it?#
Everything that happens to a model between “it exists in storage” and “it is serving traffic correctly,” and back again.
STORED → FETCHED → LOADED → INITIALIZED → WARMED → READY → DRAINING → UNLOADEDEach transition has a cost, a failure mode, and a decision to make.
2. Why it matters#
Because the lifecycle determines your cold start (which determines your autoscaling, which determines your headroom, which is a large fraction of your bill), and because most “mysterious first-request slowness” is a warmup problem.
3. The phases in detail#
Fetch#
Object storage → local disk (or straight to memory)
Decisions:
- Cache on the node? (yes: a hostPath or local PVC shared by pods)
- Parallel download? (yes: 8-32 streams)
- Verify integrity? (yes: checksums; a corrupted shard produces garbage output,
not an error)
- Which format? (safetensors — mmap-able, no pickle risk)Load#
Disk → host memory → GPU memory
Techniques:
- mmap the safetensors file (lazy, page-cache friendly)
- parallel shard reading
- pinned staging buffers (2x transfer bandwidth)
- direct-to-GPU where supported (GPUDirect Storage)
- with TP: each rank loads only its shard → N-way parallel
Layout conversion happens here:
- weight permutation for quantized kernels (Marlin etc.)
- QKV / gate-up concatenation for fused GEMMs
- dtype conversion if needed
Do it ONCE here, never per request.Initialize#
- Allocate the KV cache pool (profile activation memory first, then claim
the remainder)
- Initialize NCCL communicators (TP/PP)
- Build the tokenizer
- Set up the sampler and any grammar cachesThe KV pool sizing step deserves attention: the engine runs a memory profiling pass with a
synthetic worst-case batch to measure peak activation memory, then allocates the rest to KV.
This is why max_model_len and max_num_batched_tokens affect the reported block count.
Warm up#
Why: the FIRST request through any code path is slow.
- CUDA kernels are JIT-compiled/loaded on first use
- cuBLAS autotuning runs on first call for each shape
- Memory allocators grow their pools
- CUDA graphs must be captured
- torch.compile compiles (if not cached)
Without warmup, the first real user request can be 10-100x slower.A proper warmup exercises every code path:
// Warmup sends real requests through every path before the replica is marked ready.
func Warmup(ctx context.Context, c *Client, maxLen int, batchSizes []int) error {
// 1. Prefill of various lengths (hits different kernel selections)
for _, n := range []int{16, 128, 512, 2048, maxLen / 2, maxLen} {
if err := c.Generate(ctx, Request{Prompt: dummyPrompt(n), MaxTokens: 1}); err != nil {
return err
}
}
// 2. Decode at each captured batch size: bs concurrent requests
for _, bs := range batchSizes {
g, gctx := errgroup.WithContext(ctx)
for i := 0; i < bs; i++ {
g.Go(func() error { return c.Generate(gctx, Request{Prompt: dummyPrompt(64), MaxTokens: 8}) })
}
if err := g.Wait(); err != nil {
return err
}
}
// 3. Every sampling path you support
for _, p := range []Sampling{Greedy, TopP, TopK, WithPenalties, GuidedJSON} {
if err := c.Generate(ctx, Request{Prompt: dummyPrompt(64), MaxTokens: 4, Sampling: p}); err != nil {
return err
}
}
// 4. Long generation (exercises KV growth and block allocation)
return c.Generate(ctx, Request{Prompt: dummyPrompt(64), MaxTokens: 256})
}Most warmup implementations only do step 1. Then the first request using guided JSON, or the first one at batch 47, is slow — and it’s a real user’s request.
Ready#
Readiness probe returns 200 only after warmup completes. Traffic begins.
Drain#
1. Readiness → false (LB stops sending new requests)
2. Continue serving in-flight generations
3. After a deadline, abort whatever remains with a clear error
4. Free KV pool, destroy NCCL communicators, exitUnload#
For multi-model servers: free the weights, return memory to the pool, remove from the registry. The complication is that freeing GPU memory reliably requires that no CUDA graphs or cached allocations reference it — engines usually restart the worker process instead of trying to truly unload.
4. Tiny example — measuring the phases#
// coldstart.go — start a server process and time each phase of its cold start.
//
// go run coldstart.go vllm serve Qwen/Qwen2.5-1.5B-Instruct
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"os"
"os/exec"
"time"
)
const base = "http://localhost:8000"
func generate(model string, maxTokens int) error {
body := fmt.Sprintf(`{"model":%q,"prompt":"test","max_tokens":%d}`, model, maxTokens)
resp, err := http.Post(base+"/v1/completions", "application/json", bytes.NewBufferString(body))
if err != nil {
return err
}
defer resp.Body.Close()
_, err = io.Copy(io.Discard, resp.Body)
return err
}
func main() {
model := os.Args[len(os.Args)-1]
marks := []struct {
name string
at time.Time
}{{"start", time.Now()}}
mark := func(name string) {
marks = append(marks, struct {
name string
at time.Time
}{name, time.Now()})
}
srv := exec.Command(os.Args[1], os.Args[2:]...)
srv.Stdout, srv.Stderr = os.Stderr, os.Stderr
if err := srv.Start(); err != nil {
panic(err)
}
defer srv.Process.Kill()
mark("process started")
for { // the health endpoint answers once weights are loaded and the engine is up
if resp, err := http.Get(base + "/health"); err == nil && resp.StatusCode == 200 {
break
}
time.Sleep(200 * time.Millisecond)
}
mark("healthy (loaded)")
generate(model, 1)
mark("first request")
for i := 0; i < 10; i++ {
generate(model, 32)
}
mark("steady (10 more)")
for i := 1; i < len(marks); i++ {
fmt.Printf("%-18s %7.1f s\n", marks[i].name, marks[i].at.Sub(marks[i-1].at).Seconds())
}
}Typical output for an 8B model:
process started 0.0 s
healthy (loaded) 60.6 s ← imports + CUDA init, then weights + KV pool + graph capture
first request 6.8 s ← JIT, autotuning, allocator growth
steady (10 more) 1.2 s ← 10 requestsThe 6.8-second first request is why warmup exists. Without it, that latency lands on a user.
5. Multi-model considerations#
When one server hosts several models (Section VIII.09), the lifecycle becomes a cache management problem:
Load policy:
EAGER load all models at startup. Simple, needs memory for all.
LAZY load on first request. Saves memory, first request pays the cost.
LRU keep the N most recently used loaded; evict the rest.
Eviction is expensive: reloading costs the full load time.
So: hysteresis (don't evict a model used in the last T minutes),
and prefer evicting models with cheap reload cost.The economics: if a model takes 60 s to load and is requested every 5 minutes, keeping it loaded costs GPU memory; evicting it costs 60 s of latency every 5 minutes. Compute the threshold explicitly.
6-9. Under the hood, performance, production, mistakes#
Under the hood — why the first request is slow:
1. Kernel module loading: CUDA loads cubins lazily per kernel on first launch.
A model with 200 distinct kernels loads them all on the first pass.
2. cuBLAS heuristic/autotuning: first call for each (M,N,K,dtype) shape.
3. Allocator: PyTorch's caching allocator calls cudaMalloc (slow, ~100 µs
and synchronizing) until its pool is large enough.
4. NCCL: first collective initializes the communication rings.
5. Any lazily-imported Python module.CUDA_MODULE_LOADING=EAGER forces all kernels to load at init rather than lazily — trades
startup time for a faster first request. Worth setting if warmup is thorough anyway.
Performance:
Without warmup: first request 5-30x slower than steady state
With warmup: first request within 10% of steady state
Warmup cost: 5-30 seconds of startupProduction:
- Warm up every code path, not just a single prefill.
- Gate readiness on warmup completion.
- Cache weights on the node. Biggest cold-start lever.
- Do layout conversion at load time, never per request.
- Verify checksums. A corrupted shard produces plausible garbage, not an error.
- Instrument each phase and alarm on regression.
- Set a generous
terminationGracePeriodSecondsand implement a drain endpoint. - For multi-model: implement LRU with hysteresis and measure the eviction rate.
Mistakes:
- No warmup. First user request is 10x slow.
- Incomplete warmup. The first guided-JSON request is slow.
- Readiness true before warmup. Traffic arrives too early.
- Loading weights from object storage per pod. Minutes and money.
- Quantizing or compiling at startup.
- No checksum verification.
- Aggressive model eviction in multi-model servers, causing reload thrashing.
10. Hands-on exercise#
A. Measure the phases. Instrument a real server’s startup as in section 4. Produce a breakdown. Which phase would you attack first?
B. Prove warmup matters. Measure the first request’s latency with and without warmup, for several code paths (short prefill, long prefill, batch 32 decode, guided JSON). Which paths need explicit warmup?
C. Optimize loading. Implement parallel shard loading with pinned buffers. Compare to the naive path. How close to disk bandwidth do you get?
D. Node cache. Set up a local weight cache (a hostPath populated by an init container). Measure cold start with a cold and a warm node cache.
E. Drain test. Implement a drain endpoint and test it: start long generations, trigger a drain, verify none are truncated and that new requests are refused.
11. Interview questions#
- Why is the first request to a freshly-loaded model slow? List the causes.
- What should a complete warmup exercise?
- Why does the engine run a memory profiling pass before allocating the KV pool?
- How would you make model loading faster? Rank the techniques.
- What’s the correct drain procedure for an LLM replica?
- For a multi-model server, how would you decide when to evict a model?
- Why verify checksums on model weights?
12. Further reading#
- [REFERENCE] safetensors documentation
- [REFERENCE] vLLM startup logs and
--load-formatoptions - [REFERENCE]
CUDA_MODULE_LOADINGdocumentation - Next: 09 — Multi-model serving