1. The problem#
GIVEN:
a cluster of N nodes with GPUs of various types
M models, each with resource requirements and traffic demand
tenant quotas and priorities
DECIDE:
which models run on which nodes, and how many replicas of each
CONSTRAINTS:
GPU count, type, and NVLink topology per model
changing a placement costs minutes (model loading)
traffic demand shifts continuouslyThis is a bin-packing problem with expensive moves and a moving target. The expensive-moves property is what makes it different from ordinary Kubernetes scheduling.
2. Why the default scheduler isn’t enough#
KUBERNETES DEFAULT SCHEDULER
✓ places pods on nodes with enough free resources
✗ doesn't know that moving a model costs 5 minutes
✗ doesn't know about NVLink topology (without the Topology Manager)
✗ doesn't know traffic demand per model
✗ doesn't know that consolidating traffic improves efficiency (IX.12)
✗ doesn't know model loading is faster on a node with the weights cached
✗ spreads by default, which is often wrong for inferenceThe “spread by default” behavior is actively harmful here. Spreading replicas across nodes divides traffic, lowers batch size, and reduces throughput per GPU (Section IX.12, way 5). For inference you often want to concentrate traffic on fewer, better-utilized replicas.
3. The placement objective#
MINIMIZE: total GPUs used
SUBJECT TO: every model's SLO is met at its demand
tenant quotas are respected
placement changes are minimized
Equivalently:
MAXIMIZE: Σ (useful tokens produced) / (GPUs allocated)The scoring function:
score(model, node) =
+ w1 × (weights already cached on this node) ← fast startup
+ w2 × (NVLink domain fits the model's TP degree) ← performance
+ w3 × (GPU type matches preference)
+ w4 × (node's NUMA topology aligns)
- w5 × (fragmentation created by this placement)
- w6 × (co-location with a noisy neighbour)
+ w7 × (locality to the model's traffic source, if multi-region)w1 (cached weights) is worth more than people expect. Placing a model on a node that
already has its weights cached turns a 6-minute start into a 40-second one — which changes what
placement decisions you can make dynamically.
4. Placement strategies#
STATIC PLACEMENT
Assign models to nodes manually or by a one-time plan.
✓ simple, predictable
✗ doesn't adapt to demand
→ correct for a small number of high-traffic models
DEMAND-PROPORTIONAL
replicas(model) = ceil(demand(model) / capacity_per_replica / target_util)
Recompute hourly; enact changes gradually.
✓ adapts; simple to reason about
→ THE USUAL ANSWER for a medium platform
BIN-PACKING WITH HYSTERESIS
Solve the packing problem, but only enact a change if the improvement
exceeds a threshold (because moves cost minutes).
✓ efficient
✗ more complex; needs careful hysteresis or it thrashes
TIERED
hot models: dedicated replicas, always resident
warm models: shared pool with LRU model swapping
cold models: load on demand, accept the latency
✓ handles a long tail of rarely-used models well
→ THE ANSWER when you have 50+ models with skewed trafficThe tiered approach reflects reality: in almost every multi-model platform, 5 models are 95% of traffic and 45 models are 5%. Treating them uniformly wastes GPUs on the tail.
5. The hot/warm/cold tiering, concretely#
MEASURE: requests per model per hour over a week.
Model req/hour tier placement
chat-large 42,000 hot 6 dedicated replicas
chat-small 18,000 hot 3 dedicated
code-assist 9,500 hot 2 dedicated
summarize 2,100 warm 1 dedicated (shares a node)
translate 340 warm shared pool
legal-v2 45 cold load on demand
...40 more < 50 cold shared pool, LRU
CAPACITY
hot: 11 dedicated replicas × their GPU counts
warm: 2 nodes running a shared pool with 3-4 models resident
cold: 1 node, LRU with ~6 models resident, load-on-demand for the rest
COLD-TIER SLA
"first request to a cold model may take up to 90 seconds"
→ published, and acceptable for models used 45 times/hourPublishing a different SLA for cold models is the key enabler. Without it, you must keep everything hot and the tail dominates your GPU count.
6. Model swapping in the shared pool#
When a request arrives for a model not currently loaded:
1. Is there a free GPU slot? → load there
2. Else: evict the LRU model
- only if it has no in-flight requests
- and hasn't been used in the last T minutes (hysteresis)
- prefer evicting models with cheap reload (cached weights, small)
3. Load the requested model (from the node cache if possible)
4. Queue the request; serve when ready
VIABILITY (Section VIII.09):
load_time vs inter-arrival time for that model
load 40 s, requests every 5 min → 13% of time loading. Acceptable.
load 40 s, requests every 60 s → 67% loading. Thrashing. Pin it instead.Add hysteresis and a pin list. A model that would thrash should be promoted to the warm tier permanently, and the swap logic should detect and report that.
7. Fragmentation#
PROBLEM: 8-GPU nodes, and models needing 8, 4, 2, and 1 GPUs.
node0: [8-GPU model ] perfect
node1: [4-GPU][2][1][1] perfect
node2: [4-GPU][2] [ 2 free ] a 4-GPU model can't fit
node3: [1][1][1][1][1] [ 3 free ] fragmented
Total free: 5 GPUs. Largest contiguous NVLink-domain-aligned block: 2.
→ an 8-GPU model cannot be placed despite 5 free GPUs.Mitigations:
1. STANDARDIZE ON A FEW SIZES
Prefer models that need 1, 2, 4, or 8 GPUs. Avoid 3, 5, 6, 7.
(This is also what NVLink topology and TP head-count constraints want.)
2. NODE SPECIALIZATION
Designate some nodes for 8-GPU models, others for smaller ones.
Reduces flexibility, eliminates fragmentation.
3. PACK LARGEST FIRST
Place the 8-GPU models before the 1-GPU ones.
4. PERIODIC DEFRAGMENTATION
Rebalance during low-traffic windows. Expensive (model reloads)
but occasionally necessary.Point 1 is the cheapest fix and it’s a policy decision, not an algorithm. Requiring power-of-two GPU counts eliminates most fragmentation.
8. Implementing it#
OPTION A — Kubernetes-native
a controller that watches the registry and traffic metrics,
and adjusts Deployment replica counts and node affinities.
✓ uses existing machinery
✗ limited control over exact placement
OPTION B — custom scheduler plugin
a Kubernetes scheduler plugin implementing the scoring function.
✓ real placement control
✗ more work; must track scheduler API changes
OPTION C — external placement controller
computes the desired state, writes it as node affinities / labels,
and lets Kubernetes enact it.
✓ decoupled, testable
→ THE PRAGMATIC CHOICE for most platforms// Sketch of the controller loop (Option C)
func reconcile(ctx context.Context) {
demand := metrics.DemandPerModel(time.Hour)
current := k8s.CurrentPlacement(ctx)
capacity := registry.CapacityPerReplica() // measured, per model
desired := map[string]int{}
for model, d := range demand {
targetUtil := tierConfig(model).TargetUtilization
desired[model] = int(math.Ceil(d / capacity[model] / targetUtil))
}
plan := pack(desired, cluster.Nodes(), scoringFn)
changes := 0
for _, change := range diff(current, plan) {
if change.Benefit < hysteresisThreshold {
continue // not worth a 5-minute move
}
if changes++; changes > maxChurnPerHour {
break // rate-limit churn
}
enact(ctx, change) // gradually
}
}The two guards — a benefit threshold and a churn limit — are what prevent the controller from thrashing. Without them, normal demand variation causes continuous replacement.
9. Production implications#
- Measure per-model traffic first. The distribution determines the strategy.
- Tier your models. Hot/warm/cold with different SLAs.
- Standardize on power-of-two GPU counts to avoid fragmentation.
- Node-local weight caches make placement changes cheap enough to do dynamically.
- Add hysteresis and churn limits. A thrashing placement controller is worse than a static plan.
- Publish a different SLA for cold models. It’s what makes the long tail affordable.
- Prefer concentration over spreading for the same model (Section IX.12).
- Rebalance during low-traffic windows, not at peak.
10. Common mistakes#
Using default Kubernetes spreading. Divides traffic, lowers batch size.
Treating all models the same. The tail wastes GPUs.
No hysteresis. The controller thrashes.
Odd GPU counts. Fragmentation.
Not caching weights on nodes. Every placement change costs minutes.
Placing without checking NVLink topology.
Rebalancing at peak traffic.
No pin list. A hot model gets evicted from the shared pool and thrashes.
11. Hands-on exercise#
A. Measure the distribution. For a multi-model workload (real or synthetic), plot requests per model. What fraction of models account for 95% of traffic? Design the tiering.
B. Implement the controller. Build the reconcile loop from section 8 in simulation. Include hysteresis and churn limits. Feed it a week of synthetic demand and measure: GPUs used, SLO violations, and placement changes.
C. Fragmentation. Simulate placement with models needing {1,2,3,4,5,8} GPUs on 8-GPU nodes. Measure the wasted GPUs. Then restrict to powers of two and re-measure.
D. Swapping viability. For a set of cold models with measured load times and request rates, determine which are viable for LRU swapping and which must be pinned. Simulate and measure time spent loading.
E. Scoring function. Implement the scoring function from section 3. Tune the weights on a simulated cluster. Which weight matters most?
12. Interview questions#
- Why isn’t the default Kubernetes scheduler sufficient for model placement?
- Why is “spread” the wrong default for inference replicas?
- Design a scoring function for placing a model on a node.
- What is the hot/warm/cold tiering and why does it matter?
- When is LRU model swapping viable? Give the arithmetic.
- How does GPU count fragmentation arise, and what’s the cheapest fix?
- What prevents a placement controller from thrashing?
13. Further reading#
- [REFERENCE] Kubernetes scheduler framework documentation
- [REFERENCE] Volcano, Kueue for batch and gang scheduling
- [FUNDAMENTAL] Bin-packing heuristics (first-fit-decreasing, best-fit)
- [ESTABLISHED] Borg / Omega / Kubernetes scheduling papers
- Next: 05 — Multi-tenancy and quotas