★ Most inference platforms run on Kubernetes. Most Kubernetes GPU configurations are subtly wrong in ways that cost 20-40% of performance.
1. How GPUs get into pods#
1. NVIDIA drivers on the node (via the GPU Operator or a node image)
2. NVIDIA Container Toolkit (the runtime hook that injects devices)
3. NVIDIA Device Plugin (a DaemonSet) advertises nvidia.com/gpu
4. Pod requests the resource; kubelet allocates whole GPUs
5. The container runtime injects /dev/nvidia* and driver librariesresources:
limits:
nvidia.com/gpu: 8 # GPUs are limits-only; no fractional requestsGPUs are allocated as whole units and are not oversubscribable. A pod requesting 8 GPUs gets 8 exclusively. (MIG and time-slicing change this — see section 6.)
Diagram — How a GPU reaches a container#
flowchart TB HW["GPU on the node"] --> DRV["NVIDIA driver + container toolkit<br/>installed by the GPU Operator"] DRV --> ADV["Device plugin or DRA driver<br/>advertises GPUs to the kubelet"] ADV --> API["API server: node capacity"] POD["Pod requests a GPU"] --> SCH["Scheduler picks a node"] API --> SCH SCH --> KUB["Kubelet allocates the device"] KUB --> CTR["Container starts with<br/>device + driver libraries mounted"] class HW,DRV compute class ADV,API,SCH,KUB queue class POD neutral class CTR memory
2. The four configurations that matter#
These are the ones that are usually wrong.
(a) Guaranteed QoS + CPU Manager static policy#
resources:
requests:
cpu: "64" # INTEGER, and requests == limits
memory: "512Gi"
nvidia.com/gpu: 8
limits:
cpu: "64"
memory: "512Gi"
nvidia.com/gpu: 8Requires on the kubelet:
--cpu-manager-policy=static
--reserved-cpus=0-3 (reserve some for system daemons)
WHAT THIS BUYS YOU
✓ a dedicated cpuset — no CFS quota throttling (Section II.10, II.12)
✓ no noisy-neighbour CPU contention
✓ a prerequisite for NUMA alignment
WITHOUT IT
✗ CFS bandwidth throttling → 100 ms latency spikes (Section II.12)
✗ the engine loop competes with other podsThis is the single most impactful Kubernetes setting for inference latency, and it’s off by default.
(b) Topology Manager#
kubelet:
--topology-manager-policy=single-numa-node
--topology-manager-scope=podWHAT IT DOES
Ensures CPUs, memory, GPUs, and NICs allocated to a pod come from
the SAME NUMA node.
WHY IT MATTERS (Section II.03)
✗ without it: pinned host buffers may be on the wrong NUMA node
→ 1.5-2x slower host-to-device transfers
✗ GPUs assigned across NUMA nodes → NCCL crosses the CPU interconnect
→ the Section IX.12 case-5 disaster
CAVEAT: single-numa-node can cause pods to be UNSCHEDULABLE if no single
NUMA node has enough resources. Use "best-effort" if that's a
problem, but understand you've given up the guarantee.(c) /dev/shm sizing#
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 16Gi
volumeMounts:
- name: dshm
mountPath: /dev/shmDefault is 64 MB. NCCL's shared-memory transport, multi-process engines,
and PyTorch's data structures all use it.
Symptom of too little: "Bus error", NCCL init failures, mysterious crashes.(d) Thread counts from the quota#
env:
- name: OMP_NUM_THREADS
value: "16" # NOT the node's core count
- name: MKL_NUM_THREADS
value: "16"
- name: TOKENIZERS_PARALLELISM
value: "false" # if you parallelize at the request levelSection II.12’s lie: os.cpu_count() returns the node’s cores, not your quota. Libraries
that use it spawn far too many threads.
3. Multi-GPU pods and NVLink topology#
PROBLEM: a pod requesting 8 GPUs might get GPUs 0,1,2,3,8,9,10,11 —
spanning two NVLink domains on a 16-GPU node.
SOLUTIONS
1. Request ALL the node's GPUs (nvidia.com/gpu: 8 on an 8-GPU node)
→ simple, guarantees the whole NVLink domain, wastes nothing if
your model needs 8
2. Topology Manager with single-numa-node
→ aligns GPUs with NUMA, which usually aligns with NVLink domains
3. NVIDIA GPU Operator's topology-aware allocation
4. Custom scheduler plugin scoring by NVLink connectivity
VERIFY inside the pod:
nvidia-smi topo -m
→ all pairs should show NV#, not SYSAlways verify inside the pod, not on the node. The container sees a renumbered subset.
4. Multi-node deployments#
apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
name: llama-405b
spec:
replicas: 2
leaderWorkerTemplate:
size: 2 # 2 pods per instance
restartPolicy: RecreateGroupOnPodRestart # ← all-or-nothing (Section IX.10)
leaderTemplate:
spec:
schedulerName: volcano # gang scheduling
containers:
- name: vllm-leader
resources:
limits:
nvidia.com/gpu: 8
rdma/hca_shared_devices_a: 1 # RDMA device plugin
workerTemplate: {...}REQUIRED FOR MULTI-NODE
□ LeaderWorkerSet (or a JobSet / MPI operator)
□ RecreateGroupOnPodRestart — one pod failing must restart the group
□ Gang scheduling (Volcano, Kueue, or scheduling gates) — all pods
schedulable together, or none
□ RDMA device plugin + GPUDirect RDMA verified
□ Headless Service for pod DNS
□ NCCL environment configured for the right interfacesWithout gang scheduling, a partially-placed multi-node instance occupies GPUs and never becomes ready — a silent capacity leak.
5. Probes#
startupProbe: # for the long model load
httpGet: {path: /health, port: 8000}
failureThreshold: 60 # 60 × 10s = 10 minutes allowed
periodSeconds: 10
livenessProbe: # is the process alive?
httpGet: {path: /health, port: 8000}
periodSeconds: 30
failureThreshold: 3
timeoutSeconds: 5
readinessProbe: # can it serve? (false during load and warmup)
httpGet: {path: /health/ready, port: 8000}
periodSeconds: 5
failureThreshold: 2The startup probe is essential (Section VIII.02): without it, the liveness probe kills the pod during a 5-minute model load and it never starts. This exact failure is extremely common.
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "curl -X POST localhost:8000/drain; sleep 5"]
terminationGracePeriodSeconds: 900 # ≥ p99 generation time (Section VIII.07)6. Sharing GPUs#
TIME-SLICING (device plugin config)
Advertise N virtual GPUs per physical GPU. The driver time-slices.
✓ simple, no hardware requirement
✗ no isolation; context switching overhead
✗ memory is NOT partitioned — one pod can OOM another
→ for development and small models only
MPS (Multi-Process Service)
A daemon merges processes into one context → concurrent kernels.
✓ better utilization than time-slicing
✗ weak isolation; a fault in one process can affect others
→ good for cooperative small models (Section VIII.09)
MIG (Multi-Instance GPU)
Hardware partitioning. The device plugin advertises MIG profiles.
✓ hard isolation (SMs, L2, memory)
✗ fixed profiles; reconfiguration requires draining the whole GPU
→ for multi-tenant isolation requirements
DRA (Dynamic Resource Allocation) [ESTABLISHED core, EMERGING extras]
Kubernetes' attribute-based resource model: pods claim devices by
constraint ("a GPU with >= 40 GB", "two on one NVLink domain")
through ResourceClaims instead of an opaque count.
status (October 2026):
core API stable since Kubernetes 1.34
device taints/tolerations stable in 1.37 — take one faulty GPU
out of scheduling without the node
prioritized alternatives stable in 1.36 ("this GPU, else that")
partitionable devices still maturing (dynamic MIG-style splits)
→ new clusters should plan on DRA; the device plugin remains
supported and is still what most existing fleets run
→ your allocation dashboards must count claims and devices, not
only nvidia.com/gpu requests7. The complete pod spec#
apiVersion: apps/v1
kind: Deployment
metadata:
name: llama-70b
spec:
replicas: 6
template:
spec:
nodeSelector:
nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
containers:
- name: vllm
image: registry/vllm:0.x.y-cuda12.4
args: ["--model", "/models/llama-3-70b-fp8",
"--tensor-parallel-size", "8",
"--max-model-len", "8192",
"--max-num-seqs", "160",
"--max-num-batched-tokens", "2048",
"--gpu-memory-utilization", "0.90",
"--enable-prefix-caching",
"--enable-chunked-prefill",
"--kv-cache-dtype", "fp8"]
resources:
requests: {cpu: "64", memory: "512Gi", nvidia.com/gpu: 8}
limits: {cpu: "64", memory: "512Gi", nvidia.com/gpu: 8}
env:
- {name: OMP_NUM_THREADS, value: "16"}
- {name: NCCL_TIMEOUT, value: "600"}
- {name: TORCH_NCCL_ASYNC_ERROR_HANDLING, value: "1"}
- {name: PYTORCH_CUDA_ALLOC_CONF, value: "expandable_segments:True"}
- {name: VLLM_WORKER_MULTIPROC_METHOD, value: "spawn"}
volumeMounts:
- {name: models, mountPath: /models, readOnly: true}
- {name: dshm, mountPath: /dev/shm}
startupProbe:
httpGet: {path: /health, port: 8000}
failureThreshold: 60
periodSeconds: 10
readinessProbe:
httpGet: {path: /health/ready, port: 8000}
periodSeconds: 5
lifecycle:
preStop:
exec: {command: ["/bin/sh","-c","curl -XPOST localhost:8000/drain; sleep 5"]}
terminationGracePeriodSeconds: 900
volumes:
- name: models
hostPath: {path: /var/cache/models} # node-local cache
- name: dshm
emptyDir: {medium: Memory, sizeLimit: 16Gi}Every line in this spec addresses something covered earlier in the curriculum. Trace each back; that’s the exercise.
8. Production implications#
- Guaranteed QoS with the static CPU Manager policy. Biggest latency win.
- Topology Manager
single-numa-nodewhere schedulable. - Size
/dev/shm. 16 GB. - Set thread counts from the quota, not the node.
- Startup probes for model loading. Otherwise pods never start.
- Long termination grace period + drain hook. Otherwise every deploy truncates generations.
- Node-local model cache via hostPath or a local PV.
- Gang scheduling for multi-node.
- Verify topology inside the pod, not on the node.
- Use the GPU Operator rather than managing drivers manually.
9. Common mistakes#
No startup probe. Liveness kills the pod during model load; it never starts. The most common Kubernetes GPU mistake.
Burstable QoS (requests ≠ limits). CFS throttling → 100 ms latency spikes.
Default 64 MB /dev/shm. NCCL failures.
Thread counts from nproc. Massive oversubscription.
Short terminationGracePeriodSeconds. Truncated generations on every deploy.
Multi-GPU pods spanning NVLink domains.
No gang scheduling for multi-node. Partial placements leak GPUs.
Pulling model weights from object storage per pod. Minutes of cold start.
CPU limits set on latency-critical pods.
10. Hands-on exercise#
A. Audit a real spec. Take an existing GPU pod spec (yours or a public example) and check it against section 8’s list. How many items are missing?
B. Demonstrate throttling. Deploy the same workload with Burstable and Guaranteed QoS.
Measure nr_throttled and p99 latency for each.
C. Verify topology. Deploy a multi-GPU pod and run nvidia-smi topo -m inside it. Are all
pairs NVLink-connected? If not, fix it with the Topology Manager and re-verify.
D. Break the probe. Deploy without a startup probe and with a liveness probe that’s too aggressive. Watch the pod crashloop during model load. Fix it.
E. Test the drain. Roll a deployment while streaming requests are in flight, with a 30 s and a 900 s grace period. Count truncated generations in each case.
F. Node cache. Implement the node-local model cache with a DaemonSet. Measure cold start with and without it.
11. Interview questions#
- How do GPUs get allocated to a pod in Kubernetes?
- Why does Guaranteed QoS matter for inference latency?
- What does the Topology Manager do and why does it matter for multi-GPU?
- Why do you need a startup probe for LLM pods?
- What
terminationGracePeriodSecondswould you set, and why? - What’s required for a multi-node deployment that isn’t for single-node?
- Compare time-slicing, MPS, and MIG for sharing a GPU.
12. Further reading#
- [REFERENCE] Kubernetes CPU Manager and Topology Manager documentation
- [REFERENCE] NVIDIA GPU Operator and device plugin documentation
- [REFERENCE] Kubernetes LeaderWorkerSet; Kueue and Volcano for gang scheduling
- [ESTABLISHED] Kubernetes Dynamic Resource Allocation (DRA) — core API stable since 1.34; check the release notes of your version for the status of each sub-feature
- Next: 04 — GPU scheduling and model placement