PidokuInfra

Kubernetes for GPUs

Expert Advanced 1h 30m Difficulty 4/5 Topic 03 of 11

Prerequisites II.12, IX.10

★ Most inference platforms run on Kubernetes. Most Kubernetes GPU configurations are subtly wrong in ways that cost 20-40% of performance.


1. How GPUs get into pods#

1. NVIDIA drivers on the node (via the GPU Operator or a node image)
2. NVIDIA Container Toolkit (the runtime hook that injects devices)
3. NVIDIA Device Plugin (a DaemonSet) advertises nvidia.com/gpu
4. Pod requests the resource; kubelet allocates whole GPUs
5. The container runtime injects /dev/nvidia* and driver libraries
YAML
resources:
  limits:
    nvidia.com/gpu: 8      # GPUs are limits-only; no fractional requests

GPUs are allocated as whole units and are not oversubscribable. A pod requesting 8 GPUs gets 8 exclusively. (MIG and time-slicing change this — see section 6.)

Diagram — How a GPU reaches a container#

flowchart TB
  HW["GPU on the node"] --> DRV["NVIDIA driver + container toolkit<br/>installed by the GPU Operator"]
  DRV --> ADV["Device plugin or DRA driver<br/>advertises GPUs to the kubelet"]
  ADV --> API["API server: node capacity"]
  POD["Pod requests a GPU"] --> SCH["Scheduler picks a node"]
  API --> SCH
  SCH --> KUB["Kubelet allocates the device"]
  KUB --> CTR["Container starts with<br/>device + driver libraries mounted"]

  class HW,DRV compute
  class ADV,API,SCH,KUB queue
  class POD neutral
  class CTR memory

2. The four configurations that matter#

These are the ones that are usually wrong.

(a) Guaranteed QoS + CPU Manager static policy#

YAML
resources:
  requests:
    cpu: "64"            # INTEGER, and requests == limits
    memory: "512Gi"
    nvidia.com/gpu: 8
  limits:
    cpu: "64"
    memory: "512Gi"
    nvidia.com/gpu: 8
Requires on the kubelet:
  --cpu-manager-policy=static
  --reserved-cpus=0-3          (reserve some for system daemons)

WHAT THIS BUYS YOU
  ✓ a dedicated cpuset — no CFS quota throttling (Section II.10, II.12)
  ✓ no noisy-neighbour CPU contention
  ✓ a prerequisite for NUMA alignment

WITHOUT IT
  ✗ CFS bandwidth throttling → 100 ms latency spikes (Section II.12)
  ✗ the engine loop competes with other pods

This is the single most impactful Kubernetes setting for inference latency, and it’s off by default.

(b) Topology Manager#

kubelet:
  --topology-manager-policy=single-numa-node
  --topology-manager-scope=pod
WHAT IT DOES
  Ensures CPUs, memory, GPUs, and NICs allocated to a pod come from
  the SAME NUMA node.

WHY IT MATTERS (Section II.03)
  ✗ without it: pinned host buffers may be on the wrong NUMA node
    → 1.5-2x slower host-to-device transfers
  ✗ GPUs assigned across NUMA nodes → NCCL crosses the CPU interconnect
    → the Section IX.12 case-5 disaster

CAVEAT: single-numa-node can cause pods to be UNSCHEDULABLE if no single
        NUMA node has enough resources. Use "best-effort" if that's a
        problem, but understand you've given up the guarantee.

(c) /dev/shm sizing#

YAML
volumes:
  - name: dshm
    emptyDir:
      medium: Memory
      sizeLimit: 16Gi
volumeMounts:
  - name: dshm
    mountPath: /dev/shm
Default is 64 MB. NCCL's shared-memory transport, multi-process engines,
and PyTorch's data structures all use it.
Symptom of too little: "Bus error", NCCL init failures, mysterious crashes.

(d) Thread counts from the quota#

YAML
env:
  - name: OMP_NUM_THREADS
    value: "16"           # NOT the node's core count
  - name: MKL_NUM_THREADS
    value: "16"
  - name: TOKENIZERS_PARALLELISM
    value: "false"        # if you parallelize at the request level

Section II.12’s lie: os.cpu_count() returns the node’s cores, not your quota. Libraries that use it spawn far too many threads.


PROBLEM: a pod requesting 8 GPUs might get GPUs 0,1,2,3,8,9,10,11 —
         spanning two NVLink domains on a 16-GPU node.

SOLUTIONS
  1. Request ALL the node's GPUs (nvidia.com/gpu: 8 on an 8-GPU node)
     → simple, guarantees the whole NVLink domain, wastes nothing if
       your model needs 8
  2. Topology Manager with single-numa-node
     → aligns GPUs with NUMA, which usually aligns with NVLink domains
  3. NVIDIA GPU Operator's topology-aware allocation
  4. Custom scheduler plugin scoring by NVLink connectivity

VERIFY inside the pod:
  nvidia-smi topo -m
  → all pairs should show NV#, not SYS

Always verify inside the pod, not on the node. The container sees a renumbered subset.


4. Multi-node deployments#

YAML
apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
  name: llama-405b
spec:
  replicas: 2
  leaderWorkerTemplate:
    size: 2                                    # 2 pods per instance
    restartPolicy: RecreateGroupOnPodRestart   # ← all-or-nothing (Section IX.10)
    leaderTemplate:
      spec:
        schedulerName: volcano                 # gang scheduling
        containers:
          - name: vllm-leader
            resources:
              limits:
                nvidia.com/gpu: 8
                rdma/hca_shared_devices_a: 1   # RDMA device plugin
    workerTemplate: {...}
REQUIRED FOR MULTI-NODE
  □ LeaderWorkerSet (or a JobSet / MPI operator)
  □ RecreateGroupOnPodRestart — one pod failing must restart the group
  □ Gang scheduling (Volcano, Kueue, or scheduling gates) — all pods
    schedulable together, or none
  □ RDMA device plugin + GPUDirect RDMA verified
  □ Headless Service for pod DNS
  □ NCCL environment configured for the right interfaces

Without gang scheduling, a partially-placed multi-node instance occupies GPUs and never becomes ready — a silent capacity leak.


5. Probes#

YAML
startupProbe:                # for the long model load
  httpGet: {path: /health, port: 8000}
  failureThreshold: 60       # 60 × 10s = 10 minutes allowed
  periodSeconds: 10

livenessProbe:               # is the process alive?
  httpGet: {path: /health, port: 8000}
  periodSeconds: 30
  failureThreshold: 3
  timeoutSeconds: 5

readinessProbe:              # can it serve? (false during load and warmup)
  httpGet: {path: /health/ready, port: 8000}
  periodSeconds: 5
  failureThreshold: 2

The startup probe is essential (Section VIII.02): without it, the liveness probe kills the pod during a 5-minute model load and it never starts. This exact failure is extremely common.

YAML
lifecycle:
  preStop:
    exec:
      command: ["/bin/sh", "-c", "curl -X POST localhost:8000/drain; sleep 5"]
terminationGracePeriodSeconds: 900     # ≥ p99 generation time (Section VIII.07)

6. Sharing GPUs#

TIME-SLICING (device plugin config)
  Advertise N virtual GPUs per physical GPU. The driver time-slices.
  ✓ simple, no hardware requirement
  ✗ no isolation; context switching overhead
  ✗ memory is NOT partitioned — one pod can OOM another
  → for development and small models only

MPS (Multi-Process Service)
  A daemon merges processes into one context → concurrent kernels.
  ✓ better utilization than time-slicing
  ✗ weak isolation; a fault in one process can affect others
  → good for cooperative small models (Section VIII.09)

MIG (Multi-Instance GPU)
  Hardware partitioning. The device plugin advertises MIG profiles.
  ✓ hard isolation (SMs, L2, memory)
  ✗ fixed profiles; reconfiguration requires draining the whole GPU
  → for multi-tenant isolation requirements

DRA (Dynamic Resource Allocation)  [ESTABLISHED core, EMERGING extras]
  Kubernetes' attribute-based resource model: pods claim devices by
  constraint ("a GPU with >= 40 GB", "two on one NVLink domain")
  through ResourceClaims instead of an opaque count.
  status (October 2026):
    core API                  stable since Kubernetes 1.34
    device taints/tolerations stable in 1.37 — take one faulty GPU
                              out of scheduling without the node
    prioritized alternatives  stable in 1.36 ("this GPU, else that")
    partitionable devices     still maturing (dynamic MIG-style splits)
  → new clusters should plan on DRA; the device plugin remains
    supported and is still what most existing fleets run
  → your allocation dashboards must count claims and devices, not
    only nvidia.com/gpu requests

7. The complete pod spec#

YAML
apiVersion: apps/v1
kind: Deployment
metadata:
  name: llama-70b
spec:
  replicas: 6
  template:
    spec:
      nodeSelector:
        nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule
      containers:
        - name: vllm
          image: registry/vllm:0.x.y-cuda12.4
          args: ["--model", "/models/llama-3-70b-fp8",
                 "--tensor-parallel-size", "8",
                 "--max-model-len", "8192",
                 "--max-num-seqs", "160",
                 "--max-num-batched-tokens", "2048",
                 "--gpu-memory-utilization", "0.90",
                 "--enable-prefix-caching",
                 "--enable-chunked-prefill",
                 "--kv-cache-dtype", "fp8"]
          resources:
            requests: {cpu: "64", memory: "512Gi", nvidia.com/gpu: 8}
            limits:   {cpu: "64", memory: "512Gi", nvidia.com/gpu: 8}
          env:
            - {name: OMP_NUM_THREADS,   value: "16"}
            - {name: NCCL_TIMEOUT,      value: "600"}
            - {name: TORCH_NCCL_ASYNC_ERROR_HANDLING, value: "1"}
            - {name: PYTORCH_CUDA_ALLOC_CONF, value: "expandable_segments:True"}
            - {name: VLLM_WORKER_MULTIPROC_METHOD, value: "spawn"}
          volumeMounts:
            - {name: models, mountPath: /models, readOnly: true}
            - {name: dshm,   mountPath: /dev/shm}
          startupProbe:
            httpGet: {path: /health, port: 8000}
            failureThreshold: 60
            periodSeconds: 10
          readinessProbe:
            httpGet: {path: /health/ready, port: 8000}
            periodSeconds: 5
          lifecycle:
            preStop:
              exec: {command: ["/bin/sh","-c","curl -XPOST localhost:8000/drain; sleep 5"]}
      terminationGracePeriodSeconds: 900
      volumes:
        - name: models
          hostPath: {path: /var/cache/models}      # node-local cache
        - name: dshm
          emptyDir: {medium: Memory, sizeLimit: 16Gi}

Every line in this spec addresses something covered earlier in the curriculum. Trace each back; that’s the exercise.


8. Production implications#

  • Guaranteed QoS with the static CPU Manager policy. Biggest latency win.
  • Topology Manager single-numa-node where schedulable.
  • Size /dev/shm. 16 GB.
  • Set thread counts from the quota, not the node.
  • Startup probes for model loading. Otherwise pods never start.
  • Long termination grace period + drain hook. Otherwise every deploy truncates generations.
  • Node-local model cache via hostPath or a local PV.
  • Gang scheduling for multi-node.
  • Verify topology inside the pod, not on the node.
  • Use the GPU Operator rather than managing drivers manually.

9. Common mistakes#

No startup probe. Liveness kills the pod during model load; it never starts. The most common Kubernetes GPU mistake.

Burstable QoS (requests ≠ limits). CFS throttling → 100 ms latency spikes.

Default 64 MB /dev/shm. NCCL failures.

Thread counts from nproc. Massive oversubscription.

Short terminationGracePeriodSeconds. Truncated generations on every deploy.

Multi-GPU pods spanning NVLink domains.

No gang scheduling for multi-node. Partial placements leak GPUs.

Pulling model weights from object storage per pod. Minutes of cold start.

CPU limits set on latency-critical pods.


10. Hands-on exercise#

A. Audit a real spec. Take an existing GPU pod spec (yours or a public example) and check it against section 8’s list. How many items are missing?

B. Demonstrate throttling. Deploy the same workload with Burstable and Guaranteed QoS. Measure nr_throttled and p99 latency for each.

C. Verify topology. Deploy a multi-GPU pod and run nvidia-smi topo -m inside it. Are all pairs NVLink-connected? If not, fix it with the Topology Manager and re-verify.

D. Break the probe. Deploy without a startup probe and with a liveness probe that’s too aggressive. Watch the pod crashloop during model load. Fix it.

E. Test the drain. Roll a deployment while streaming requests are in flight, with a 30 s and a 900 s grace period. Count truncated generations in each case.

F. Node cache. Implement the node-local model cache with a DaemonSet. Measure cold start with and without it.


11. Interview questions#

  1. How do GPUs get allocated to a pod in Kubernetes?
  2. Why does Guaranteed QoS matter for inference latency?
  3. What does the Topology Manager do and why does it matter for multi-GPU?
  4. Why do you need a startup probe for LLM pods?
  5. What terminationGracePeriodSeconds would you set, and why?
  6. What’s required for a multi-node deployment that isn’t for single-node?
  7. Compare time-slicing, MPS, and MIG for sharing a GPU.

12. Further reading#

  • [REFERENCE] Kubernetes CPU Manager and Topology Manager documentation
  • [REFERENCE] NVIDIA GPU Operator and device plugin documentation
  • [REFERENCE] Kubernetes LeaderWorkerSet; Kueue and Volcano for gang scheduling
  • [ESTABLISHED] Kubernetes Dynamic Resource Allocation (DRA) — core API stable since 1.34; check the release notes of your version for the status of each sub-feature
  • Next: 04 — GPU scheduling and model placement

↑↓ navigate↵ openesc close