The idea in one minute#
Kubernetes knows about CPUs and memory out of the box. It knows nothing about GPUs until
software on each node advertises them. That software is mostly written in Go, by NVIDIA, on
the go-nvml library you met in IV.05. Once advertised, a GPU is a countable resource a pod can
request — and the scheduler’s job becomes placing pods so that expensive GPUs are neither idle
nor fragmented.
A picture#
flowchart TB
subgraph NODE["GPU node"]
DRV["NVIDIA driver"] --> NVML["NVML"]
NVML --> DP["Device plugin (Go)<br/>advertises GPUs"]
NVML --> EXP["DCGM exporter<br/>metrics"]
TK["Container toolkit<br/>injects driver into containers"]
KUBELET["kubelet"]
DP -->|"nvidia.com/gpu: 8"| KUBELET
KUBELET --> TK --> POD["Pod with a GPU"]
end
API["API server and scheduler"]
KUBELET <--> API
USER["Pod spec<br/>limits: nvidia.com/gpu: 1"] --> API
EXP --> PROM["Prometheus"]
class DRV,NVML neutral
class DP,EXP,TK,KUBELET compute
class API,USER queue
class POD memory
class PROM ioHow it really works#
The pieces#
| Component | Job |
|---|---|
| NVIDIA driver | Must exist on the node (III.05) |
| Container toolkit | Makes the driver and devices visible inside containers |
| Device plugin | Tells the kubelet how many GPUs exist; on allocation, tells it which device to attach |
| GPU feature discovery | Labels nodes with GPU model, memory, driver version |
| DCGM exporter | Publishes per-GPU metrics for Prometheus |
| GPU operator | Installs and upgrades all of the above for you |
Requesting a GPU#
apiVersion: v1
kind: Pod
metadata:
name: llm-server
spec:
nodeSelector:
nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3 # label from feature discovery
containers:
- name: server
image: my-registry/llm-server:1.0
resources:
limits:
nvidia.com/gpu: 1 # whole GPUs only, in the classic modelIn this classic model a GPU is an opaque integer: a pod gets whole devices, cannot ask for “half”, and cannot express “these two must be NVLink neighbours”.
Sharing and richer requests#
- Time-slicing / MPS can be enabled in the device plugin so one physical GPU is advertised as several. Pods then share it with the isolation properties from lesson 03.
- MIG instances are advertised as their own resource types (for example
nvidia.com/mig-1g.10gb). - Dynamic Resource Allocation (DRA) is the newer Kubernetes mechanism in which devices are described by attributes and requested by constraint (“a GPU with at least 40 GB”, “two GPUs on the same NVLink switch”). It replaces the opaque-integer model for anything sophisticated. Its core API has been stable since Kubernetes 1.34; per-device taints (to pull one faulty GPU out of scheduling) became stable in 1.37 (August 2026); splitting one device into independently allocated partitions is still maturing. Check your cluster’s version before relying on a specific feature.
What makes GPU scheduling hard#
- GPUs are expensive and lumpy. One stranded GPU costs more per hour than a rack of CPU nodes. Bin-packing (fill nodes before starting new ones) matters more than spreading.
- Fragmentation. Eight free GPUs spread over eight nodes cannot run one eight-GPU job.
- Topology. A multi-GPU pod needs its GPUs on the same NVLink domain and NUMA node (VI.01).
- Gang scheduling. A distributed job needs all of its pods or none; starting half of them wastes GPUs while waiting. The default scheduler does not do this; add-ons (Kueue, Volcano) do.
- Cold starts. Pulling a multi-gigabyte image and loading tens of GB of weights takes minutes, so scaling up is slow and scaling to zero is costly.
Health#
Nodes should be drained automatically when a GPU reports uncorrectable memory errors or critical XID events (II.05). A GPU that is “present but broken” is worse than a missing one: the scheduler keeps sending work to it.
Code#
A bin-packing GPU scheduler in Go — the core decision the cluster makes, reduced to its logic.
// schedule.go — spread vs pack: the same pods, the same nodes, different outcomes.
package main
import (
"fmt"
"sort"
)
type Node struct {
Name string
Free int // free GPUs
}
type Pod struct {
Name string
GPUs int
}
// place returns the names of pods that could not be scheduled.
func place(nodes []Node, pods []Pod, pack bool) (pending []string) {
for _, p := range pods {
// candidates that can fit this pod
var fit []int
for i, n := range nodes {
if n.Free >= p.GPUs {
fit = append(fit, i)
}
}
if len(fit) == 0 {
pending = append(pending, p.Name)
continue
}
sort.Slice(fit, func(a, b int) bool {
if pack { // least free first: fill nodes up, keep big holes intact
return nodes[fit[a]].Free < nodes[fit[b]].Free
}
return nodes[fit[a]].Free > nodes[fit[b]].Free // most free first: spread
})
nodes[fit[0]].Free -= p.GPUs
}
return pending
}
func main() {
pods := []Pod{{"a", 1}, {"b", 1}, {"c", 1}, {"d", 1}, {"train", 8}}
for _, pack := range []bool{false, true} {
nodes := []Node{{"n1", 8}, {"n2", 8}, {"n3", 8}, {"n4", 8}}
// some existing load
nodes[0].Free, nodes[1].Free, nodes[2].Free = 5, 6, 7
pending := place(nodes, pods, pack)
fmt.Printf("pack=%-5v free GPUs per node: %v pending: %v\n", pack, nodes, pending)
}
}Spreading puts a small pod on the last empty node and the 8-GPU job can never start. Packing keeps one node whole.
Remember this#
- Kubernetes learns about GPUs from the device plugin; the GPU operator installs the whole stack.
- Classic requests are whole GPUs (
nvidia.com/gpu: N); DRA allows attribute-based requests. - Pack GPU workloads; watch for fragmentation, topology and gang-scheduling needs.
- Drain nodes on GPU hardware errors.
Try it#
- Run
schedule.go. Add apreferSameNoderule for multi-GPU pods and a second 8-GPU job. - Extend
Nodewith a GPU model label andPodwith a required model. Schedule a mix. - (Cluster)
kubectl describe node <gpu-node>— find thenvidia.com/gpucapacity and the labels feature discovery added.
Check yourself#
- What does the device plugin do?
- Why is bin-packing preferred for GPU workloads?
- What problem does gang scheduling solve?