PidokuInfra

GPUs in Kubernetes

Advanced 1h Difficulty 4/5 Topic 04 of 04

Prerequisites 03, III.05

The idea in one minute#

Kubernetes knows about CPUs and memory out of the box. It knows nothing about GPUs until software on each node advertises them. That software is mostly written in Go, by NVIDIA, on the go-nvml library you met in IV.05. Once advertised, a GPU is a countable resource a pod can request — and the scheduler’s job becomes placing pods so that expensive GPUs are neither idle nor fragmented.

A picture#

flowchart TB
  subgraph NODE["GPU node"]
    DRV["NVIDIA driver"] --> NVML["NVML"]
    NVML --> DP["Device plugin (Go)<br/>advertises GPUs"]
    NVML --> EXP["DCGM exporter<br/>metrics"]
    TK["Container toolkit<br/>injects driver into containers"]
    KUBELET["kubelet"]
    DP -->|"nvidia.com/gpu: 8"| KUBELET
    KUBELET --> TK --> POD["Pod with a GPU"]
  end
  API["API server and scheduler"]
  KUBELET <--> API
  USER["Pod spec<br/>limits: nvidia.com/gpu: 1"] --> API
  EXP --> PROM["Prometheus"]
  class DRV,NVML neutral
  class DP,EXP,TK,KUBELET compute
  class API,USER queue
  class POD memory
  class PROM io

How it really works#

The pieces#

ComponentJob
NVIDIA driverMust exist on the node (III.05)
Container toolkitMakes the driver and devices visible inside containers
Device pluginTells the kubelet how many GPUs exist; on allocation, tells it which device to attach
GPU feature discoveryLabels nodes with GPU model, memory, driver version
DCGM exporterPublishes per-GPU metrics for Prometheus
GPU operatorInstalls and upgrades all of the above for you

Requesting a GPU#

YAML
apiVersion: v1
kind: Pod
metadata:
  name: llm-server
spec:
  nodeSelector:
    nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3   # label from feature discovery
  containers:
    - name: server
      image: my-registry/llm-server:1.0
      resources:
        limits:
          nvidia.com/gpu: 1        # whole GPUs only, in the classic model

In this classic model a GPU is an opaque integer: a pod gets whole devices, cannot ask for “half”, and cannot express “these two must be NVLink neighbours”.

Sharing and richer requests#

  • Time-slicing / MPS can be enabled in the device plugin so one physical GPU is advertised as several. Pods then share it with the isolation properties from lesson 03.
  • MIG instances are advertised as their own resource types (for example nvidia.com/mig-1g.10gb).
  • Dynamic Resource Allocation (DRA) is the newer Kubernetes mechanism in which devices are described by attributes and requested by constraint (“a GPU with at least 40 GB”, “two GPUs on the same NVLink switch”). It replaces the opaque-integer model for anything sophisticated. Its core API has been stable since Kubernetes 1.34; per-device taints (to pull one faulty GPU out of scheduling) became stable in 1.37 (August 2026); splitting one device into independently allocated partitions is still maturing. Check your cluster’s version before relying on a specific feature.

What makes GPU scheduling hard#

  1. GPUs are expensive and lumpy. One stranded GPU costs more per hour than a rack of CPU nodes. Bin-packing (fill nodes before starting new ones) matters more than spreading.
  2. Fragmentation. Eight free GPUs spread over eight nodes cannot run one eight-GPU job.
  3. Topology. A multi-GPU pod needs its GPUs on the same NVLink domain and NUMA node (VI.01).
  4. Gang scheduling. A distributed job needs all of its pods or none; starting half of them wastes GPUs while waiting. The default scheduler does not do this; add-ons (Kueue, Volcano) do.
  5. Cold starts. Pulling a multi-gigabyte image and loading tens of GB of weights takes minutes, so scaling up is slow and scaling to zero is costly.

Health#

Nodes should be drained automatically when a GPU reports uncorrectable memory errors or critical XID events (II.05). A GPU that is “present but broken” is worse than a missing one: the scheduler keeps sending work to it.

Code#

A bin-packing GPU scheduler in Go — the core decision the cluster makes, reduced to its logic.

Go
// schedule.go — spread vs pack: the same pods, the same nodes, different outcomes.
package main

import (
	"fmt"
	"sort"
)

type Node struct {
	Name string
	Free int // free GPUs
}

type Pod struct {
	Name string
	GPUs int
}

// place returns the names of pods that could not be scheduled.
func place(nodes []Node, pods []Pod, pack bool) (pending []string) {
	for _, p := range pods {
		// candidates that can fit this pod
		var fit []int
		for i, n := range nodes {
			if n.Free >= p.GPUs {
				fit = append(fit, i)
			}
		}
		if len(fit) == 0 {
			pending = append(pending, p.Name)
			continue
		}
		sort.Slice(fit, func(a, b int) bool {
			if pack { // least free first: fill nodes up, keep big holes intact
				return nodes[fit[a]].Free < nodes[fit[b]].Free
			}
			return nodes[fit[a]].Free > nodes[fit[b]].Free // most free first: spread
		})
		nodes[fit[0]].Free -= p.GPUs
	}
	return pending
}

func main() {
	pods := []Pod{{"a", 1}, {"b", 1}, {"c", 1}, {"d", 1}, {"train", 8}}
	for _, pack := range []bool{false, true} {
		nodes := []Node{{"n1", 8}, {"n2", 8}, {"n3", 8}, {"n4", 8}}
		// some existing load
		nodes[0].Free, nodes[1].Free, nodes[2].Free = 5, 6, 7
		pending := place(nodes, pods, pack)
		fmt.Printf("pack=%-5v  free GPUs per node: %v  pending: %v\n", pack, nodes, pending)
	}
}

Spreading puts a small pod on the last empty node and the 8-GPU job can never start. Packing keeps one node whole.

Remember this#

  • Kubernetes learns about GPUs from the device plugin; the GPU operator installs the whole stack.
  • Classic requests are whole GPUs (nvidia.com/gpu: N); DRA allows attribute-based requests.
  • Pack GPU workloads; watch for fragmentation, topology and gang-scheduling needs.
  • Drain nodes on GPU hardware errors.

Try it#

  1. Run schedule.go. Add a preferSameNode rule for multi-GPU pods and a second 8-GPU job.
  2. Extend Node with a GPU model label and Pod with a required model. Schedule a mix.
  3. (Cluster) kubectl describe node <gpu-node> — find the nvidia.com/gpu capacity and the labels feature discovery added.

Check yourself#

  1. What does the device plugin do?
  2. Why is bin-packing preferred for GPU workloads?
  3. What problem does gang scheduling solve?

↑↓ navigate↵ openesc close