The idea in one minute#
A GPU reports four kinds of thing: how busy it is, how much memory is in use, its power and temperature, and its errors. The first two are the ones people look at, and both mislead for LLM serving: “utilization” only means something was running, and memory is almost fully allocated by the engine at startup.
The numbers worth watching are the less famous ones: power draw against the limit, throttling, memory errors, Xid events, link errors, and — if you want a real activity measure — the profiling counters that say how much of the time the tensor cores and the memory interface were actually active.
An analogy#
A factory’s “machine on” light tells you the machine is powered, not how many parts it is making. The electricity meter, the temperature gauge and the fault log tell you far more about whether it is working hard and whether it is about to break.
A picture#
flowchart LR
GPU[("GPU")] --> DRV["Driver"]
DRV --> NVML["NVML library<br/>what nvidia-smi reads"]
NVML --> DCGM["DCGM host engine<br/>sampling, health, profiling counters"]
DCGM --> EXP["dcgm-exporter<br/>/metrics, adds pod and container labels"]
NVML --> GO["go-nvml<br/>your own exporter"]
DRV --> KLOG["Kernel log<br/>Xid events"]
EXP --> PROM["Prometheus or<br/>OTel Collector"]
GO --> PROM
KLOG --> LOGS["Log pipeline"]
class GPU memory
class DRV,NVML,DCGM neutral
class EXP,GO,PROM io
class KLOG,LOGS warnHow it really works#
Where the numbers come from#
| Layer | What it is | Use it for |
|---|---|---|
| NVML | NVIDIA’s management library; nvidia-smi is a front end to it | Ad-hoc inspection; your own exporter via go-nvml |
| DCGM | A daemon on top of NVML adding sampling, health checks, diagnostics and hardware profiling counters | Fleet monitoring |
| dcgm-exporter | Exposes DCGM fields as Prometheus metrics; in Kubernetes it labels each GPU with the pod using it | The standard path; installed by the GPU Operator |
| Kernel log | Where the driver reports Xid events | Hardware and driver faults |
Other vendors follow the same shape: AMD’s device metrics exporter for Instinct GPUs, Intel’s tooling for Gaudi, and cloud-provider metrics for their own accelerators. The fields differ; the categories below do not.
The fields, and how much to trust each#
Names are DCGM field names as exposed by dcgm-exporter. The exporter’s default-counters.csv
decides which are enabled — several of the most useful ones are off by default.
| Category | Field | Meaning | Trust |
|---|---|---|---|
| Utilization | DCGM_FI_DEV_GPU_UTIL | % of the sample period in which any kernel was running | Low. 100% says “not idle”, nothing more |
DCGM_FI_DEV_MEM_COPY_UTIL | % of time memory was being read or written — not how full it is | Low, and widely misread | |
| Activity (profiling) | DCGM_FI_PROF_GR_ENGINE_ACTIVE | Fraction of time the compute engine was active | Better |
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE | Fraction of time the tensor cores were active | Good signal of real matrix work | |
DCGM_FI_PROF_DRAM_ACTIVE | Fraction of time the memory interface was active | Good signal of memory-bound work | |
| Memory | DCGM_FI_DEV_FB_USED, _FREE, _RESERVED | Device memory in MiB | True, but an engine pre-allocates: it is high from the first second |
| Power | DCGM_FI_DEV_POWER_USAGE | Watts now | High. The most honest “how hard is it working” gauge |
DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION | Millijoules since driver load (a counter) | High. rate() gives average watts; the basis of tokens per joule (lesson 06) | |
| Thermal | DCGM_FI_DEV_GPU_TEMP, DCGM_FI_DEV_MEMORY_TEMP | °C | High |
| Throttling | DCGM_FI_DEV_POWER_VIOLATION, DCGM_FI_DEV_THERMAL_VIOLATION | Time spent throttled for power or heat (off by default) | High. Enable them |
| Clocks | DCGM_FI_DEV_SM_CLOCK | Core clock in MHz | A drop means throttling |
| Faults | DCGM_FI_DEV_XID_ERRORS | The last Xid code seen (a gauge, not a count) | Use for alerting on change; get details from the kernel log |
| Memory health | ECC single/double-bit counters (off by default); DCGM_FI_DEV_CORRECTABLE_REMAPPED_ROWS, DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS, DCGM_FI_DEV_ROW_REMAP_FAILURE | Memory cells failing and being remapped | High. Rising counts predict failure |
| Links | DCGM_FI_DEV_PCIE_REPLAY_COUNTER; NVLink CRC / replay / recovery counters (off by default); DCGM_FI_DEV_NVLINK_BANDWIDTH_TOTAL | Retries on PCIe; errors and traffic on NVLink | High. Errors here hurt multi-GPU jobs first |
| Host link traffic | DCGM_FI_PROF_PCIE_TX_BYTES, _RX_BYTES | Bytes/s between host and GPU | Useful: steady-state decoding should move little |
Why “GPU utilization” misleads#
The classic figure is sampled: was at least one kernel executing during the interval? An LLM server decoding for a single user launches kernels continuously, so it reads ~100% while serving one-fiftieth of the tokens the GPU could deliver with a full batch. The GPU is busy but not saturated.
What to use instead, in order of usefulness:
- The engine’s own metrics: tokens per second, batch size, queue depth (lesson 03).
- Power draw relative to the card’s limit: light work draws visibly less.
- The profiling counters: tensor and memory-interface activity.
- Recent vLLM versions also export estimated FLOPs and bytes moved per GPU, which can be compared with the hardware’s peak — a direct “how much of the roofline are we using” figure.
GPU Engineering IV.05 and Inference Engineering X.04 explain the mechanism.
Xid events#
An Xid is a numbered error report from the driver in the kernel log. A short field guide:
| Xid | Meaning | Usual response |
|---|---|---|
| 13, 31 | Application-level fault (bad memory access in a kernel) | Usually software; investigate the workload |
| 48 | Double-bit ECC error | Drain the node; the GPU needs a reset or replacement |
| 63, 64 | Row remapping recorded / failed | Reset at next opportunity / replace if failed |
| 74 | NVLink error | Check link counters; affects tensor-parallel jobs first |
| 79 | GPU fell off the bus | Node needs a reboot; check power and seating |
| 94, 95 | Contained / uncontained ECC error | Restart the affected process / drain the node |
Scrape Xids from the kernel log into your log pipeline as structured events with the node and GPU UUID, and alert on the hardware ones. NVIDIA’s documentation maintains the full list.
In Kubernetes#
- The GPU Operator installs the driver, the container toolkit, device discovery and dcgm-exporter together.
- dcgm-exporter joins each GPU to its
pod,namespaceandcontainer, which is the join key to engine metrics. Check that this mapping works with your allocation mode: with Dynamic Resource Allocation, sharing or MIG, verify the labels you get rather than assuming them. - Node health should feed scheduling: a GPU with rising errors is tainted and drained before it fails a request. Node Problem Detector and NVIDIA’s open-source NVSentinel automate detect → cordon → drain → remediate; with DRA, device taints (stable in Kubernetes 1.37) let a single bad GPU be excluded without removing the node.
A minimal alert set#
# Hardware fault reported (the gauge changed to a non-zero Xid)
DCGM_FI_DEV_XID_ERRORS != 0
# Running hot for 10 minutes
avg_over_time(DCGM_FI_DEV_GPU_TEMP[10m]) > 85
# Memory cells being remapped
increase(DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS[1h]) > 0
# A GPU allocated to a pod but drawing idle power for 30 minutes: paid for, not working
avg_over_time(DCGM_FI_DEV_POWER_USAGE{pod!=""}[30m]) < 100
# One GPU slower than its peers in the same node: clock well below the group's maximum
DCGM_FI_DEV_SM_CLOCK < 0.8 * on (Hostname) group_left max by (Hostname) (DCGM_FI_DEV_SM_CLOCK)Thresholds are starting points; the idle-power floor differs by card. Adapt them to your hardware and check label names against your exporter version.
Code#
The utilization illusion, simulated: one busy flag versus the measures that track real work.
// gpuutil.go — why "GPU utilization" reads 100% long before the GPU is doing all it can.
package main
import "fmt"
func main() {
// One decode step reads all the weights once, whatever the batch size (memory-bound),
// until arithmetic becomes the limit. Illustrative figures for one model on one GPU.
const (
stepMsMemoryBound = 20.0 // time for one step when limited by reading weights
msPerSeqCompute = 0.25 // extra arithmetic time per sequence in the batch
idleWatts = 90.0
maxWatts = 700.0
)
fmt.Println("batch step ms tokens/s 'GPU util' tensor active power W tokens per joule")
for _, batch := range []int{1, 4, 16, 64, 128, 256} {
compute := msPerSeqCompute * float64(batch)
step := stepMsMemoryBound
if compute > step {
step = compute // past the ridge point: arithmetic is the limit
}
tps := float64(batch) / step * 1000
util := 100.0 // a kernel is always running: the classic gauge is pinned
tensorActive := compute / step
power := idleWatts + (maxWatts-idleWatts)*(0.35+0.65*tensorActive)
fmt.Printf("%5d %7.1f %8.0f %9.0f%% %12.0f%% %7.0f %12.2f\n",
batch, step, tps, util, tensorActive*100, power, tps/power)
}
fmt.Println("\nUtilization never moved. Throughput, tensor activity, power and tokens per joule did.")
}Remember this#
GPU_UTILmeans “a kernel was running”. It is not load, and not saturation.- GPU memory is pre-allocated by the engine; watch the engine’s KV-cache usage instead.
- Power, energy, throttling, ECC/remap counts, Xids and link errors are the trustworthy hardware signals — and several are disabled by default.
- Join GPU metrics to pods so they can be read next to engine metrics.
Try it#
- Run
gpuutil.go. At which batch size does tensor activity reach 100%? Compare with the ridge-point idea in GPU Engineering IV.01. - On a machine with an NVIDIA GPU, run
nvidia-smi dmon -s pucvmetwhile a model serves one user, then many. Which columns change? No GPU: read the dcgm-exporter default counter list and mark the fields you would enable. - Write an alert that fires when one GPU in a node draws 30% less power than its siblings under the same workload.
Check yourself#
- What does
DCGM_FI_DEV_GPU_UTILactually measure? - Why is
DCGM_FI_DEV_FB_USEDnear its maximum on an idle inference server? - Which hardware counters predict a failure before it happens?