PidokuInfra

GPU Telemetry

Advanced 1h Difficulty 3/5 Topic 02 of 07

Prerequisites 01; helpful: GPU Engineering II.05 and IV.05

The idea in one minute#

A GPU reports four kinds of thing: how busy it is, how much memory is in use, its power and temperature, and its errors. The first two are the ones people look at, and both mislead for LLM serving: “utilization” only means something was running, and memory is almost fully allocated by the engine at startup.

The numbers worth watching are the less famous ones: power draw against the limit, throttling, memory errors, Xid events, link errors, and — if you want a real activity measure — the profiling counters that say how much of the time the tensor cores and the memory interface were actually active.

An analogy#

A factory’s “machine on” light tells you the machine is powered, not how many parts it is making. The electricity meter, the temperature gauge and the fault log tell you far more about whether it is working hard and whether it is about to break.

A picture#

flowchart LR
  GPU[("GPU")] --> DRV["Driver"]
  DRV --> NVML["NVML library<br/>what nvidia-smi reads"]
  NVML --> DCGM["DCGM host engine<br/>sampling, health, profiling counters"]
  DCGM --> EXP["dcgm-exporter<br/>/metrics, adds pod and container labels"]
  NVML --> GO["go-nvml<br/>your own exporter"]
  DRV --> KLOG["Kernel log<br/>Xid events"]
  EXP --> PROM["Prometheus or<br/>OTel Collector"]
  GO --> PROM
  KLOG --> LOGS["Log pipeline"]
  class GPU memory
  class DRV,NVML,DCGM neutral
  class EXP,GO,PROM io
  class KLOG,LOGS warn

How it really works#

Where the numbers come from#

LayerWhat it isUse it for
NVMLNVIDIA’s management library; nvidia-smi is a front end to itAd-hoc inspection; your own exporter via go-nvml
DCGMA daemon on top of NVML adding sampling, health checks, diagnostics and hardware profiling countersFleet monitoring
dcgm-exporterExposes DCGM fields as Prometheus metrics; in Kubernetes it labels each GPU with the pod using itThe standard path; installed by the GPU Operator
Kernel logWhere the driver reports Xid eventsHardware and driver faults

Other vendors follow the same shape: AMD’s device metrics exporter for Instinct GPUs, Intel’s tooling for Gaudi, and cloud-provider metrics for their own accelerators. The fields differ; the categories below do not.

The fields, and how much to trust each#

Names are DCGM field names as exposed by dcgm-exporter. The exporter’s default-counters.csv decides which are enabled — several of the most useful ones are off by default.

CategoryFieldMeaningTrust
UtilizationDCGM_FI_DEV_GPU_UTIL% of the sample period in which any kernel was runningLow. 100% says “not idle”, nothing more
DCGM_FI_DEV_MEM_COPY_UTIL% of time memory was being read or written — not how full it isLow, and widely misread
Activity (profiling)DCGM_FI_PROF_GR_ENGINE_ACTIVEFraction of time the compute engine was activeBetter
DCGM_FI_PROF_PIPE_TENSOR_ACTIVEFraction of time the tensor cores were activeGood signal of real matrix work
DCGM_FI_PROF_DRAM_ACTIVEFraction of time the memory interface was activeGood signal of memory-bound work
MemoryDCGM_FI_DEV_FB_USED, _FREE, _RESERVEDDevice memory in MiBTrue, but an engine pre-allocates: it is high from the first second
PowerDCGM_FI_DEV_POWER_USAGEWatts nowHigh. The most honest “how hard is it working” gauge
DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTIONMillijoules since driver load (a counter)High. rate() gives average watts; the basis of tokens per joule (lesson 06)
ThermalDCGM_FI_DEV_GPU_TEMP, DCGM_FI_DEV_MEMORY_TEMP°CHigh
ThrottlingDCGM_FI_DEV_POWER_VIOLATION, DCGM_FI_DEV_THERMAL_VIOLATIONTime spent throttled for power or heat (off by default)High. Enable them
ClocksDCGM_FI_DEV_SM_CLOCKCore clock in MHzA drop means throttling
FaultsDCGM_FI_DEV_XID_ERRORSThe last Xid code seen (a gauge, not a count)Use for alerting on change; get details from the kernel log
Memory healthECC single/double-bit counters (off by default); DCGM_FI_DEV_CORRECTABLE_REMAPPED_ROWS, DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS, DCGM_FI_DEV_ROW_REMAP_FAILUREMemory cells failing and being remappedHigh. Rising counts predict failure
LinksDCGM_FI_DEV_PCIE_REPLAY_COUNTER; NVLink CRC / replay / recovery counters (off by default); DCGM_FI_DEV_NVLINK_BANDWIDTH_TOTALRetries on PCIe; errors and traffic on NVLinkHigh. Errors here hurt multi-GPU jobs first
Host link trafficDCGM_FI_PROF_PCIE_TX_BYTES, _RX_BYTESBytes/s between host and GPUUseful: steady-state decoding should move little

Why “GPU utilization” misleads#

The classic figure is sampled: was at least one kernel executing during the interval? An LLM server decoding for a single user launches kernels continuously, so it reads ~100% while serving one-fiftieth of the tokens the GPU could deliver with a full batch. The GPU is busy but not saturated.

What to use instead, in order of usefulness:

  1. The engine’s own metrics: tokens per second, batch size, queue depth (lesson 03).
  2. Power draw relative to the card’s limit: light work draws visibly less.
  3. The profiling counters: tensor and memory-interface activity.
  4. Recent vLLM versions also export estimated FLOPs and bytes moved per GPU, which can be compared with the hardware’s peak — a direct “how much of the roofline are we using” figure.

GPU Engineering IV.05 and Inference Engineering X.04 explain the mechanism.

Xid events#

An Xid is a numbered error report from the driver in the kernel log. A short field guide:

XidMeaningUsual response
13, 31Application-level fault (bad memory access in a kernel)Usually software; investigate the workload
48Double-bit ECC errorDrain the node; the GPU needs a reset or replacement
63, 64Row remapping recorded / failedReset at next opportunity / replace if failed
74NVLink errorCheck link counters; affects tensor-parallel jobs first
79GPU fell off the busNode needs a reboot; check power and seating
94, 95Contained / uncontained ECC errorRestart the affected process / drain the node

Scrape Xids from the kernel log into your log pipeline as structured events with the node and GPU UUID, and alert on the hardware ones. NVIDIA’s documentation maintains the full list.

In Kubernetes#

  • The GPU Operator installs the driver, the container toolkit, device discovery and dcgm-exporter together.
  • dcgm-exporter joins each GPU to its pod, namespace and container, which is the join key to engine metrics. Check that this mapping works with your allocation mode: with Dynamic Resource Allocation, sharing or MIG, verify the labels you get rather than assuming them.
  • Node health should feed scheduling: a GPU with rising errors is tainted and drained before it fails a request. Node Problem Detector and NVIDIA’s open-source NVSentinel automate detect → cordon → drain → remediate; with DRA, device taints (stable in Kubernetes 1.37) let a single bad GPU be excluded without removing the node.

A minimal alert set#

PromQL
# Hardware fault reported (the gauge changed to a non-zero Xid)
DCGM_FI_DEV_XID_ERRORS != 0

# Running hot for 10 minutes
avg_over_time(DCGM_FI_DEV_GPU_TEMP[10m]) > 85

# Memory cells being remapped
increase(DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS[1h]) > 0

# A GPU allocated to a pod but drawing idle power for 30 minutes: paid for, not working
avg_over_time(DCGM_FI_DEV_POWER_USAGE{pod!=""}[30m]) < 100

# One GPU slower than its peers in the same node: clock well below the group's maximum
DCGM_FI_DEV_SM_CLOCK < 0.8 * on (Hostname) group_left max by (Hostname) (DCGM_FI_DEV_SM_CLOCK)

Thresholds are starting points; the idle-power floor differs by card. Adapt them to your hardware and check label names against your exporter version.

Code#

The utilization illusion, simulated: one busy flag versus the measures that track real work.

Go
// gpuutil.go — why "GPU utilization" reads 100% long before the GPU is doing all it can.
package main

import "fmt"

func main() {
	// One decode step reads all the weights once, whatever the batch size (memory-bound),
	// until arithmetic becomes the limit. Illustrative figures for one model on one GPU.
	const (
		stepMsMemoryBound = 20.0 // time for one step when limited by reading weights
		msPerSeqCompute   = 0.25 // extra arithmetic time per sequence in the batch
		idleWatts         = 90.0
		maxWatts          = 700.0
	)
	fmt.Println("batch  step ms  tokens/s  'GPU util'  tensor active  power W  tokens per joule")
	for _, batch := range []int{1, 4, 16, 64, 128, 256} {
		compute := msPerSeqCompute * float64(batch)
		step := stepMsMemoryBound
		if compute > step {
			step = compute // past the ridge point: arithmetic is the limit
		}
		tps := float64(batch) / step * 1000

		util := 100.0 // a kernel is always running: the classic gauge is pinned
		tensorActive := compute / step
		power := idleWatts + (maxWatts-idleWatts)*(0.35+0.65*tensorActive)
		fmt.Printf("%5d  %7.1f  %8.0f  %9.0f%%  %12.0f%%  %7.0f  %12.2f\n",
			batch, step, tps, util, tensorActive*100, power, tps/power)
	}
	fmt.Println("\nUtilization never moved. Throughput, tensor activity, power and tokens per joule did.")
}

Remember this#

  • GPU_UTIL means “a kernel was running”. It is not load, and not saturation.
  • GPU memory is pre-allocated by the engine; watch the engine’s KV-cache usage instead.
  • Power, energy, throttling, ECC/remap counts, Xids and link errors are the trustworthy hardware signals — and several are disabled by default.
  • Join GPU metrics to pods so they can be read next to engine metrics.

Try it#

  1. Run gpuutil.go. At which batch size does tensor activity reach 100%? Compare with the ridge-point idea in GPU Engineering IV.01.
  2. On a machine with an NVIDIA GPU, run nvidia-smi dmon -s pucvmet while a model serves one user, then many. Which columns change? No GPU: read the dcgm-exporter default counter list and mark the fields you would enable.
  3. Write an alert that fires when one GPU in a node draws 30% less power than its siblings under the same workload.

Check yourself#

  1. What does DCGM_FI_DEV_GPU_UTIL actually measure?
  2. Why is DCGM_FI_DEV_FB_USED near its maximum on an idle inference server?
  3. Which hardware counters predict a failure before it happens?

↑↓ navigate↵ openesc close