PidokuInfra

Measuring a GPU

Intermediate 1h Difficulty 3/5 Topic 05 of 05

Prerequisites 01, III.05

The idea in one minute#

You cannot improve what you have not measured, and GPU metrics are easy to misread. The most famous one — “GPU utilization” — does not mean what almost everyone assumes. This lesson covers what each common metric really measures, how to read them from Go, and a method for going from “it is slow” to a cause.

A picture#

flowchart TB
  S["It is slow"] --> Q1{"GPU utilization low<br/>and a CPU core at 100%?"}
  Q1 -->|"yes"| R1["Launch-bound or host-bound<br/>fusion, graphs, faster host code"]
  Q1 -->|"no"| Q2{"Memory nearly full<br/>or OOM errors?"}
  Q2 -->|"yes"| R2["Capacity-bound<br/>smaller batch, quantize, more GPUs"]
  Q2 -->|"no"| Q3{"Intensity below<br/>the ridge?"}
  Q3 -->|"yes"| R3["Memory-bound<br/>batch more, move fewer bytes"]
  Q3 -->|"no"| R4["Compute-bound<br/>lower precision, tensor cores"]
  Q1 -.-> T["Also check: temperature,<br/>clock throttling, PCIe link"]
  class S neutral
  class Q1,Q2,Q3 queue
  class R1,R4 compute
  class R2,R3 memory
  class T warn

How it really works#

What “GPU utilization” means#

nvidia-smi reports utilization as the percentage of the sample period during which at least one kernel was running. That is all.

It says nothing about how many SMs were busy, how much of the arithmetic was used, or how much bandwidth. A single thread running a trivial kernel continuously shows 100%. From IV.01, a batch-1 LLM decode uses well under 1% of the arithmetic and also shows 100%.

So:

  • Low utilization is informative: the GPU is idle part of the time (starved by the host, by data loading, or by launch overhead).
  • High utilization is not: it means “something was running”, not “the GPU is being used well”.

For capacity decisions use throughput (tokens/s, requests/s) against the roofline, not utilization.

The metrics worth watching#

MetricMeaningUse it to detect
Utilization (GPU)% of time any kernel ranIdle gaps, starvation
Utilization (memory)% of time memory was being read/writtenNot how full memory is
Memory used / totalBytes allocated by processesHeadroom, leaks, OOM risk
Power drawWattsReal load (a better activity signal than utilization)
Temperature, clock°C, MHzThermal throttling
ECC errors, XID eventsHardware faultsFailing GPUs
PCIe link gen/widthNegotiated host linkA card stuck at x4 or gen 1

Power draw is underrated: a GPU at 100% utilization drawing 150 W of a possible 700 W is not working hard.

The tools#

ToolLevelTells you
nvidia-smiWhole GPU, sampledState at a glance
NVML / go-nvmlSame data, programmaticFeed dashboards and schedulers
DCGM + dcgm-exporterFleetPrometheus metrics for every GPU in a cluster
Nsight SystemsTimeline of one processWhen each kernel, copy and CPU function ran; idle gaps
Nsight ComputeInside one kernelOccupancy, coalescing, cache behaviour
Framework profilersPer operationWhich layer or op is slow

Work top-down: system metrics first, then a timeline, and a kernel-level profiler only once a specific kernel is the suspect.

Benchmarking rules#

  1. Synchronize before starting and stopping the clock (III.04).
  2. Warm up. The first calls include one-time costs: library initialization, memory pool growth, kernel compilation.
  3. Repeat and report a distribution, not one number: median and p99.
  4. Hold everything else still: same input sizes, nothing else on the GPU, a steady temperature.
  5. Compare with the roofline. A measurement without an expected value is just a number.

Code#

The idiomatic way to read GPU state from Go is NVIDIA’s go-nvml. This is a minimal sampler — the seed of every Go GPU monitor, scheduler extension and autoscaler signal.

Go
// sampler.go — sample GPU state with NVIDIA's official Go bindings.
//
//	go mod init sampler && go get github.com/NVIDIA/go-nvml/pkg/nvml && go run .
package main

import (
	"fmt"
	"log"
	"time"

	"github.com/NVIDIA/go-nvml/pkg/nvml"
)

func main() {
	if ret := nvml.Init(); ret != nvml.SUCCESS {
		log.Fatalf("NVML init failed: %v", nvml.ErrorString(ret))
	}
	defer nvml.Shutdown()

	count, _ := nvml.DeviceGetCount()
	for range time.Tick(time.Second) {
		for i := 0; i < count; i++ {
			dev, _ := nvml.DeviceGetHandleByIndex(i)
			name, _ := dev.GetName()
			mem, _ := dev.GetMemoryInfo()
			util, _ := dev.GetUtilizationRates()
			temp, _ := dev.GetTemperature(nvml.TEMPERATURE_GPU)
			milliwatts, _ := dev.GetPowerUsage()
			limit, _ := dev.GetEnforcedPowerLimit()

			fmt.Printf("gpu%d %-20s util %3d%%  mem %5.1f/%5.1f GB  %2d C  power %3.0f/%3.0f W (%2.0f%%)\n",
				i, name, util.Gpu,
				float64(mem.Used)/1e9, float64(mem.Total)/1e9, temp,
				float64(milliwatts)/1000, float64(limit)/1000,
				100*float64(milliwatts)/float64(limit))
		}
	}
}

The last column — power as a fraction of the limit — is the honest “how hard is it working” signal that utilization fails to give.

And a benchmark harness that follows the rules above, for any function, with no GPU needed:

Go
// bench.go — warm up, repeat, report a distribution.
package main

import (
	"fmt"
	"slices"
	"time"
)

type Result struct{ Median, P99, Min time.Duration }

// Bench assumes fn is synchronous: it returns only when its work is finished.
// For GPU work, fn must end with a synchronize.
func Bench(warmup, iters int, fn func()) Result {
	for i := 0; i < warmup; i++ {
		fn()
	}
	samples := make([]time.Duration, iters)
	for i := range samples {
		t0 := time.Now()
		fn()
		samples[i] = time.Since(t0)
	}
	slices.Sort(samples)
	return Result{samples[iters/2], samples[iters*99/100], samples[0]}
}

func main() {
	buf := make([]float32, 4<<20)
	r := Bench(5, 200, func() {
		for i := range buf {
			buf[i] = buf[i]*1.0001 + 1
		}
	})
	fmt.Printf("median %v  p99 %v  min %v\n", r.Median, r.P99, r.Min)
}

Remember this#

  • “GPU utilization” = fraction of time any kernel was running. 100% does not mean fully used.
  • Low utilization is a real signal; high utilization is not.
  • Watch memory used, power draw and temperature alongside it.
  • Benchmark with sync, warm-up, repetition, and a roofline expectation.

Try it#

  1. Run bench.go. How far is p99 from the median? What causes the spread on your machine?
  2. (GPU) Run sampler.go while a workload runs. Compare utilization with power fraction.
  3. A dashboard shows GPU utilization 100%, power 30% of limit, one CPU core at 100%. Using the decision tree, what is the likely cause?

Check yourself#

  1. Define GPU utilization precisely.
  2. Why is power draw a useful complement to utilization?
  3. List three rules for a trustworthy GPU benchmark.

↑↓ navigate↵ openesc close