★ This one misconception has probably wasted more GPU-hours than any other in the field.
1. What nvidia-smi actually reports#
$ nvidia-smi
| GPU Name | GPU-Util |
| 0 H100 | 100% |“GPU-Util” is the percentage of time in the last sampling period during which at least one kernel was executing.
That’s it. Not:
- ✗ how many SMs were busy
- ✗ how many arithmetic units were active
- ✗ how much of the memory bandwidth was used
- ✗ whether any useful work was done
A kernel using 1 of 132 SMs, stalled on memory 99% of the time,
computing something you don't need:
GPU-Util: 100%2. Why this matters so much#
Because it is the metric everyone reaches for, and it is systematically misleading in the direction of complacency.
"GPU is at 100%, so we're maxed out — we need more GPUs."
Reality: the GPU is at 100% "utilization", 8% SM throughput,
22% memory throughput, and running at batch 6 when memory
allows 60.
Actual headroom: 5-10x.Teams buy hardware because of this metric. It is worth being emphatic about.
3. Simple analogy#
A factory’s “lights on” metric.
The lights are on 100% of the time. Is the factory productive? The metric cannot distinguish:
- every machine running at capacity
- one machine running, 131 idle
- every machine running, producing scrap
“Lights on” tells you the building is occupied. It tells you nothing about output.
4. What to measure instead#
LEVEL 1 — Is work happening? (nvidia-smi's answer)
utilization.gpu % of time a kernel was resident
LEVEL 2 — How much of the chip? (DCGM / Nsight)
DCGM_FI_PROF_SM_ACTIVE fraction of SMs with at least one warp
DCGM_FI_PROF_SM_OCCUPANCY average resident warps / max
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE fraction of cycles tensor cores active
DCGM_FI_PROF_DRAM_ACTIVE fraction of cycles the memory interface
was transferring
LEVEL 3 — Is the work useful? (application metrics)
average running batch size / max_num_seqs
KV cache utilization
tokens per second
fraction of generated tokens actually delivered (vs aborted)Level 3 is what actually matters, and it’s application-specific, so no GPU monitoring tool provides it. You must emit it.
5. The DCGM metrics that are worth collecting#
METRIC MEANING HEALTHY (LLM decode)
DCGM_FI_PROF_SM_ACTIVE SMs with resident warps 0.8 - 1.0
DCGM_FI_PROF_SM_OCCUPANCY warp slots filled 0.2 - 0.6
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE tensor core cycles 0.02 - 0.15 ← low is
NORMAL for decode
DCGM_FI_PROF_DRAM_ACTIVE memory interface busy 0.6 - 0.9 ← the key one
DCGM_FI_PROF_PCIE_TX/RX_BYTES PCIe traffic near zero in steady state
DCGM_FI_PROF_NVLINK_TX/RX_BYTES NVLink traffic scales with TP
DCGM_FI_DEV_POWER_USAGE watts 400-700 (H100)
DCGM_FI_DEV_SM_CLOCK actual clock check for throttlingDCGM_FI_PROF_DRAM_ACTIVE is the single best GPU-side health metric for LLM decode. It
directly measures the binding constraint. A well-tuned decode-heavy system sits at 0.7-0.9.
Note the tensor-core row: low tensor activity during decode is correct, not a problem. Decode is memory-bound; the tensor cores have nothing to do. Alerting on low tensor utilization for a decode workload generates false alarms.
Deploy dcgm-exporter as a DaemonSet and scrape these into Prometheus.
6. The composite metric that actually means something#
EFFECTIVE UTILIZATION =
(fraction of time doing useful work)
× (fraction of the machine used while doing it)
× (fraction of the output actually consumed)
Practically, for LLM serving:
effective = (busy_time / wall_time)
× (avg_running_batch / max_possible_batch)
× (1 - aborted_token_fraction)
Example:
busy 92% of the time
average batch 14 of a possible 72
11% of tokens generated for clients who disconnected
effective = 0.92 × 0.194 × 0.89 = 16%
nvidia-smi says: 100%.16% versus 100%. That gap is where the money is, and no GPU monitoring tool will show it to you.
7. Worked example — the false capacity crisis#
REPORT: "All 64 GPUs at 100% utilization. We need to double the fleet."
INVESTIGATION:
DCGM_FI_PROF_DRAM_ACTIVE: 0.24 ← 24% of memory bandwidth
DCGM_FI_PROF_SM_ACTIVE: 0.97 ← SMs have work resident
DCGM_FI_PROF_PIPE_TENSOR: 0.03 ← normal for decode
Application metrics:
average running batch: 7 (max_num_seqs = 96)
KV cache utilization: 9%
queue depth: 0
DIAGNOSIS
Not capacity-limited. Traffic is spread across 16 replicas, each getting
4 req/s. Each replica reads the full model per step to serve 7 sequences.
FIX
Consolidate 16 replicas → 5 replicas with prefix-aware routing.
Average batch: 7 → 26. DRAM_ACTIVE: 0.24 → 0.71.
Throughput per GPU: 3.4x.
RESULT: freed 44 of 64 GPUs. No hardware purchased.This is Section IX.12 way 5, diagnosed through metrics. The 100% utilization number actively concealed the problem.
8. Production implications#
- Never autoscale on
utilization.gpu. It’s ~100% whenever anything runs. Use KV utilization or queue wait (Section VIII.07). - Put
DCGM_FI_PROF_DRAM_ACTIVEand average running batch on your primary dashboard. - Compute and track effective utilization. It’s the number that connects engineering to cost.
- Don’t alert on low tensor-core utilization for decode workloads. It’s expected.
- Educate stakeholders. “The GPU is at 100%” will come up in every capacity conversation. Have the two-sentence explanation ready.
- Deploy dcgm-exporter. The Level 2 metrics are not available any other way.
9. Common mistakes#
Treating nvidia-smi utilization as a capacity signal. The central error.
Autoscaling on it. It never varies, so the autoscaler never acts (or always acts).
Alerting on low tensor utilization during decode. False alarms.
Not measuring average running batch size. It’s the most informative single number for LLM serving efficiency and almost nobody collects it.
Concluding “we’re maxed out” without checking DRAM_ACTIVE.
Ignoring the aborted-token fraction. 10-30% of your “utilization” may be work nobody wanted.
10. Hands-on exercise#
A. Demonstrate the lie. Write a kernel that uses one SM and spins. Run it. Observe
nvidia-smi reporting 100% utilization. Then measure DCGM_FI_PROF_SM_ACTIVE.
B. Deploy DCGM. Set up dcgm-exporter and Prometheus. Collect the metrics from section 5
for a real workload. Build a dashboard.
C. Compute effective utilization. For a real system, compute all three factors and the
product. How far is it from nvidia-smi’s number?
D. The consolidation test. Split a fixed request rate across 1, 2, 4, 8 replicas. For each,
record utilization.gpu, DRAM_ACTIVE, average batch, and total throughput. Show that
utilization.gpu is constant while the others vary enormously.
E. Write the explainer. Draft the two-paragraph explanation you’d send to a manager who says “the GPUs are at 100%, we need more.” Include the metrics you’d point them to instead.
11. Interview questions#
- What does
nvidia-smi’s GPU-Util actually measure? - Give an example where it reads 100% and the GPU is doing almost nothing useful.
- What metrics would you use instead, and what does each tell you?
- Why is low tensor-core utilization normal during LLM decode?
- Define effective utilization for an LLM service.
- Why can’t you autoscale on GPU utilization?
- A team says all their GPUs are at 100% and they need more. What do you check?
12. Further reading#
- [REFERENCE] NVIDIA DCGM documentation and the DCGM field identifiers reference
- [REFERENCE]
dcgm-exporterfor Prometheus - [REFERENCE]
nvidia-smimanual — read the definition of utilization.gpu yourself - Next: 05 — Memory fragmentation and allocators