Characterize a GPU in ten minutes: produce the handful of numbers from which you can predict any model’s performance on it.
1. What you build#
A single command — python -m gpubench — that emits a JSON report and a roofline plot for
whatever GPU it runs on: memory bandwidth, peak FLOPs per dtype, PCIe transfer rates, kernel
launch overhead, the decode ridge point, and a predicted-vs-measured tokens/s for a real model.
Diagram — From measurements to a prediction#
flowchart LR BW["Measure bandwidth"] --> RF["Roofline"] FL["Measure peak FLOPs"] --> RF RF --> RP["Ridge point"] HO["Host costs<br/>PCIe, launch, allocation"] --> PR["Predict tokens/s<br/>for a real model"] RP --> PR PR --> CMP["Compare with measured"] class BW,RP memory class FL compute class HO io class RF,PR,CMP neutral
2. Why it matters#
Datasheets give peak numbers under ideal conditions; your cloud instance, driver, and neighbours give something else. Engineers who carry their own measured constants make capacity predictions that hold; engineers who quote datasheets get surprised.
This is the numbers.md habit turned into a tool. Run it on every new instance type before
you trust it.
3. Read first#
- VI.04 — Bandwidth and roofline
- VI.12 — Profiling CUDA
- X.02 — Roofline in practice
- X.04 — GPU utilization myths
- X.07 — Benchmarking
4. Spec#
Bench What Unit
bandwidth large tensor copy / add (≥ 25% of VRAM) GB/s
gemm_peak square GEMM n=8192, per dtype fp32/fp16/bf16 TFLOP/s
gemm_sweep n = 64 … 16384 TFLOP/s vs n
gemv_decode (B × d) @ (d × d), d=4096, B = 1 … 512 TFLOP/s, GB/s vs B
h2d / d2h pageable and pinned, 1 MB … 1 GB GB/s
launch_overhead empty-ish kernel, 10k launches µs / launch
alloc torch.empty of varied sizes, cold and cached µs
attention scaled_dot_product_attention, T = 128 … 8192 ms, tokens/s
model_decode real model, batch 1 … 64 tok/s, ms/step
Output: report.json (+ env: GPU name, driver, CUDA, torch, power limit, temperature)
roofline.png with every benchmark placed on itEvery benchmark: warm up, ≥20 repetitions, torch.cuda.synchronize() around the timed region,
report median and p95.
5. Milestones#
- Timer harness with warmup, sync, and repetition. Test it on a
sleepto make sure it reports what you think. - Bandwidth and peak FLOPs. These two define the roofline. Compare both to the datasheet; explain any gap over 15%.
- Roofline plot. Compute ceiling, bandwidth slope, ridge point at
peak_flops / bandwidthFLOPs per byte. - Decode-shaped sweep. Batched GEMV at model-like width. Find the batch size at which throughput stops scaling linearly — the measured ridge. This answers Checkpoint C question 5 for your hardware.
- Host-side costs. PCIe, launch overhead, allocation.
- Prediction. For a real model:
step_time ≈ weight_bytes / bandwidthat batch 1. Predict tokens/s, then measure. Report the ratio. - Environment capture. Without it the report is not reproducible and therefore not useful.
6. Starter skeleton#
def bench(fn, warmup=5, reps=20):
for _ in range(warmup): fn()
torch.cuda.synchronize()
ts = []
for _ in range(reps):
t0 = time.perf_counter(); fn(); torch.cuda.synchronize()
ts.append(time.perf_counter() - t0)
return {"median": statistics.median(ts), "p95": sorted(ts)[int(0.95 * len(ts)) - 1]}
def bandwidth_gbs(nbytes=8 << 30):
a = torch.empty(nbytes // 4, dtype=torch.float32, device="cuda").normal_()
b = torch.empty_like(a)
t = bench(lambda: b.copy_(a))["median"]
return 2 * nbytes / t / 1e9 # read a + write b
def gemm_tflops(n=8192, dtype=torch.float16):
A = torch.randn(n, n, device="cuda", dtype=dtype); B = torch.randn_like(A)
t = bench(lambda: A @ B)["median"]
return 2 * n**3 / t / 1e12
def decode_sweep(d=4096, layers=8, dtype=torch.float16):
Ws = [torch.randn(d, d, device="cuda", dtype=dtype) for _ in range(layers)]
for B in (1, 2, 4, 8, 16, 32, 64, 128, 256, 512):
x = torch.randn(B, d, device="cuda", dtype=dtype)
def f():
y = x
for W in Ws: y = y @ W
t = bench(f)["median"]
yield B, B / t, layers * d * d * Ws[0].element_size() / t / 1e9 # rows/s, GB/s of weights7. What to measure#
| Measurement | Expectation to write down first |
|---|---|
| Bandwidth vs datasheet | 80-95% of rated |
| Peak FP16/BF16 TFLOP/s vs datasheet | Tensor cores; often well below the marketing number |
| FP32 vs FP16 GEMM | Large gap — tensor cores (VI.11) |
| GEMM TFLOP/s vs n | Tiny n is dominated by launch overhead |
| Decode sweep: rows/s vs B | Linear, then bends at the ridge |
| Pinned vs pageable H2D | 2-3× |
| Launch overhead | ~5-15 µs; × kernels per step = host floor |
| Predicted vs measured tok/s at batch 1 | Within ~1.5× if the model is sound |
nvidia-smi utilization during the batch-1 decode | ~100% while mostly idle on FLOPs (X.04) |
8. Done when#
- One command produces
report.jsonandroofline.png. - Bandwidth and peak FLOPs are within a defensible distance of the datasheet, or you know why not (power cap, thermal throttling, shared host, old driver).
- You located the decode ridge point by measurement.
- Your batch-1 tokens/s prediction and measurement are both in the report.
- You ran it on two different GPUs (or instance types) and compared.
9. Common pitfalls#
No synchronize. You measure launch, not execution.
Tensors too small for the bandwidth test. They fit in cache-like structures or are dominated by overhead. Go big.
Thermal or power throttling mid-run. Log nvidia-smi --query-gpu=temperature.gpu,power.draw,clocks.sm.
Including RNG or allocation in the timed region.
Reporting one run. Variance on shared cloud GPUs is real; report the spread.
Trusting nvidia-smi utilization as a measure of how hard the GPU is working. It is the
fraction of time any kernel was running.
10. No GPU? Do this instead#
Run the same suite against the CPU (device="cpu") and against a Colab T4, and compare the two
rooflines. The method is the deliverable; the device is incidental.
11. Stretch goals#
- Multi-GPU: add
nccl-tests-style all-reduce bandwidth and P2P copy between devices (IX.07). - Add Nsight Compute collection for one kernel and report
dram__throughputdirectly. - Track results over time in CI to catch driver or image regressions.
- Publish a tokens/s-per-dollar column using current on-demand prices (XI.03).
12. Interview questions this project answers#
- How do you measure a GPU’s memory bandwidth?
- Where is the ridge point for decode on your hardware, and what does it tell you about batch size?
- Given weight bytes and bandwidth, estimate tokens/s at batch 1.
nvidia-smisays 100%. What do you measure next?