PidokuInfra

Project 11 — GPU Benchmark Suite

Advanced 8h Difficulty 3/5 Topic 11 of 15

Prerequisites Project 03, Section VI (03-07, 11-12), Section X (01-04, 07)

Characterize a GPU in ten minutes: produce the handful of numbers from which you can predict any model’s performance on it.


1. What you build#

A single command — python -m gpubench — that emits a JSON report and a roofline plot for whatever GPU it runs on: memory bandwidth, peak FLOPs per dtype, PCIe transfer rates, kernel launch overhead, the decode ridge point, and a predicted-vs-measured tokens/s for a real model.

Diagram — From measurements to a prediction#

flowchart LR
  BW["Measure bandwidth"] --> RF["Roofline"]
  FL["Measure peak FLOPs"] --> RF
  RF --> RP["Ridge point"]
  HO["Host costs<br/>PCIe, launch, allocation"] --> PR["Predict tokens/s<br/>for a real model"]
  RP --> PR
  PR --> CMP["Compare with measured"]

  class BW,RP memory
  class FL compute
  class HO io
  class RF,PR,CMP neutral

2. Why it matters#

Datasheets give peak numbers under ideal conditions; your cloud instance, driver, and neighbours give something else. Engineers who carry their own measured constants make capacity predictions that hold; engineers who quote datasheets get surprised.

This is the numbers.md habit turned into a tool. Run it on every new instance type before you trust it.


3. Read first#


4. Spec#

Bench                What                                        Unit
bandwidth            large tensor copy / add (≥ 25% of VRAM)       GB/s
gemm_peak            square GEMM n=8192, per dtype fp32/fp16/bf16  TFLOP/s
gemm_sweep           n = 64 … 16384                                TFLOP/s vs n
gemv_decode          (B × d) @ (d × d), d=4096, B = 1 … 512        TFLOP/s, GB/s vs B
h2d / d2h            pageable and pinned, 1 MB … 1 GB              GB/s
launch_overhead      empty-ish kernel, 10k launches                µs / launch
alloc                torch.empty of varied sizes, cold and cached   µs
attention            scaled_dot_product_attention, T = 128 … 8192  ms, tokens/s
model_decode         real model, batch 1 … 64                      tok/s, ms/step

Output: report.json (+ env: GPU name, driver, CUDA, torch, power limit, temperature)
        roofline.png with every benchmark placed on it

Every benchmark: warm up, ≥20 repetitions, torch.cuda.synchronize() around the timed region, report median and p95.


5. Milestones#

  1. Timer harness with warmup, sync, and repetition. Test it on a sleep to make sure it reports what you think.
  2. Bandwidth and peak FLOPs. These two define the roofline. Compare both to the datasheet; explain any gap over 15%.
  3. Roofline plot. Compute ceiling, bandwidth slope, ridge point at peak_flops / bandwidth FLOPs per byte.
  4. Decode-shaped sweep. Batched GEMV at model-like width. Find the batch size at which throughput stops scaling linearly — the measured ridge. This answers Checkpoint C question 5 for your hardware.
  5. Host-side costs. PCIe, launch overhead, allocation.
  6. Prediction. For a real model: step_time ≈ weight_bytes / bandwidth at batch 1. Predict tokens/s, then measure. Report the ratio.
  7. Environment capture. Without it the report is not reproducible and therefore not useful.

6. Starter skeleton#

Python
def bench(fn, warmup=5, reps=20):
    for _ in range(warmup): fn()
    torch.cuda.synchronize()
    ts = []
    for _ in range(reps):
        t0 = time.perf_counter(); fn(); torch.cuda.synchronize()
        ts.append(time.perf_counter() - t0)
    return {"median": statistics.median(ts), "p95": sorted(ts)[int(0.95 * len(ts)) - 1]}

def bandwidth_gbs(nbytes=8 << 30):
    a = torch.empty(nbytes // 4, dtype=torch.float32, device="cuda").normal_()
    b = torch.empty_like(a)
    t = bench(lambda: b.copy_(a))["median"]
    return 2 * nbytes / t / 1e9                    # read a + write b

def gemm_tflops(n=8192, dtype=torch.float16):
    A = torch.randn(n, n, device="cuda", dtype=dtype); B = torch.randn_like(A)
    t = bench(lambda: A @ B)["median"]
    return 2 * n**3 / t / 1e12

def decode_sweep(d=4096, layers=8, dtype=torch.float16):
    Ws = [torch.randn(d, d, device="cuda", dtype=dtype) for _ in range(layers)]
    for B in (1, 2, 4, 8, 16, 32, 64, 128, 256, 512):
        x = torch.randn(B, d, device="cuda", dtype=dtype)
        def f():
            y = x
            for W in Ws: y = y @ W
        t = bench(f)["median"]
        yield B, B / t, layers * d * d * Ws[0].element_size() / t / 1e9   # rows/s, GB/s of weights

7. What to measure#

MeasurementExpectation to write down first
Bandwidth vs datasheet80-95% of rated
Peak FP16/BF16 TFLOP/s vs datasheetTensor cores; often well below the marketing number
FP32 vs FP16 GEMMLarge gap — tensor cores (VI.11)
GEMM TFLOP/s vs nTiny n is dominated by launch overhead
Decode sweep: rows/s vs BLinear, then bends at the ridge
Pinned vs pageable H2D2-3×
Launch overhead~5-15 µs; × kernels per step = host floor
Predicted vs measured tok/s at batch 1Within ~1.5× if the model is sound
nvidia-smi utilization during the batch-1 decode~100% while mostly idle on FLOPs (X.04)

8. Done when#

  • One command produces report.json and roofline.png.
  • Bandwidth and peak FLOPs are within a defensible distance of the datasheet, or you know why not (power cap, thermal throttling, shared host, old driver).
  • You located the decode ridge point by measurement.
  • Your batch-1 tokens/s prediction and measurement are both in the report.
  • You ran it on two different GPUs (or instance types) and compared.

9. Common pitfalls#

No synchronize. You measure launch, not execution.

Tensors too small for the bandwidth test. They fit in cache-like structures or are dominated by overhead. Go big.

Thermal or power throttling mid-run. Log nvidia-smi --query-gpu=temperature.gpu,power.draw,clocks.sm.

Including RNG or allocation in the timed region.

Reporting one run. Variance on shared cloud GPUs is real; report the spread.

Trusting nvidia-smi utilization as a measure of how hard the GPU is working. It is the fraction of time any kernel was running.


10. No GPU? Do this instead#

Run the same suite against the CPU (device="cpu") and against a Colab T4, and compare the two rooflines. The method is the deliverable; the device is incidental.


11. Stretch goals#

  • Multi-GPU: add nccl-tests-style all-reduce bandwidth and P2P copy between devices (IX.07).
  • Add Nsight Compute collection for one kernel and report dram__throughput directly.
  • Track results over time in CI to catch driver or image regressions.
  • Publish a tokens/s-per-dollar column using current on-demand prices (XI.03).

12. Interview questions this project answers#

  1. How do you measure a GPU’s memory bandwidth?
  2. Where is the ridge point for decode on your hardware, and what does it tell you about batch size?
  3. Given weight bytes and bandwidth, estimate tokens/s at batch 1.
  4. nvidia-smi says 100%. What do you measure next?

13. Next#

Project 12 — Inference gateway

↑↓ navigate↵ openesc close