Sections I.08 and VI.04 introduced the roofline. This file is about using it on a real system: how to place your kernels on it, and what each position tells you to do.
1. The practical roofline#
achieved
TFLOP/s
▲
│ ┌────────────────────── peak compute
│ ╱│
│ ╱ │
│ ●prefill ╱ │ ← you want kernels HERE (near the roof)
│ ╱ │
│ ╱ │
│ ╱ │
│ ● decode ╱ │ ← a kernel BELOW the sloped line
│ (b=64) ╱ │ has headroom
│ ╱
│ ●decode(b=1) ╱ slope = memory bandwidth
│ ●rmsnorm ╱
└──────────────────────────────────────► arithmetic intensity
1 10 100 1000Your job: for each significant kernel, find its (intensity, achieved) point and its distance from the roof.
2. Getting the numbers#
ncu --metrics \
sm__sass_thread_inst_executed_op_ffma_pred_on.sum,\
sm__sass_thread_inst_executed_op_fadd_pred_on.sum,\
sm__sass_thread_inst_executed_op_fmul_pred_on.sum,\
sm__inst_executed_pipe_tensor.sum,\
dram__bytes_read.sum,\
dram__bytes_write.sum,\
gpu__time_duration.sum \
-k regex:"your_kernel" ./appFLOPs = 2×ffma + fadd + fmul (for non-tensor-core kernels)
= tensor_inst × FLOPs_per_MMA (for tensor-core kernels;
e.g. m16n8k16 = 4096 FLOPs)
Bytes = dram_bytes_read + dram_bytes_write
Intensity = FLOPs / Bytes
Achieved = FLOPs / durationOr let Nsight do it:
ncu --set roofline -k regex:"your_kernel" -o report ./app
ncu-ui report.ncu-rep # the Roofline section has the chartThe GUI roofline chart is the fastest path and shows both the HBM and the L2/shared-memory rooflines (the hierarchical roofline, Section VI.04).
3. Reading the position#
POSITION DIAGNOSIS ACTION
On the sloped line, low I memory bound, optimal reduce BYTES:
quantize, fuse,
better algorithm
(kernel tuning won't help)
Below the sloped line, low I memory bound, SUBOPTIMAL fix ACCESS:
coalescing, layout,
vectorized loads,
more parallelism
On the flat roof, high I compute bound, optimal reduce FLOPs, lower
precision, or a
different algorithm
Below the flat roof, high I compute bound, SUBOPTIMAL use TENSOR CORES,
better tiling,
reduce divergence
Far below both, middling I LATENCY bound more occupancy,
more ILP,
reduce dependenciesThe middle column is the important distinction: “optimal” versus “suboptimal” at the same regime. A kernel at 90% of the memory roofline is done — go reduce bytes elsewhere. A kernel at 30% of it has a fixable problem.
4. Worked example — placing an LLM’s kernels#
Llama-3-8B decode, batch 32, H100 (peak 990 TF/s BF16, 3.35 TB/s):
Ridge point = 296 FLOP/byte
Kernel FLOPs Bytes I Achieved % of roof
qkv_proj GEMM 1.2 G 0.61 G 2.0 6.5 TF/s 97% of memory roof
o_proj GEMM 1.1 G 0.55 G 2.0 6.4 TF/s 96%
gate_up GEMM 7.5 G 3.8 G 2.0 6.4 TF/s 96%
down GEMM 3.8 G 1.9 G 2.0 6.5 TF/s 97%
paged_attention 0.5 G 0.55 G 0.9 2.9 TF/s 96%
add_rms_norm (fused) 0.02 G 0.016 G 1.25 1.4 TF/s 33% ← LOOK
silu_and_mul (fused) 0.03 G 0.05 G 0.6 2.0 TF/s 99%
lm_head GEMM 1.05 G 0.53 G 2.0 6.4 TF/s 96%
sampling 0.01 G 0.03 G 0.3 0.6 TF/s 60% ← LOOK
Total step: 15.2 GFLOP, 8.0 GB, 2.4 msTwo kernels stand out. add_rms_norm at 33% of the memory roofline has a fixable problem —
likely poor vectorization or low occupancy for a small tensor. sampling at 60% likewise.
Everything else is at 96-99% of the memory roofline: those kernels are done. No amount of tuning improves them. The only lever is reducing bytes — which means quantization.
That table is the output of a proper performance investigation, and it tells you exactly where to spend effort: fix the two outliers (small gain), then quantize (large gain).
5. The hierarchical roofline#
A kernel can appear to exceed the HBM roofline. It isn’t cheating — it’s getting cache hits.
TFLOP/s
▲ ┌───────────────── peak compute
│ ╱╱│
│ ╱╱╱ │ ← shared memory roofline (19 TB/s)
│ ╱╱╱ │ ← L2 roofline (7 TB/s)
│ ╱╱╱ │ ← HBM roofline (3.35 TB/s)
│╱╱╱
└────────────────────────► intensity (computed against DRAM bytes)Kernel with high L2/shared reuse:
DRAM intensity: 50 FLOP/byte → appears memory-bound
L2 intensity: 8 FLOP/byte → but it's hitting L2, not DRAM
→ the relevant roofline is L2's, and it may be near itCheck both: dram__bytes.sum and lts__t_bytes.sum (L2) and l1tex__t_bytes.sum.
The ratio l1tex_bytes / dram_bytes is your reuse factor.
FlashAttention is the canonical example: it moved attention from the HBM roofline to the shared-memory roofline by tiling.
6. What the roofline doesn’t tell you#
Be honest about its limits:
✗ Launch overhead. A kernel can be at 99% of the roofline and still be
irrelevant if 60% of your step time is gaps between kernels.
✗ Whether the algorithm is right. A kernel at the roof might be computing
something you don't need.
✗ Tail effects. Wave and tile quantization show as reduced achieved
performance without explaining why.
✗ Multi-kernel interactions. Cache pollution between kernels.
✗ Anything CPU-side.The roofline is a per-kernel tool. It comes at step 4-5 of the methodology (file 01), after you’ve established that the GPU is the bottleneck and which kernels matter.
7. Using it for capacity and procurement#
The roofline is also a planning tool:
"We're at 94% of the HBM roofline for decode."
→ software optimization is exhausted for this configuration
→ the remaining levers are: fewer bytes (quantization, GQA, KV quant)
or more bandwidth (different hardware)
→ this is a defensible procurement argument
"We're at 30% of the roofline."
→ don't buy hardware yet; there's 3x available in software“What fraction of the roofline are we achieving?” is the single best summary number for an inference system’s efficiency, and it belongs in your regular reporting.
8. Production implications#
- Establish the roofline position of your top 5 kernels once, and record it.
- Track achieved-vs-roofline as a regression metric. A drop from 90% to 55% means something changed: a kernel fell back, a layout broke, a library updated.
- Use it to reject optimizations. “This reduces FLOPs 40%” is worth zero for a kernel at the memory roof.
- Use it to justify hardware. It’s the quantitative version of “we’ve optimized as far as we can.”
- Recompute it when the workload changes. Batch size and context length move kernels along the intensity axis.
9. Common mistakes#
Using peak FLOPs that include sparsity. Halves your apparent efficiency.
Ignoring the hierarchical roofline. A cache-resident kernel looks impossible.
Computing intensity from the algorithm, not the implementation. Unfused chains have much lower effective intensity.
Forgetting writes. Bytes = read + write.
Applying it to the whole model. It’s per-kernel; a model contains kernels in both regimes.
Using it too early. It’s step 4-5, not step 1.
10. Hands-on exercise#
A. Build the table. For a real model’s decode step, produce the table in section 4: every significant kernel with its FLOPs, bytes, intensity, achieved TFLOP/s, and % of the appropriate roofline. This is the single most informative artifact you can produce about an inference system.
B. Find the outliers. In your table, which kernels are furthest below their roofline? Investigate one: is it coalescing, occupancy, or something else?
C. Hierarchical. For a tiled kernel (a GEMM or FlashAttention), measure both DRAM and L2/L1 bytes. Compute both intensities. Which roofline is it actually near?
D. Track the change. Measure the roofline position at batch 1, 8, 64, 256. Plot how the kernels move along the intensity axis. At what batch does the decode GEMM approach the ridge point?
E. Predict a quantization gain. Using the roofline, predict the decode speedup from FP8. Then measure it. How close was your prediction?
11. Interview questions#
- How do you compute a kernel’s position on the roofline from profiler metrics?
- A kernel is at 95% of the memory roofline. What optimizations remain?
- A kernel is at 30% of the memory roofline. What do you investigate?
- What is the hierarchical roofline and when does it matter?
- What does the roofline not tell you?
- How would you use the roofline in a hardware procurement argument?
- Where does roofline analysis fit in your debugging methodology?
12. Further reading#
- [FUNDAMENTAL] Williams, Waterman, Patterson, “Roofline” (CACM 2009)
- [REFERENCE] Nsight Compute roofline documentation
- [ESTABLISHED] Yang et al., “Hierarchical Roofline Analysis for GPUs” (NERSC)
- Next: 03 — Memory-bound vs compute-bound