1. What is it?#
A kernel is a function that runs on the GPU. A launch is the CPU telling the GPU to run one. Launches have a fixed cost — roughly 3-10 µs of CPU time and ~1-3 µs of GPU-side scheduling — regardless of how much work the kernel does.
An LLM decode step issues 200-2,000 launches. At batch 1, that overhead can exceed the actual computation.
2. Why does it exist as a problem?#
Because the operation granularity of a neural network framework (one op = one kernel) is much finer than the granularity at which launch overhead is amortized.
Prefill, 2048 tokens: each kernel does ~40 ms of work. 5 µs launch = 0.01%. Irrelevant.
Decode, batch 1: each kernel does ~10 µs of work. 5 µs launch = 50%. Fatal.Same model, same kernels, completely different significance.
3. Simple analogy#
Placing an order with a supplier. There’s a fixed 5 minutes of paperwork per order plus the actual delivery time. Ordering a truckload: paperwork is negligible. Ordering one screw at a time, 400 times: you spend 33 hours on paperwork and 20 minutes receiving screws.
CUDA graphs are the standing order: fill in the paperwork once for the whole sequence, then trigger it repeatedly.
4. Tiny example#
import torch, time
x = torch.randn(16, 16, device='cuda')
# Measure launch overhead with a trivially small kernel
torch.cuda.synchronize(); t0 = time.perf_counter()
for _ in range(10000): y = x + 1
torch.cuda.synchronize()
per_launch = (time.perf_counter()-t0)/10000
print(f"per-op time: {per_launch*1e6:.2f} us") # typically 5-15 usThat number is (Python dispatch + CUDA launch), essentially none of it useful arithmetic (adding 1 to 256 floats takes nanoseconds).
Now the fix:
# CUDA graph capture
static_x = torch.randn(16, 16, device='cuda')
g = torch.cuda.CUDAGraph()
# warmup on a side stream is required before capture
s = torch.cuda.Stream()
s.wait_stream(torch.cuda.current_stream())
with torch.cuda.stream(s):
for _ in range(3): y = static_x + 1
torch.cuda.current_stream().wait_stream(s)
with torch.cuda.graph(g):
static_y = static_x + 1
torch.cuda.synchronize(); t0 = time.perf_counter()
for _ in range(10000): g.replay()
torch.cuda.synchronize()
print(f"graph replay: {(time.perf_counter()-t0)/10000*1e6:.2f} us") # typically 2-4 usFor a single op the win is modest. For a 350-kernel decode step captured as one graph, the win is dramatic — one replay instead of 350 launches.
5. Technical explanation#
The launch path#
Python: y = a @ b
↓ ~2-5 µs
PyTorch dispatcher → backend → cuBLAS
↓
cudaLaunchKernel:
- validate arguments
- write a command packet into a pinned ring buffer
- ring the doorbell (a write to a mapped GPU register)
↓ ~1-3 µs GPU-side
GPU front-end reads the packet, allocates blocks to SMs
↓
kernel executesNote the CPU work happens asynchronously with respect to GPU execution — the CPU returns immediately. So launch overhead only hurts when the CPU cannot stay ahead of the GPU. At batch 1 with fast kernels, it can’t.
GPU: |kern|gap|kern|gap|kern|gap| ← launch-bound: gaps between kernels
CPU: |--launch--|--launch--|... ← CPU is the critical path
GPU: |kern|kern|kern|kern| ← compute-bound: CPU stays ahead
CPU: |l|l|l|l| idleDiagnosing this is easy: look at an Nsight Systems timeline. Gaps between kernels with a busy CPU row = launch-bound.
CUDA graphs#
A CUDA graph captures a sequence of operations and their dependencies as a DAG, then replays it with a single launch.
Capture: record all kernels, memcpies, and their dependencies
Instantiate: the driver builds an optimized executable graph
Replay: one API call launches the entire DAGBenefits:
- One launch instead of N. CPU overhead drops ~10-50x.
- The driver can pre-resolve dependencies and pre-allocate.
- GPU-side scheduling gaps shrink.
Constraints — and these are the reason it isn’t universal:
- Shapes must be static. Every tensor’s shape is baked in at capture.
- Memory addresses must be static. Input/output buffers are fixed; you copy into them.
- No CPU-dependent control flow inside the captured region.
- No dynamic allocation during replay.
For LLM decode this is workable because decode shapes are almost static: (batch, 1, d). The
batch size varies, so engines capture graphs for a set of batch sizes (1, 2, 4, 8, 16, …) and
pad up to the nearest. That padding is a real cost — running batch 9 as batch 16 wastes 44%
of that step — so engines capture many sizes, trading capture time and memory for less padding.
vLLM does exactly this: --enforce-eager disables it, and you can see the capture happening at
startup (it takes 20-60 seconds and a GB or two of memory).
Why graphs can’t cover everything#
Prefill: variable sequence lengths → different shapes every time → no graph
(some engines bucket prefill lengths and capture graphs per bucket)
Decode: nearly static → graphs work well
Sampling: variable (different requests have different sampling params) → often outside
Scheduling: CPU logic → outsideTypical result: decode is graph-captured, prefill and everything else is eager.
6. Under the hood#
# Count kernel launches
nsys profile --stats=true python your_script.py
# look at "CUDA API Summary": cudaLaunchKernel count and total time
# In PyTorch
from torch.profiler import profile, ProfilerActivity
with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA]) as prof:
model(x)
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=20))
# compare "Self CPU time total" to "Self CUDA time total"If CPU total ≈ or > CUDA total, you are launch/dispatch bound.
7. Performance implications#
Typical measured gains from CUDA graphs on LLM decode:
Batch 1: 25-45% faster
Batch 8: 15-30%
Batch 32: 8-15%
Batch 128: 3-8%
Prefill: ~0% (not captured; kernels are large anyway)The benefit shrinks as batch grows because the GPU work per kernel grows while launch cost stays fixed. CUDA graphs are a small-batch optimization — which is exactly the regime of latency-sensitive serving.
Other launch-overhead reducers:
- Fusion (file 09): fewer kernels to launch at all.
- Larger batch: more work per launch.
- Moving the loop out of Python: a C++ or Rust engine loop reduces dispatch cost even without graphs.
8. Production implications#
- Enable CUDA graphs in production. In vLLM they’re on by default;
--enforce-eagerturns them off (useful for debugging, costly in production). - Budget for capture. 20-60 seconds of startup and 1-3 GB of memory for the captured graphs and their static buffers.
- Understand the padding cost. Check which batch sizes your engine captures. If it captures {1,2,4,8,16,32,…} and your typical batch is 20, you’re padding to 32 — 60% waste on that dimension. Some engines let you configure the capture sizes.
- Graphs and dynamic features conflict. LoRA switching, variable sampling parameters, and speculative decoding all complicate capture. Check that your feature set is compatible.
9. Common mistakes#
Ignoring launch overhead at small batch. It’s often 30-50% of your ITL.
Disabling CUDA graphs and forgetting. --enforce-eager left on in production.
Capturing too few batch sizes. Excessive padding.
Capturing too many. Long startup, high memory.
Assuming graphs help prefill. They generally don’t.
Timing a graph without warmup. The first replay includes instantiation.
10. Hands-on exercise#
A. Measure your launch overhead. Run the microbenchmark in section 4. Record µs per launch
in numbers.md.
B. Count kernels per token. Profile a decode step of a real model. How many kernels? Multiply by your launch overhead. What fraction of ITL is that?
C. Graph a model. Take a small transformer and capture its decode step as a CUDA graph. Measure before and after at batch 1, 8, 32. Plot the speedup vs batch size. Does it match the table in section 7?
D. Padding cost. With vLLM (or your engine), determine which batch sizes are graph-captured. Compute the average padding waste for a realistic batch-size distribution.
E. Find the gaps. Run nsys on a batch-1 decode with and without graphs. Look at the
timeline. Measure the total gap time in each.
11. Interview questions#
- What is kernel launch overhead and when does it matter?
- Why does it matter at batch 1 but not at batch 256?
- What are CUDA graphs and what constraints do they impose?
- Why can’t prefill usually be graph-captured?
- How do engines handle variable batch sizes with CUDA graphs, and what does it cost?
- How would you determine whether a workload is launch-bound?
12. Further reading#
- [REFERENCE] CUDA Programming Guide, “CUDA Graphs”
- [REFERENCE] PyTorch
torch.cuda.CUDAGraphandmake_graphed_callablesdocs - [REFERENCE] vLLM’s
cuda_graphcapture implementation - Next: 09 — Operator fusion