PidokuInfra

Streams and Synchronization

Intermediate 1h Difficulty 3/5 Topic 06 of 12

Prerequisites 05


1. What is it?#

A stream is an ordered queue of GPU operations. Operations in the same stream execute in order; operations in different streams may overlap.

Stream 0:  [kernel A]──[kernel B]──[kernel C]        in order
Stream 1:      [memcpy H2D]──[kernel D]              in order, may overlap with stream 0
Stream 2:              [memcpy D2H]                   may overlap with both

Streams are how you get concurrency between independent work: overlapping data transfer with compute, running a small kernel alongside a large one, or serving two models on one GPU.


2. Why does it exist?#

Because a single stream serializes everything, and the GPU has independent hardware units that could be working simultaneously:

Hardware units on an H100:
  - SMs (compute)
  - Copy engines (H2D and D2H, separate)
  - NVLink/network engines

With one stream: copy, then compute, then copy. Serial.
With three streams: copy(i+1) while compute(i) while copy-back(i-1). Pipelined.

3. Simple analogy#

A kitchen with one cook versus a kitchen with lanes.

One stream: the cook does everything in strict order — fetch ingredients, chop, cook, plate, serve, then fetch the next order’s ingredients. The oven sits idle while fetching.

Multiple streams: while dish A is in the oven, fetch ingredients for dish B and plate dish C. The equipment is used concurrently.

The synchronization primitives are the rules: “don’t plate dish A until it’s out of the oven.”


4. Tiny example#

Overlapping transfer and compute:

Python
import torch, time

N = 100_000_000
h = torch.empty(N, dtype=torch.float32, pin_memory=True)   # pinned!
d = torch.empty(N, dtype=torch.float32, device='cuda')
W = torch.randn(4096, 4096, device='cuda', dtype=torch.float16)
x = torch.randn(4096, 4096, device='cuda', dtype=torch.float16)

# --- SERIAL: same stream ---
torch.cuda.synchronize(); t0 = time.perf_counter()
d.copy_(h)
y = x @ W
torch.cuda.synchronize()
t_serial = time.perf_counter()-t0

# --- OVERLAPPED: separate streams ---
s_copy = torch.cuda.Stream()
torch.cuda.synchronize(); t0 = time.perf_counter()
with torch.cuda.stream(s_copy):
    d.copy_(h, non_blocking=True)
y = x @ W                                 # runs on the default stream, concurrently
torch.cuda.synchronize()
t_overlap = time.perf_counter()-t0

print(f"serial {t_serial*1e3:.1f} ms, overlapped {t_overlap*1e3:.1f} ms, "
      f"saved {(1-t_overlap/t_serial)*100:.0f}%")

Verify the overlap actually happened in Nsight Systems — the copy and the kernel should appear on different rows, temporally overlapping.


5. Technical explanation#

The default stream, and why it’s tricky#

Legacy default stream (stream 0):
  Implicitly synchronizes with all other blocking streams.
  Any operation on the default stream waits for all other streams,
  and other streams wait for it.
  → accidentally serializes everything

Per-thread default stream (compile with --default-stream per-thread):
  Each host thread gets its own default stream. No implicit sync.

Non-blocking streams (cudaStreamNonBlocking):
  Do not synchronize with the legacy default stream.

PyTorch uses a per-device default stream and creates non-blocking streams via torch.cuda.Stream(), so the classic footgun is mostly avoided — but if you mix raw CUDA code in, be careful.

Events for cross-stream dependencies#

Python
s1, s2 = torch.cuda.Stream(), torch.cuda.Stream()
ev = torch.cuda.Event()

with torch.cuda.stream(s1):
    a = produce()
    ev.record(s1)                 # mark: "a is ready"

with torch.cuda.stream(s2):
    ev.wait(s2)                   # s2 waits for the event, but the CPU does NOT block
    b = consume(a)

ev.wait(stream) inserts a GPU-side dependency. The CPU keeps running. This is how you build pipelines without CPU involvement.

Contrast with ev.synchronize(), which blocks the CPU until the event fires. Use that only when you actually need the CPU to wait.

Where streams help in inference#

1. WEIGHT LOADING at startup
   Overlap disk read, H2D copy, and layout conversion across streams.
   Cuts cold start meaningfully (Section II.07).

2. KV CACHE OFFLOAD / PREFETCH
   Copy KV blocks between GPU and CPU on a side stream while compute proceeds.
   Essential for CPU-offload designs (Section XIII.07).

3. MULTI-MODEL SERVING on one GPU
   Each model on its own stream. Works, but they compete for SMs;
   MPS or MIG gives better isolation.

4. SPECULATIVE DECODING
   Draft model on one stream, target on another — limited benefit since
   they're dependent, but the draft's small kernels can fill gaps.

5. DISAGGREGATED KV TRANSFER
   Sending KV to a decode worker while continuing prefill (Section XIII.06).

6. PREFILL/DECODE OVERLAP
   Some engines run prefill and decode on separate streams. Contended,
   but can improve utilization.

Priorities#

Python
high = torch.cuda.Stream(priority=-1)   # lower number = higher priority
low  = torch.cuda.Stream(priority=0)

Priority affects which stream’s blocks get scheduled first when SM slots free up. It does not preempt running blocks. So a long-running kernel on the low-priority stream still delays the high-priority one — priorities help at block granularity, not instruction granularity.

Practical use: put latency-critical decode on a high-priority stream and background work (prefix cache warming, metrics computation) on a low-priority one.

MPS and time-slicing#

Default:  multiple processes on one GPU are TIME-SLICED by the driver.
          Only one process's kernels run at a time. Context switches cost ~10-100 µs.

MPS (Multi-Process Service):
          A daemon merges multiple processes' work into one context, so their
          kernels can run CONCURRENTLY on different SMs.
          nvidia-cuda-mps-control -d

MIG:      Hardware partitioning. Strongest isolation, fixed sizes.

For serving several small models on one GPU, MPS is often the right answer: better utilization than time-slicing, more flexible than MIG. Downside: a fault in one process can affect the MPS server.


6. Under the hood#

Whether two kernels actually overlap depends on resources:

Kernel A uses 100% of SMs → kernel B waits regardless of streams
Kernel A uses 20% of SMs  → kernel B can fill the remaining 80%

So streams enable overlap; they don’t guarantee it. Small kernels overlap well; large ones don’t. A decode step’s small kernels can overlap with a prefill’s large ones only to the extent the prefill leaves SMs free — which it usually doesn’t.

Verify overlap in Nsight Systems: kernels on different rows should visibly overlap in time. If they’re serialized despite being on different streams, look for an implicit synchronization or resource contention.


7. Performance implications#

Use case                          Typical gain from streams
Weight loading overlap            20-40% faster cold start
KV offload prefetch               makes offload viable at all
Multi-model on one GPU (MPS)      1.5-3x utilization vs time-slicing
Prefill/decode overlap            5-15% (contended)
H2D/compute overlap in a pipeline 30-50% for transfer-heavy workloads

8. Production implications#

  • Use pinned memory pools for anything you stream.
  • Verify overlap in a profile. Assuming it happened is a common error.
  • Consider MPS for multi-model, small-model serving. Measure it against time-slicing.
  • Beware over-synchronizing. A single torch.cuda.synchronize() in a hot path serializes everything you carefully pipelined.
  • Stream priorities are a weak tool. Don’t rely on them for latency isolation; use separate GPUs or MIG for hard guarantees.

9. Common mistakes#

Assuming different streams guarantee overlap. Resources must be available.

Using the legacy default stream in code that mixes with other streams.

Forgetting pinned memory. No async transfer without it.

Over-synchronizing. torch.cuda.synchronize() in a loop.

Expecting stream priority to preempt. It doesn’t.

Not verifying with a profiler.


10. Hands-on exercise#

A. Prove overlap. Run the example in section 4. Verify in Nsight Systems that the copy and the kernel overlap. Then remove pin_memory=True and observe that they don’t.

B. Pipeline. Build a 3-stage pipeline (H2D → compute → D2H) with 3 streams and double buffering, processing 20 chunks. Compare total time to the serial version. Compute the theoretical speedup and compare.

C. Overlap limits. Launch a kernel that uses all SMs and a small kernel on another stream. Measure whether they overlap. Then reduce the first kernel’s grid size until overlap occurs.

D. MPS. If you have a spare GPU, run two model server processes with and without MPS. Compare aggregate throughput.

E. Find the sync. Take a piece of ML code, profile it, and find every synchronization point. Remove the unnecessary ones and measure.


11. Interview questions#

  1. What is a CUDA stream and what ordering does it guarantee?
  2. Why don’t kernels on different streams always overlap?
  3. What is an event and how does it differ from synchronize()?
  4. When is pinned memory required?
  5. What is MPS and when would you use it for inference?
  6. Give three places in an inference stack where streams matter.
  7. What does stream priority do and what doesn’t it do?

12. Further reading#

  • [REFERENCE] CUDA C++ Programming Guide, “Streams and Events”
  • [REFERENCE] NVIDIA MPS documentation
  • [REFERENCE] Mark Harris, “How to Overlap Data Transfers in CUDA C/C++”
  • Next: 07 — CUDA graphs

↑↓ navigate↵ openesc close