1. What is it?#
Syscalls are the boundary between your program and the kernel. Profiling is finding out where time goes.
This file teaches the profiling method on the CPU side, where the tools are mature and the concepts are clear. Section X applies the same method to GPUs with Nsight.
2. Why does it exist?#
Because “it’s slow” is not a diagnosis, and guessing is expensive. The discipline is:
1. Measure end to end → how slow, and which metric?
2. Find the biggest component → profile, don't guess
3. Form a hypothesis → a mechanism, not a feeling
4. Test it cheaply → change one thing
5. Re-measure → did the number move?Steps 2 and 5 are where tools matter.
3. Simple analogy#
A doctor’s diagnostic ladder. Vital signs first (top, nvidia-smi), then targeted tests (perf, py-spy), then imaging (flame graphs, Nsight timelines), then biopsy (tracing individual syscalls or kernels). You do not start with the biopsy — it’s slow, invasive, and produces data you can’t interpret without the earlier steps.
4. Tiny example#
Two profilers, same program, different answers — and both are right:
// slow.go — one CPU-bound function, one I/O-bound function. Which is which?
package main
import (
"crypto/sha256"
"encoding/json"
"fmt"
"os"
"strconv"
)
func compute(n int) string {
h := sha256.New()
for i := 0; i < n; i++ {
h.Write([]byte(strconv.Itoa(i)))
}
return fmt.Sprintf("%x", h.Sum(nil))
}
func ioBound() {
f, _ := os.Create(os.TempDir() + "/x")
defer f.Close()
for i := 0; i < 20000; i++ {
line, _ := json.Marshal(map[string]int{"i": i})
f.Write(append(line, '\n')) // unbuffered: one write syscall per line
}
}
func main() {
for i := 0; i < 3; i++ {
compute(2_000_000)
ioBound()
}
}go build -o slow slow.go
# Sampling profiler: where is CPU time? (Linux; on macOS use Go's own pprof, below)
perf record -g ./slow && perf report
# Syscall profiler: what is it asking the kernel to do?
strace -c -f ./slowThe sampling profiler will point at compute. strace -c will show tens of thousands of write calls.
Different questions, different tools. If your problem is “too many syscalls” (unbuffered
I/O, per-token flush(), excessive futex from lock contention), a CPU profiler may show
nothing obviously wrong.
5. Technical explanation#
The syscalls that matter in inference#
ioctl CUDA driver communication — every kernel launch, alloc, sync
futex lock/condvar waits — high counts = contention
epoll_wait the async event loop — normal for the API server
read/write network and file I/O
mmap/munmap allocator activity — churn here means allocation thrash
clone thread/process creation — should be near zero in steady state
sched_yield usually a spin-wait; suspicious in large numbersstrace -c -f -p <pid> # counts and time per syscall
strace -f -p <pid> -e trace=futex -T # who is waiting, and how longA steady-state inference server should show mostly ioctl, epoll_wait, read/write, and
futex. Large mmap/munmap counts mean allocator churn; large clone counts mean you’re
creating threads per request.
Sampling vs instrumenting#
| Sampling | Instrumenting | |
|---|---|---|
| How | interrupt periodically, record stack | wrap every function |
| Overhead | 1-5% | 10-1000% |
| Bias | misses very short functions | changes what you’re measuring |
| Use for | production, “where does time go?” | development, exact call counts |
| Tools | perf record, py-spy, async-profiler | cProfile, line_profiler |
Use sampling in production. Always. cProfile on an inference server will change the answer
by slowing Python 2-10x, which is exactly the component you’re measuring.
perf#
# Where are cycles spent, system-wide or for a process?
perf record -F 99 -g -p <pid> -- sleep 30
perf report --stdio | head -40
# Hardware counters — the fastest way to classify a workload
perf stat -e cycles,instructions,cache-misses,cache-references,\
branch-misses,dTLB-load-misses,LLC-load-misses -p <pid> -- sleep 10
# Off-CPU: where is it BLOCKED? (often the real answer)
perf record -e sched:sched_switch -g -p <pid> -- sleep 10Interpreting perf stat:
- IPC < 1 → stalling (memory, branches, or interpreter overhead)
- cache-miss rate > 20% of references → memory-bound
- branch-miss rate > 5% → branchy code (tokenizer, interpreter)
Flame graphs#
perf record -F 99 -g -p <pid> -- sleep 30
perf script | stackcollapse-perf.pl | flamegraph.pl > cpu.svg
# Python (much easier)
py-spy record -o py.svg --pid <pid> --duration 30
py-spy record -o py.svg --pid <pid> --duration 30 --idle # include blocked threadsRead a flame graph: width = time, height = stack depth. Look for wide plateaus. Ignore height — a deep stack that’s narrow costs nothing.
The --idle flag matters for inference servers: much of the interesting time is spent waiting
(on the GPU, on a lock, on the network), and the default excludes it.
eBPF / bpftrace#
For questions no existing tool answers:
# Distribution of read() latencies
bpftrace -e 'tracepoint:syscalls:sys_enter_read /pid == PID/ { @start[tid] = nsecs; }
tracepoint:syscalls:sys_exit_read /@start[tid]/ {
@us = hist((nsecs - @start[tid])/1000); delete(@start[tid]); }'
# Off-CPU time by stack
bpftrace -e 'kprobe:finish_task_switch { @[kstack] = count(); }'
# Which files is a process opening?
bpftrace -e 'tracepoint:syscalls:sys_enter_openat { printf("%s %s\n", comm, str(args->filename)); }'eBPF runs safely in the kernel with ~1-2% overhead and no restarts. It is the right tool when you need a custom measurement in production.
The USE method#
For every resource, check three things:
Utilization — % of time busy
Saturation — queued work waiting
Errors — error countsApplied to an inference host:
| Resource | Utilization | Saturation | Errors |
|---|---|---|---|
| CPU | top, vmstat us+sy | run queue r, nr_throttled | — |
| Memory | free, RSS | swapping, memory.events | OOM kills |
| GPU compute | nvidia-smi util | queued kernels, launch gaps | Xid errors in dmesg |
| GPU memory | memory.used | preemption/swapping in engine | CUDA OOM |
| Disk | iostat %util | await, queue depth | I/O errors |
| Network | sar -n DEV | retransmits, backlog drops | interface errors |
| KV cache | occupancy % | queue depth, preemptions | — |
That last row is inference-specific and belongs on your dashboard (Section XI.08).
6. Under the hood#
Sampling profilers work by setting a hardware performance counter (or a timer) to fire an
interrupt every N events; the handler walks the stack. Stack walking needs either frame pointers
(-fno-omit-frame-pointer) or DWARF unwind info. Python builds and many distro packages omit
frame pointers, which is why perf sometimes shows you nothing but [unknown] — use
--call-graph dwarf or use py-spy, which understands CPython’s own frame objects directly.
7. Performance implications of profiling itself#
perf record -F 99— ~1% overhead, safe in production.py-spy— ~1-2%, safe, no restart, no instrumentation.strace— 10-100x slowdown. Never leave it attached in production; use-cbriefly or use eBPF instead.cProfile— 2-10x on Python. Development only.- Nsight Systems — 5-20% depending on trace scope.
8. Production implications#
- Ship
py-spyandperfin your production image. The time to install tools is the time the incident is ongoing. - Continuous profiling (Parca, Pyroscope, Grafana Phlare) gives you the flame graph from before the incident, which is usually what you actually need.
- Instrument phase timings in the app: queue wait, prefill, decode, detokenize, network. Application-level timings answer 80% of questions without any profiler.
- Record a baseline profile for every release. Comparing two flame graphs is far more informative than reading one.
9. Common mistakes#
Guessing instead of measuring. The most expensive mistake. Intuition about performance is reliably wrong, including yours and including mine.
Profiling the wrong thing. If the GPU is the bottleneck, a CPU flame graph shows an idle event loop and tells you nothing. Determine the bottleneck layer first.
Ignoring off-CPU time. For inference servers, most wall-clock time is waiting. Use
py-spy --idle and off-CPU profiling.
Using cProfile on a server. Distorts what you’re measuring.
Profiling without a warm-up. The first iterations include compilation, allocation, and cache warming.
Optimizing what the profiler shows without checking it’s on the critical path. A wide plateau in a background thread doesn’t affect latency.
10. Hands-on exercise#
A. Both profilers. Run the slow.go example. Produce a CPU flame graph (perf, or add
runtime/pprof and use go tool pprof -http=:0 cpu.prof) and a strace -c summary. Write down what each tells you that the other doesn’t.
B. Classify with perf stat. Run perf stat on: a pure Python loop, a NumPy matmul, a
random-access gather, and your model server. Build a table of IPC, cache-miss rate, and
branch-miss rate. Explain each row.
C. Off-CPU profile. Profile a model server under load with py-spy record --idle. What
fraction of time is spent waiting? On what?
D. Write a bpftrace one-liner. Measure the distribution of ioctl durations for your
inference process (a proxy for CUDA driver call latency). What is the p99?
E. Build the USE checklist. Turn the table in section 5 into a script for your environment.
11. Interview questions#
- Describe your methodology for diagnosing a slow inference service.
- Sampling vs instrumenting profilers — when do you use each?
- What does a low IPC tell you? What about a high branch-miss rate?
- Why is off-CPU profiling important for inference servers?
- What is the USE method and how would you apply it to a GPU?
- How would you profile a production server without restarting it?
- Your flame graph shows 40% of time in a function, but removing it doesn’t help. Why might that be?
12. Further reading#
- [FUNDAMENTAL] Brendan Gregg, Systems Performance (2nd ed.) and BPF Performance Tools
- [REFERENCE]
perfwiki; FlameGraph repository (brendangregg/FlameGraph) - [REFERENCE] py-spy, bpftrace documentation
- Next: Section III — ML Fundamentals