PidokuInfra

Syscalls and Profiling Basics

Foundations Intermediate 1h 15m Difficulty 3/5 Topic 13 of 13

Prerequisites 05, 10, 11


1. What is it?#

Syscalls are the boundary between your program and the kernel. Profiling is finding out where time goes.

This file teaches the profiling method on the CPU side, where the tools are mature and the concepts are clear. Section X applies the same method to GPUs with Nsight.


2. Why does it exist?#

Because “it’s slow” is not a diagnosis, and guessing is expensive. The discipline is:

1. Measure end to end          →  how slow, and which metric?
2. Find the biggest component  →  profile, don't guess
3. Form a hypothesis           →  a mechanism, not a feeling
4. Test it cheaply             →  change one thing
5. Re-measure                  →  did the number move?

Steps 2 and 5 are where tools matter.


3. Simple analogy#

A doctor’s diagnostic ladder. Vital signs first (top, nvidia-smi), then targeted tests (perf, py-spy), then imaging (flame graphs, Nsight timelines), then biopsy (tracing individual syscalls or kernels). You do not start with the biopsy — it’s slow, invasive, and produces data you can’t interpret without the earlier steps.


4. Tiny example#

Two profilers, same program, different answers — and both are right:

Go
// slow.go — one CPU-bound function, one I/O-bound function. Which is which?
package main

import (
	"crypto/sha256"
	"encoding/json"
	"fmt"
	"os"
	"strconv"
)

func compute(n int) string {
	h := sha256.New()
	for i := 0; i < n; i++ {
		h.Write([]byte(strconv.Itoa(i)))
	}
	return fmt.Sprintf("%x", h.Sum(nil))
}

func ioBound() {
	f, _ := os.Create(os.TempDir() + "/x")
	defer f.Close()
	for i := 0; i < 20000; i++ {
		line, _ := json.Marshal(map[string]int{"i": i})
		f.Write(append(line, '\n')) // unbuffered: one write syscall per line
	}
}

func main() {
	for i := 0; i < 3; i++ {
		compute(2_000_000)
		ioBound()
	}
}
Shell
go build -o slow slow.go

# Sampling profiler: where is CPU time?   (Linux; on macOS use Go's own pprof, below)
perf record -g ./slow && perf report

# Syscall profiler: what is it asking the kernel to do?
strace -c -f ./slow

The sampling profiler will point at compute. strace -c will show tens of thousands of write calls. Different questions, different tools. If your problem is “too many syscalls” (unbuffered I/O, per-token flush(), excessive futex from lock contention), a CPU profiler may show nothing obviously wrong.


5. Technical explanation#

The syscalls that matter in inference#

ioctl        CUDA driver communication — every kernel launch, alloc, sync
futex        lock/condvar waits — high counts = contention
epoll_wait   the async event loop — normal for the API server
read/write   network and file I/O
mmap/munmap  allocator activity — churn here means allocation thrash
clone        thread/process creation — should be near zero in steady state
sched_yield  usually a spin-wait; suspicious in large numbers
Shell
strace -c -f -p <pid>          # counts and time per syscall
strace -f -p <pid> -e trace=futex -T   # who is waiting, and how long

A steady-state inference server should show mostly ioctl, epoll_wait, read/write, and futex. Large mmap/munmap counts mean allocator churn; large clone counts mean you’re creating threads per request.

Sampling vs instrumenting#

SamplingInstrumenting
Howinterrupt periodically, record stackwrap every function
Overhead1-5%10-1000%
Biasmisses very short functionschanges what you’re measuring
Use forproduction, “where does time go?”development, exact call counts
Toolsperf record, py-spy, async-profilercProfile, line_profiler

Use sampling in production. Always. cProfile on an inference server will change the answer by slowing Python 2-10x, which is exactly the component you’re measuring.

perf#

Shell
# Where are cycles spent, system-wide or for a process?
perf record -F 99 -g -p <pid> -- sleep 30
perf report --stdio | head -40

# Hardware counters — the fastest way to classify a workload
perf stat -e cycles,instructions,cache-misses,cache-references,\
branch-misses,dTLB-load-misses,LLC-load-misses -p <pid> -- sleep 10

# Off-CPU: where is it BLOCKED? (often the real answer)
perf record -e sched:sched_switch -g -p <pid> -- sleep 10

Interpreting perf stat:

  • IPC < 1 → stalling (memory, branches, or interpreter overhead)
  • cache-miss rate > 20% of references → memory-bound
  • branch-miss rate > 5% → branchy code (tokenizer, interpreter)

Flame graphs#

Shell
perf record -F 99 -g -p <pid> -- sleep 30
perf script | stackcollapse-perf.pl | flamegraph.pl > cpu.svg

# Python (much easier)
py-spy record -o py.svg --pid <pid> --duration 30
py-spy record -o py.svg --pid <pid> --duration 30 --idle   # include blocked threads

Read a flame graph: width = time, height = stack depth. Look for wide plateaus. Ignore height — a deep stack that’s narrow costs nothing.

The --idle flag matters for inference servers: much of the interesting time is spent waiting (on the GPU, on a lock, on the network), and the default excludes it.

eBPF / bpftrace#

For questions no existing tool answers:

Shell
# Distribution of read() latencies
bpftrace -e 'tracepoint:syscalls:sys_enter_read /pid == PID/ { @start[tid] = nsecs; }
             tracepoint:syscalls:sys_exit_read  /@start[tid]/ {
               @us = hist((nsecs - @start[tid])/1000); delete(@start[tid]); }'

# Off-CPU time by stack
bpftrace -e 'kprobe:finish_task_switch { @[kstack] = count(); }'

# Which files is a process opening?
bpftrace -e 'tracepoint:syscalls:sys_enter_openat { printf("%s %s\n", comm, str(args->filename)); }'

eBPF runs safely in the kernel with ~1-2% overhead and no restarts. It is the right tool when you need a custom measurement in production.

The USE method#

For every resource, check three things:

Utilization  — % of time busy
Saturation   — queued work waiting
Errors       — error counts

Applied to an inference host:

ResourceUtilizationSaturationErrors
CPUtop, vmstat us+syrun queue r, nr_throttled—
Memoryfree, RSSswapping, memory.eventsOOM kills
GPU computenvidia-smi utilqueued kernels, launch gapsXid errors in dmesg
GPU memorymemory.usedpreemption/swapping in engineCUDA OOM
Diskiostat %utilawait, queue depthI/O errors
Networksar -n DEVretransmits, backlog dropsinterface errors
KV cacheoccupancy %queue depth, preemptions—

That last row is inference-specific and belongs on your dashboard (Section XI.08).


6. Under the hood#

Sampling profilers work by setting a hardware performance counter (or a timer) to fire an interrupt every N events; the handler walks the stack. Stack walking needs either frame pointers (-fno-omit-frame-pointer) or DWARF unwind info. Python builds and many distro packages omit frame pointers, which is why perf sometimes shows you nothing but [unknown] — use --call-graph dwarf or use py-spy, which understands CPython’s own frame objects directly.


7. Performance implications of profiling itself#

  • perf record -F 99 — ~1% overhead, safe in production.
  • py-spy — ~1-2%, safe, no restart, no instrumentation.
  • strace — 10-100x slowdown. Never leave it attached in production; use -c briefly or use eBPF instead.
  • cProfile — 2-10x on Python. Development only.
  • Nsight Systems — 5-20% depending on trace scope.

8. Production implications#

  • Ship py-spy and perf in your production image. The time to install tools is the time the incident is ongoing.
  • Continuous profiling (Parca, Pyroscope, Grafana Phlare) gives you the flame graph from before the incident, which is usually what you actually need.
  • Instrument phase timings in the app: queue wait, prefill, decode, detokenize, network. Application-level timings answer 80% of questions without any profiler.
  • Record a baseline profile for every release. Comparing two flame graphs is far more informative than reading one.

9. Common mistakes#

Guessing instead of measuring. The most expensive mistake. Intuition about performance is reliably wrong, including yours and including mine.

Profiling the wrong thing. If the GPU is the bottleneck, a CPU flame graph shows an idle event loop and tells you nothing. Determine the bottleneck layer first.

Ignoring off-CPU time. For inference servers, most wall-clock time is waiting. Use py-spy --idle and off-CPU profiling.

Using cProfile on a server. Distorts what you’re measuring.

Profiling without a warm-up. The first iterations include compilation, allocation, and cache warming.

Optimizing what the profiler shows without checking it’s on the critical path. A wide plateau in a background thread doesn’t affect latency.


10. Hands-on exercise#

A. Both profilers. Run the slow.go example. Produce a CPU flame graph (perf, or add runtime/pprof and use go tool pprof -http=:0 cpu.prof) and a strace -c summary. Write down what each tells you that the other doesn’t.

B. Classify with perf stat. Run perf stat on: a pure Python loop, a NumPy matmul, a random-access gather, and your model server. Build a table of IPC, cache-miss rate, and branch-miss rate. Explain each row.

C. Off-CPU profile. Profile a model server under load with py-spy record --idle. What fraction of time is spent waiting? On what?

D. Write a bpftrace one-liner. Measure the distribution of ioctl durations for your inference process (a proxy for CUDA driver call latency). What is the p99?

E. Build the USE checklist. Turn the table in section 5 into a script for your environment.


11. Interview questions#

  1. Describe your methodology for diagnosing a slow inference service.
  2. Sampling vs instrumenting profilers — when do you use each?
  3. What does a low IPC tell you? What about a high branch-miss rate?
  4. Why is off-CPU profiling important for inference servers?
  5. What is the USE method and how would you apply it to a GPU?
  6. How would you profile a production server without restarting it?
  7. Your flame graph shows 40% of time in a function, but removing it doesn’t help. Why might that be?

12. Further reading#

  • [FUNDAMENTAL] Brendan Gregg, Systems Performance (2nd ed.) and BPF Performance Tools
  • [REFERENCE] perf wiki; FlameGraph repository (brendangregg/FlameGraph)
  • [REFERENCE] py-spy, bpftrace documentation
  • Next: Section III — ML Fundamentals

↑↓ navigate↵ openesc close