The idea in one minute#
A profile answers “which code is using the resource?”. A CPU profiler interrupts the program about a hundred times a second and records the call stack; stacks that appear often are where time goes. Continuous profiling leaves that running in production at low overhead, so you can compare this week against last week or one version against another.
eBPF lets small, verified programs run inside the Linux kernel. It is how modern tools profile every process on a machine and produce traces and metrics without changing or even restarting the application.
An analogy#
To find out how office staff spend their day you could ask everyone to log every task (instrumentation: accurate, intrusive). Or you could walk through the office at random moments and note what each person is doing (sampling). After a few hundred walks you know where the hours go, and nobody had to change how they work.
A picture#
flowchart TB
P["Running process"] -->|"interrupt 100 times per second"| S["Capture the call stack"]
S --> AGG["Count identical stacks"]
AGG --> FG["Flame graph<br/>width = share of samples"]
FG --> DIFF["Diff two profiles<br/>before and after a deploy"]
subgraph K["Linux kernel"]
EB["eBPF program<br/>attached to a timer, syscall or function"]
end
EB -->|"stacks from every process,<br/>no code changes"| AGG
class P compute
class S,AGG queue
class FG,DIFF memory
class EB ioHow it really works#
Kinds of profile#
| Profile | Samples | Answers |
|---|---|---|
| CPU | Stacks that were on-CPU | What burns CPU? |
| Heap / allocations | Stacks that allocated memory | What allocates, what is retained? |
| Goroutine | All goroutine stacks | What is everyone waiting on? |
| Block / mutex | Stacks that waited | Where is the contention? |
| Off-CPU / wall clock | Stacks that were not running | Why is it slow while the CPU is idle? |
Profiling Go#
Add one import and the profiles are served over HTTP:
import _ "net/http/pprof" // registers /debug/pprof/* on the default muxgo tool pprof -http=:0 http://localhost:6060/debug/pprof/profile?seconds=30 # CPU
go tool pprof -http=:0 http://localhost:6060/debug/pprof/heap
go test -bench . -cpuprofile cpu.out && go tool pprof -http=:0 cpu.outReading a flame graph#
- Each box is a function; the box below is its caller.
- Width is the share of samples — the only thing that matters. Left-to-right order is alphabetical, not time.
- Look for wide plateaus at the top: functions that are themselves expensive.
- A diff flame graph colours what grew and what shrank between two profiles — the fastest way to find a regression.
Continuous profiling#
Collect a short profile from every process every few seconds, tag it with service and version, and store it. Overhead is typically a percent or two of CPU. What it buys:
- “CPU per request rose 12% in v1.8.2” — diff the two versions.
- “Which function costs the most across the whole fleet?” — the answer is often a serializer, a logger or a regular expression, not business logic.
- With span IDs attached to samples, a slow span links straight to the code that ran in it.
Tools: Grafana Pyroscope, Parca, Polar Signals, and the profilers built into the large commercial platforms. OpenTelemetry’s Profiles signal entered public alpha in March 2026; it defines a common format based on pprof and ships an eBPF profiling agent (donated by Elastic) that runs as part of the Collector.
eBPF in one page#
An eBPF program is a small piece of bytecode loaded into the kernel, checked by a verifier (it must terminate and cannot touch arbitrary memory), and attached to a hook:
| Hook | Fires on | Used for |
|---|---|---|
| Timer / perf event | A fixed frequency | CPU profiling of every process |
| kprobe / tracepoint | A kernel function or event | Syscalls, scheduling, disk and network I/O |
| uprobe | A function in a user program or library | TLS reads/writes, HTTP handlers, GPU runtime calls |
| Socket / TC / XDP | Network packets | Flow metrics, service maps |
Programs write results to maps that a user-space agent reads.
What it gives observability:
- Zero-code instrumentation. OpenTelemetry eBPF Instrumentation (OBI, derived from Grafana Beyla) watches HTTP, gRPC and SQL calls and emits spans and RED metrics for any language — useful for services you cannot or will not modify.
- Whole-machine profiling, including the kernel and native libraries.
- Network observability without sidecars (Cilium Hubble, Pixie, Coroot).
Its limits: you get what can be seen from outside — protocol-level spans, not your business attributes. It needs a recent kernel and elevated privileges. Encrypted traffic needs uprobes on the TLS library. Use it for breadth, and SDK instrumentation for depth.
Profiling and GPUs#
A CPU profiler shows the host side of an inference server: tokenization, Python or Go overhead, scheduling — and stacks that sit in a driver call waiting for the GPU. What happens on the device needs GPU tools (Nsight Systems, the PyTorch profiler, CUPTI). eBPF tools can attach uprobes to the CUDA runtime to time kernel launches and memory copies per process; this is an active and fast-moving area (V.02, VI.02).
Code#
A program with an obvious hotspot that profiles itself and prints where the samples went.
// hot.go — profile a program from inside and print the top of the CPU profile.
package main
import (
"fmt"
"os"
"os/exec"
"regexp"
"runtime/pprof"
"strings"
)
var re = regexp.MustCompile(`^[a-z]+_[0-9]+$`)
// slowValidate compiles nothing new, but regex matching is far costlier than it looks.
func slowValidate(keys []string) int {
n := 0
for _, k := range keys {
if re.MatchString(k) {
n++
}
}
return n
}
func fastValidate(keys []string) int {
n := 0
for _, k := range keys {
if i := strings.IndexByte(k, '_'); i > 0 && i < len(k)-1 {
n++
}
}
return n
}
func main() {
keys := make([]string, 200000)
for i := range keys {
keys[i] = fmt.Sprintf("tenant_%d", i)
}
f, err := os.CreateTemp("", "cpu-*.prof")
if err != nil {
panic(err)
}
defer os.Remove(f.Name())
pprof.StartCPUProfile(f)
total := 0
for i := 0; i < 20; i++ {
total += slowValidate(keys) + fastValidate(keys)
}
pprof.StopCPUProfile()
f.Close()
fmt.Println("validated:", total)
// Ask the Go toolchain to summarize the profile: flat = time in the function itself.
out, err := exec.Command("go", "tool", "pprof", "-top", "-nodecount=8", f.Name()).CombinedOutput()
if err != nil {
fmt.Println("could not run go tool pprof:", err)
return
}
fmt.Println(string(out))
}Both functions do the same job. The profile shows almost all samples under slowValidate and
the regexp package — which is the kind of finding continuous profiling turns up fleet-wide.
Remember this#
- A profile is aggregated stack samples. Width in a flame graph is share of the resource.
- Continuous profiling makes “what changed between versions?” a diff.
- eBPF observes from the kernel: every process, no code changes, protocol-level detail only.
- CPU profiles show the host side of GPU workloads; device time needs GPU-aware tools.
Try it#
- Run
hot.go. Then addimport _ "net/http/pprof", serve on:6060, and open the flame graph withgo tool pprof -http=:0. - Take a heap profile of a program that builds a large slice. Find the allocating line.
- Write down three things an eBPF-generated span cannot contain that an SDK span can.
Check yourself#
- What does the width of a box in a flame graph mean?
- Why is continuous profiling cheap enough to leave on?
- What does the eBPF verifier guarantee?