PidokuInfra

Linux for Inference Engineers

Foundations Beginner 1h 15m Difficulty 2/5 Topic 11 of 13

Prerequisites 01-10 (skimmable)

This is a practical command reference organized by the question you’re trying to answer. Work through it at a terminal, not by reading.


1. What is it?#

The set of Linux commands and files that answer the questions inference engineers actually ask about a running system. Not a general Linux tutorial — a targeted toolkit.


2. Why does it exist?#

Because when an inference service is slow at 2 a.m., the difference between fixing it in ten minutes and three hours is knowing which four commands to run first.


3. The five-minute triage#

When something is wrong, run these in order. This is the inference-specific version of Brendan Gregg’s “USE method” checklist.

Shell
# 1. What is the GPU doing?
nvidia-smi
nvidia-smi dmon -s pucm -c 20        # power, util, clocks, memory over time

# 2. Is the CPU the bottleneck?
top -H -p $(pgrep -f 'vllm|sglang|triton' | head -1)
vmstat 1 5                            # r (runqueue), cs (switches), us/sy/wa

# 3. Where is memory going?
free -h
nvidia-smi --query-gpu=memory.used,memory.total --format=csv

# 4. Is it I/O?
iostat -x 1 5

# 5. Is it the network?
ss -s
sar -n DEV 1 5

# 6. What is Python actually doing?
py-spy dump --pid <pid>
py-spy top --pid <pid>

If you learn nothing else from this file, learn py-spy dump. For a hung or slow Python inference server it tells you, in one command and without a restart, exactly which line every thread is on.


4. GPU inspection#

Shell
# Full inventory
nvidia-smi -q | less

# The essentials, scriptable
nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,\
utilization.memory,temperature.gpu,power.draw,clocks.sm --format=csv

# Per-process GPU memory
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv

# Live monitoring (better than watch nvidia-smi)
nvidia-smi dmon -s pucvmet
#  p=power u=utilization c=clocks v=violations m=memory e=ecc t=pcie throughput

# Topology (critical before choosing parallelism)
nvidia-smi topo -m

# Is something throttling?
nvidia-smi -q -d PERFORMANCE | grep -A10 "Clocks Event Reasons"
#  SwPowerCap / HwSlowdown / SwThermalSlowdown → you are being throttled

# MIG status
nvidia-smi mig -lgi

# Reset a wedged GPU (drains all processes first!)
nvidia-smi --gpu-reset -i 0

Interpretation notes:

  • utilization.gpu = percentage of time at least one kernel was resident. Not a measure of useful work. See Section X.04.
  • utilization.memory = percentage of time the memory interface was busy. Closer to useful but still coarse.
  • Persistent SwPowerCap under load means you are power-limited; check power.draw vs power.limit.

5. Process and thread inspection#

Shell
# Threads of a process with per-thread CPU
top -H -p <pid>
ps -T -p <pid>

# What syscalls is it making?
strace -f -p <pid> -c            # summary counts
strace -f -p <pid> -e trace=ioctl -T   # CUDA goes through ioctl on /dev/nvidia*

# Open files and sockets
lsof -p <pid> | wc -l
ls -l /proc/<pid>/fd | wc -l

# Memory map
pmap -x <pid> | tail -3
grep -E 'VmRSS|VmHWM|VmSwap' /proc/<pid>/status

# Where is it blocked?
cat /proc/<pid>/wchan; echo
cat /proc/<pid>/stack 2>/dev/null

# Python-specific
py-spy dump --pid <pid>          # all threads' Python stacks
py-spy top --pid <pid>           # live profile
py-spy record -o prof.svg --pid <pid> --duration 30   # flame graph

6. Memory#

Shell
free -h
cat /proc/meminfo | grep -E 'MemAvailable|Dirty|Writeback|Huge'

# Page cache usage per file
vmtouch -v /models/model.safetensors

# Who is using memory?
ps aux --sort=-rss | head

# Was something OOM-killed?
dmesg -T | grep -i -E 'oom|killed process'
journalctl -k | grep -i oom

# cgroup memory (inside a container)
cat /sys/fs/cgroup/memory.current /sys/fs/cgroup/memory.max          # v2
cat /sys/fs/cgroup/memory/memory.usage_in_bytes                       # v1

Exit code 137 = SIGKILL = almost always the OOM killer. Check dmesg first.


7. Storage#

Shell
iostat -x 1                # %util, await, r/s, w/s, rMB/s
iotop -oPa                 # which process is doing the I/O
df -h; du -sh /models/*

# Benchmark
fio --name=seqread --rw=read --bs=1M --size=8G --direct=1 \
    --numjobs=4 --group_reporting --filename=/models/testfile

await (average I/O wait in ms) is the number to watch. NVMe should be < 1 ms; if you see 20 ms, you are queueing or on network storage.


8. Network#

Shell
ss -s                                    # summary of socket states
ss -tan state established | wc -l        # concurrent connections
ss -tin | head                           # per-socket rtt, cwnd, retransmits
sar -n DEV 1 5                           # per-interface throughput
netstat -s | grep -iE 'retrans|overflow|drop'
ethtool -S eth0 | grep -iE 'drop|error'

# End-to-end client view (the real TTFT)
curl -N -o /dev/null -s -w 'dns=%{time_namelookup} conn=%{time_connect} \
tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total}\n' \
  -X POST https://endpoint/v1/chat/completions -d @req.json -H 'Content-Type: application/json'

9. cgroups and containers#

Shell
# Am I in a container? What are my limits?
cat /proc/self/cgroup
cat /sys/fs/cgroup/cpu.max          # "quota period" — e.g. "400000 100000" = 4 CPUs
cat /sys/fs/cgroup/memory.max
cat /sys/fs/cgroup/cpu.stat         # nr_throttled, throttled_usec  ← check this!

# The lie os.cpu_count() tells
nproc                                # host CPUs, not your quota
python -c "import os; print(os.cpu_count(), len(os.sched_getaffinity(0)))"

nr_throttled climbing means CFS bandwidth throttling — see files 10 and 12.


10. System configuration for inference hosts#

Shell
# Limits
ulimit -n                     # file descriptors — raise to 65535+
ulimit -l                     # locked memory — needs to be unlimited for pinned buffers + RDMA

# In /etc/security/limits.conf:
#   * soft nofile 65535
#   * hard nofile 65535
#   * soft memlock unlimited
#   * hard memlock unlimited

# Swap: turn it off on GPU hosts
swapoff -a
cat /proc/sys/vm/swappiness

# Huge pages
cat /sys/kernel/mm/transparent_hugepage/enabled

# NUMA
numactl --hardware

# Persistence mode (avoids driver re-init latency)
nvidia-smi -pm 1

# Lock clocks for reproducible benchmarking
nvidia-smi -lgc <min>,<max>      # then -rgc to reset

nvidia-smi -pm 1 and locked clocks are essential for reproducible benchmarking (Section X.07). Without them, your numbers vary 10-20% run to run.


11. Handy one-liners#

Shell
# Watch GPU memory of a specific process over time
watch -n1 'nvidia-smi --query-compute-apps=pid,used_memory --format=csv'

# Correlate GPU util with CPU util
paste <(nvidia-smi dmon -c 60 -s u | awk '{print $2}') <(vmstat 1 60 | awk '{print $13+$14}')

# Find the biggest tensors in a safetensors file
python - <<'PY'
from safetensors import safe_open
import functools, operator
with safe_open("model.safetensors","pt") as f:
    sizes = {k: functools.reduce(operator.mul, f.get_slice(k).get_shape(), 1) for k in f.keys()}
for k,v in sorted(sizes.items(), key=lambda x:-x[1])[:10]:
    print(f"{v/1e6:10.1f}M  {k}")
PY

# Tail structured logs for slow requests
journalctl -u inference -f | jq 'select(.ttft_ms > 1000)'

12. Hands-on exercise#

A. Build your own triage script. Write triage.sh that collects, in one run: GPU state, top CPU threads, memory, I/O, network sockets, cgroup limits and throttling, and a py-spy dump. Make it output a timestamped file. Put it in your production image.

B. Detect throttling. Run a GPU-heavy workload and watch nvidia-smi -q -d PERFORMANCE. Can you induce a power cap? What happens to throughput?

C. Break something on purpose. On a test box: exhaust file descriptors, fill the GPU memory, saturate the disk. For each, determine which command in this file would have identified it fastest.

D. Reproducible benchmark setup. Enable persistence mode, lock clocks, disable swap, and run the same benchmark 10 times. Compare variance to the unlocked case.


13. Interview questions#

  1. An inference server is slow. Walk me through your first five commands.
  2. How do you tell whether a GPU is being thermally or power throttled?
  3. Your pod restarted with exit code 137. What happened and how do you confirm?
  4. Why is nproc misleading inside a container, and what should you use instead?
  5. How would you find out which Python line a hung inference server is stuck on?
  6. What system settings would you change on a fresh GPU host before serving traffic?

14. Further reading#

  • [FUNDAMENTAL] Brendan Gregg, Systems Performance; the USE method
  • [REFERENCE] nvidia-smi manual; NVIDIA DCGM documentation
  • [REFERENCE] py-spy, perf, bpftrace documentation
  • Next: 12 — Containers, namespaces, cgroups

↑↓ navigate↵ openesc close