PidokuInfra

OS Scheduling

Foundations Intermediate 45 min Difficulty 2/5 Topic 10 of 13

Prerequisites 05


1. What is it?#

The kernel decides which runnable thread gets a CPU and for how long. On Linux this is CFS (Completely Fair Scheduler) by default, with real-time classes available for latency-critical work.

For inference, one thread matters more than the others: the engine loop. If it is descheduled for 5 ms, every request in the batch experiences a 5 ms ITL hiccup.


2. Why does it exist?#

There are always more runnable threads than cores. The scheduler’s job is to share fairly while keeping interactive work responsive. “Fairly” is exactly the wrong policy when one of your threads is driving a $30,000 GPU and the others are logging.


3. Simple analogy#

A shared workshop with one supervisor allocating benches. By default they rotate everyone fairly. But if one worker is operating a machine that must be fed continuously or it stalls, a fair rotation is wrong — that worker needs priority. Telling the supervisor which worker that is is what nice, chrt, and cgroup CPU shares do.


4. Tiny example#

Watch the effect of contention on a latency-sensitive loop:

Go
// jitter.go — how late does a 1 ms timer fire when the CPUs are oversubscribed?
package main

import (
	"fmt"
	"runtime"
	"slices"
	"sync/atomic"
	"time"
)

func measure(n int) (p50, p99 time.Duration) {
	gaps := make([]time.Duration, n)
	prev := time.Now()
	for i := range gaps {
		time.Sleep(time.Millisecond) // target: 1 ms period
		now := time.Now()
		gaps[i], prev = now.Sub(prev), now
	}
	slices.Sort(gaps)
	return gaps[n/2], gaps[n*99/100]
}

func main() {
	p50, p99 := measure(2000)
	fmt.Printf("idle:   p50=%v p99=%v\n", p50.Round(10*time.Microsecond), p99.Round(10*time.Microsecond))

	var stop atomic.Bool
	for i := 0; i < 4*runtime.NumCPU(); i++ { // 4 CPU hogs per core
		go func() {
			for !stop.Load() {
			}
		}()
	}
	p50, p99 = measure(2000)
	fmt.Printf("loaded: p50=%v p99=%v\n", p50.Round(10*time.Microsecond), p99.Round(10*time.Microsecond))
	stop.Store(true)
}

You’ll see p99 jump from ~1.1 ms to tens of milliseconds. That jitter lands directly in your ITL distribution. The engine loop is exactly this kind of periodic, latency-sensitive task.


5. Technical explanation#

CFS in one paragraph#

Each thread accumulates vruntime proportional to CPU time consumed, scaled by its weight (from nice). The scheduler always runs the thread with the smallest vruntime. Fair-share by construction; no notion of deadlines. Timeslices adapt to the number of runnable threads — with many runnable threads, each gets a shorter slice and switches more often.

Key tunables:

Shell
# Minimum timeslice granularity (older kernels; EEVDF in 6.6+ differs)
cat /proc/sys/kernel/sched_min_granularity_ns
cat /proc/sys/kernel/sched_latency_ns

# Per-thread scheduling info
chrt -p <tid>
cat /proc/<pid>/task/<tid>/sched

Scheduling classes#

SCHED_OTHER (CFS)   default, fair share, nice -20..19
SCHED_BATCH         like OTHER but assumed non-interactive; fewer preemptions
SCHED_IDLE          only runs when nothing else wants the CPU
SCHED_FIFO (RT)     priority 1-99, runs until it yields or a higher-priority RT thread runs
SCHED_RR (RT)       FIFO with timeslices among equal priorities
SCHED_DEADLINE      EDF with runtime/period/deadline guarantees

Real-time classes are powerful and dangerous: a runaway SCHED_FIFO thread can lock up a core. sched_rt_runtime_us (default 950000 of 1000000) reserves 5% for non-RT work as a safety valve.

For inference, SCHED_FIFO on the engine loop is occasionally justified in ultra-low-latency deployments, but the usual answer is simpler: give it a dedicated core and reduce contention.

CPU affinity and isolation#

Shell
# Pin a process to specific CPUs
taskset -c 0-7 python server.py

# Pin an existing thread
taskset -pc 8 <tid>

# Boot-time isolation (removes CPUs from the general scheduler)
# kernel cmdline: isolcpus=8-15 nohz_full=8-15 rcu_nocbs=8-15

isolcpus + nohz_full gives you cores with almost no kernel interference — the standard technique in low-latency trading, occasionally worth it for inference engine loops.

Interrupts#

Network and NVMe interrupts land on specific CPUs. If they land on your engine loop’s core, you get jitter.

Shell
cat /proc/interrupts | head -20
# Move IRQ affinity away from your critical cores
echo 3 > /proc/irq/<N>/smp_affinity_list
systemctl status irqbalance     # often better to configure than to disable

6. Under the hood#

Shell
# Involuntary context switches for the engine thread
pidstat -w -t -p <pid> 1

# Scheduler latency (how long runnable threads waited)
perf sched latency -p <pid>
perf sched record -- sleep 10 && perf sched latency

# Run queue length
vmstat 1     # 'r' column: runnable threads. If r >> cores, you're oversubscribed.

perf sched latency gives you per-thread max and average wait time. If your engine thread shows a 20 ms max wait, you have found the source of an ITL outlier.


7. Performance implications#

  • Scheduler jitter appears directly in ITL percentiles. p99 ITL is often a scheduling story, not a GPU story.
  • Oversubscription costs twice: switching overhead plus cache pollution.
  • cgroup CPU throttling (file 12) causes 100 ms stalls — CFS bandwidth control has a 100 ms default period, and hitting the quota parks the thread until the period rolls over. This is a spectacularly common cause of latency spikes in Kubernetes.

8. Production implications#

  • Reserve cores for the engine loop. Use taskset/cgroup cpusets so the tokenizer pool and HTTP workers do not compete with it.
  • Keep the run queue shorter than the core count. Set thread counts explicitly (file 05).
  • Avoid CPU limits (quotas) on latency-critical containers; use requests + guaranteed QoS instead so you get a cpuset rather than throttling (Section XII.03).
  • Move IRQs off critical cores on dedicated hosts.
  • Monitor involuntary context switches and run-queue length as leading indicators of ITL degradation.

9. Common mistakes#

Setting a CPU limit on an inference pod. Guarantees periodic 100 ms stalls under load.

Running everything as one big thread pool. No isolation between latency-critical and throughput work.

Using SCHED_FIFO casually. One bug and the machine is unresponsive.

Assuming nice is enough. Under CFS, nice changes shares, not guarantees. A nice -20 thread still waits behind many runnable threads.

Blaming the GPU for ITL jitter without checking the scheduler.


10. Hands-on exercise#

A. Measure jitter. Run the example in section 4 with increasing numbers of burner threads. Plot p50 and p99 wake-up latency vs load. Then repeat with the measuring process pinned to an isolated core.

B. perf sched. Record scheduler events while a model server is under load. Find the maximum wait time for the engine thread. Correlate with your ITL outliers.

C. Throttling. Run a container with --cpus=2 and a workload wanting 4 CPUs. Observe nr_throttled and throttled_time in cpu.stat. Measure the resulting latency spikes.

D. Affinity. Pin an inference server’s engine process to dedicated cores and its HTTP workers elsewhere. Measure ITL p99 before and after.


11. Interview questions#

  1. How does CFS decide what to run, and why is “fair” wrong for an inference engine loop?
  2. What is cgroup CPU throttling and why does it cause 100 ms latency spikes?
  3. When would you use SCHED_FIFO, and what are the risks?
  4. How do you find out whether ITL jitter is caused by scheduling?
  5. Why can CPU limits be worse than no limits for a latency-sensitive service?

12. Further reading#

  • [FUNDAMENTAL] Gregg, Systems Performance, ch. 6
  • [REFERENCE] sched(7), chrt(1), taskset(1), perf-sched(1)
  • [REFERENCE] Kernel docs: Documentation/scheduler/
  • Next: 11 — Linux for inference engineers

↑↓ navigate↵ openesc close