PidokuInfra
02Computer Systems Foundations
On this page

Computer Systems Foundations

FoundationsModule 0213 topics~13h 45m

Topics, in order

01 CPU ArchitectureA CPU is a machine that fetches instructions, decodes them, executes them, and writes results. Modern CPUs do all four stages for dozens of instructions simultaneously, out of order, … Beginner 1h 02 Caches and the Memory HierarchyA hierarchy of progressively larger and slower memories, arranged so that frequently used data sits close to the compute units. Intermediate 1h 15m 03 RAM, DRAM, and NUMADRAM is main memory: cheap, large, and slow relative to caches. NUMA (Non-Uniform Memory Access) is what happens on multi-socket servers: each CPU socket has its own memory controller and … Intermediate 1h 04 SIMD and VectorizationSIMD = Single Instruction, Multiple Data. One instruction operates on a whole vector of values at once. Intermediate 1h 05 Threads, Processes, and Context Switching- Process — an isolated address space with its own memory, file descriptors, and at least one thread. Isolation is strong; communication is expensive. - Thread — an execution context … Beginner 1h 06 Memory Allocation and Virtual MemoryThis file is the conceptual ancestor of PagedAttention (Section V.10). Read it carefully; when you meet paged KV cache you will recognize every idea here. Intermediate 1h 15m 07 Storage and Model LoadingBefore a model can serve anything, tens to hundreds of gigabytes must travel from wherever they are stored into GPU memory. That journey is the dominant term in cold-start time, and … Intermediate 1h 08 Networking FundamentalsThe network appears in inference in three distinct roles, with wildly different requirements: Intermediate 1h 09 PCIe, DMA, and InterconnectsThe GPU is not in the CPU's memory. It is a peripheral at the end of a bus. Everything that moves between them crosses PCIe, and it is moved by a DMA engine, not by the CPU. Intermediate 1h 10 OS SchedulingThe kernel decides which runnable thread gets a CPU and for how long. On Linux this is CFS (Completely Fair Scheduler) by default, with real-time classes available for latency-critical work. Intermediate 45 min 11 Linux for Inference EngineersThis is a practical command reference organized by the question you're trying to answer. Work through it at a terminal, not by reading. Beginner 1h 15m 12 Containers, Namespaces, and cgroupsA container is not a virtual machine. It is a normal Linux process with three things applied: Intermediate 1h 13 Syscalls and Profiling BasicsSyscalls are the boundary between your program and the kernel. Profiling is finding out where time goes. Intermediate 1h 15m

About this module

Goal: give you the systems knowledge that inference engineering assumes and rarely teaches. Every file answers “why does an inference engineer care about this?” — if a systems topic doesn’t affect inference, it isn’t here.

You can serve a model without this section. You cannot debug one.

Files#

#FileLevelTime
01CPU architectureBeginner60 min
02Caches and the memory hierarchyIntermediate75 min
03RAM, DRAM, and NUMAIntermediate60 min
04SIMD and vectorizationIntermediate60 min
05Threads, processes, context switchingBeginner60 min
06Memory allocation and virtual memoryIntermediate75 min
07Storage and model loadingIntermediate60 min
08Networking fundamentalsIntermediate60 min
09PCIe, DMA, and interconnectsIntermediate60 min
10OS schedulingIntermediate45 min
11Linux for inference engineersBeginner75 min
12Containers, namespaces, cgroupsIntermediate60 min
13Syscalls and profiling basicsIntermediate75 min

The thread#

flowchart TD
  N0["Compute is fast; memory is slow<br/><b>01, 02</b>"]
  N1["So hardware hides latency with caches, prefetch, SIMD, threads<br/><b>02, 04, 05</b>"]
  N2["And the OS multiplexes all of it<br/><b>05, 06, 10</b>"]
  N3["Data must cross buses to reach the GPU (09) and networks to reach other nodes<br/><b>08</b>"]
  N4["And models must be read from storage before any of it starts<br/><b>07</b>"]
  N5["All of it inside containers with limits you must understand<br/><b>12</b>"]
  N6["And when it's slow, you need tools to see it<br/><b>11, 13</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6

  class N0,N1 neutral
  class N2 io
  class N3,N4 queue
  class N5 compute
  class N6 memory

Why this matters (concrete payoffs)#

  • Section 02 explains why tokenization is slow and how to fix it.
  • Section 03 explains a class of mysterious 2x slowdowns on dual-socket boxes.
  • Section 06 is the direct conceptual ancestor of PagedAttention (Section V.10).
  • Section 07 explains your 8-minute cold start and how to cut it to 40 seconds.
  • Section 09 explains why you never stream weights over PCIe per request.
  • Section 12 explains why your container sees 128 CPUs and gets throttled at 4.
  • Section 13 gives you the tools you will use in Section X.

↑↓ navigate↵ openesc close