PidokuInfra

Inference Engineering

How modern LLM inference works — from a single matrix multiply to a multi-tenant serving platform.

Start reading
  • 5 levels
  • 14 modules
  • 186 topics
  • ~354h 45m

Contents

  1. Foundations

    Build the mental model.

    36 topics · ~40h 45m

    1. 01 FundamentalsGoal of this section: give you the vocabulary and the physics of inference. By the end you will be able to read any inference discussion without getting lost, and — more importantly — you … 11 topics · ~11h 30m
      1. 00Module overview
      2. 01What Is Inference?Beginner45 min
      3. 02Training vs InferenceBeginner45 min
      4. 03Model Anatomy: Parameters, Weights, ActivationsBeginner1h
      5. 04Tensors and the Forward PassBeginner1h
      6. 05Latency, Throughput, and the Metrics That MatterBeginner1h 15m
      7. 06Batching: The Central TradeoffBeginner1h
      8. 07Compute vs MemoryIntermediate1h 15m
      9. 08FLOPs, Bandwidth, and Arithmetic IntensityIntermediate1h 30m
      10. 09CPU vs GPUBeginner1h
      11. 10Why Inference Is Not Normal Backend EngineeringIntermediate1h
      12. 11Why Inference Gets Expensive at ScaleIntermediate1h
    2. 02 Computer Systems FoundationsGoal: give you the systems knowledge that inference engineering assumes and rarely teaches. Every file answers "why does an inference engineer care about this?" — if a systems topic doesn't … 13 topics · ~13h 45m
      1. 00Module overview
      2. 01CPU ArchitectureBeginner1h
      3. 02Caches and the Memory HierarchyIntermediate1h 15m
      4. 03RAM, DRAM, and NUMAIntermediate1h
      5. 04SIMD and VectorizationIntermediate1h
      6. 05Threads, Processes, and Context SwitchingBeginner1h
      7. 06Memory Allocation and Virtual MemoryIntermediate1h 15m
      8. 07Storage and Model LoadingIntermediate1h
      9. 08Networking FundamentalsIntermediate1h
      10. 09PCIe, DMA, and InterconnectsIntermediate1h
      11. 10OS SchedulingIntermediate45 min
      12. 11Linux for Inference EngineersBeginner1h 15m
      13. 12Containers, Namespaces, and cgroupsIntermediate1h
      14. 13Syscalls and Profiling BasicsIntermediate1h 15m
    3. 03 Machine Learning FundamentalsGoal: teach exactly the ML you need to reason about inference — no more, no less. 12 topics · ~15h 30m
      1. 00Module overview
      2. 01Vectors and MatricesBeginner1h
      3. 02Matrix Multiplication by HandBeginner1h 15m
      4. 03Tensors, Shapes, and BroadcastingBeginner1h
      5. 04Neural Networks from ScratchBeginner1h 30m
      6. 05Activation FunctionsBeginner45 min
      7. 06EmbeddingsBeginner1h
      8. 07Softmax and Numerical StabilityIntermediate1h 15m
      9. 08Attention from First PrinciplesIntermediate2h
      10. 09The TransformerIntermediate2h
      11. 10Training vs Inference: The MathIntermediate45 min
      12. 11Number Formats: FP32, FP16, BF16, FP8, INT8, INT4Intermediate1h 30m
      13. 12Quantization FundamentalsIntermediate1h 30m
  2. Intermediate

    Learn the optimization techniques.

    40 topics · ~49h

    1. 06 GPU ComputingGoal: understand the machine well enough to predict its behavior, read a profiler, and write a kernel when you need to. 12 topics · ~15h
      1. 00Module overview
      2. 01GPU ArchitectureBeginner1h 15m
      3. 02Execution Model: Threads, Warps, Blocks, GridsIntermediate1h 15m
      4. 03GPU Memory HierarchyIntermediate1h 15m
      5. 04Bandwidth and the Roofline on GPUIntermediate1h
      6. 05Kernel Launch and Host-Device InteractionIntermediate1h
      7. 06Streams and SynchronizationIntermediate1h
      8. 07CUDA GraphsAdvanced1h
      9. 08CUDA Programming FundamentalsAdvanced2h
      10. 09Coalescing and Access PatternsAdvanced1h 15m
      11. 10Occupancy and Warp DivergenceAdvanced1h 15m
      12. 11Tensor CoresAdvanced1h 15m
      13. 12Profiling CUDAAdvanced1h 30m
    2. 07 Inference OptimizationGoal: a catalogue of the techniques that make inference faster, each analyzed the same way: 14 topics · ~17h 30m
      1. 00Module overview
      2. 01Optimization MethodologyIntermediate1h
      3. 02Quantization Overview and Decision GuideIntermediate1h 30m
      4. 03Weight-Only Quantization: GPTQ and AWQAdvanced1h 30m
      5. 04Activation Quantization and SmoothQuantAdvanced1h 15m
      6. 05FP8 and INT8 InferenceAdvanced1h 15m
      7. 06INT4 and Low-Bit InferenceAdvanced1h 15m
      8. 07Kernel and Operator FusionAdvanced1h
      9. 08CUDA Graphs in ServingAdvanced45 min
      10. 09TensorRT and TensorRT-LLMAdvanced1h 15m
      11. 10FlashAttentionAdvanced2h
      12. 11KV Cache OptimizationAdvanced1h 30m
      13. 12KV Cache QuantizationAdvanced1h
      14. 13Speculative Decoding in PracticeAdvanced1h 15m
      15. 14Pruning, Distillation, and Architecture OptimizationAdvanced1h
    3. 08 Inference Serving SystemsGoal: understand how models become services — and, more importantly, understand why real inference servers are built the way they are. 14 topics · ~16h 30m
      1. 00Module overview
      2. 01Anatomy of an Inference ServerIntermediate1h 15m
      3. 02APIs: HTTP, gRPC, StreamingIntermediate1h
      4. 03Queues, Scheduling, and Admission ControlAdvanced1h 30m
      5. 04Batching in ServersAdvanced1h
      6. 05Backpressure, Timeouts, and RetriesAdvanced1h 15m
      7. 06Load Balancing and RoutingAdvanced1h 15m
      8. 07Autoscaling and Cold StartsAdvanced1h 15m
      9. 08Model Lifecycle: Loading, Warmup, UnloadingIntermediate1h
      10. 09Multi-Model ServingAdvanced1h 15m
      11. 10Versioning, Canary Deployments, and A/B TestingIntermediate1h
      12. 11vLLM ArchitectureAdvanced1h 30m
      13. 12SGLang ArchitectureAdvanced1h
      14. 13TGI, Triton, and ONNX RuntimeAdvanced1h 15m
      15. 14Choosing a Serving StackIntermediate1h
  3. Advanced

    Study systems at production scale.

    34 topics · ~41h 45m

    1. 09 Distributed InferenceGoal: understand what happens when one GPU isn't enough — and, equally important, when adding GPUs makes things worse. 12 topics · ~14h 45m
      1. 00Module overview
      2. 01Why Distribute (and When Not To)Intermediate1h
      3. 02Data ParallelismIntermediate45 min
      4. 03Tensor ParallelismAdvanced2h
      5. 04Pipeline ParallelismAdvanced1h 15m
      6. 05Expert ParallelismAdvanced1h 15m
      7. 06Sequence and Context ParallelismAdvanced1h
      8. 07Collectives and NCCLAdvanced1h 15m
      9. 08Interconnects: NVLink, PCIe, InfiniBand, RDMAAdvanced1h 15m
      10. 09Communication Cost MathAdvanced1h 30m
      11. 10Multi-Node InferenceAdvanced1h 15m
      12. 11Distributed KV CacheAdvanced1h
      13. 12When More GPUs Make Things WorseAdvanced1h 15m
    2. 10 Memory and Performance EngineeringGoal: answer "why is my inference system slow?" systematically, every time, without guessing. 11 topics · ~13h 30m
      1. 00Module overview
      2. 01The Performance Debugging MethodologyAdvanced1h 30m
      3. 02The Roofline Model in PracticeAdvanced1h
      4. 03Memory-Bound vs Compute-BoundAdvanced1h
      5. 04GPU Utilization Is a LieIntermediate1h
      6. 05Memory Fragmentation and AllocatorsAdvanced1h 15m
      7. 06OOM: Causes and CuresAdvanced1h
      8. 07Benchmarking Inference CorrectlyAdvanced1h 30m
      9. 08GPU Profiling with NsightAdvanced1h 30m
      10. 09CPU and Python ProfilingIntermediate1h
      11. 10Metrics, Prometheus, and DashboardsIntermediate1h 15m
      12. 11Case StudiesAdvanced1h 30m
    3. 11 Production Inference EngineeringGoal: run an inference service that meets commitments, costs what you planned, and doesn't wake you at 3 a.m. 11 topics · ~13h 30m
      1. 00Module overview
      2. 01SLOs, SLIs, and Latency BudgetsIntermediate1h 15m
      3. 02Capacity PlanningAdvanced1h 30m
      4. 03Cost Per TokenAdvanced1h 15m
      5. 04Autoscaling in ProductionAdvanced1h
      6. 05Fault Tolerance and Failure ModesAdvanced1h 15m
      7. 06Multi-Region and Disaster RecoveryAdvanced1h
      8. 07Model RolloutsIntermediate1h
      9. 08Observability for LLM ServicesIntermediate1h
      10. 09Security and IsolationAdvanced1h 15m
      11. 10Abuse Prevention, Rate Limiting, and QuotasAdvanced1h
      12. 11Scenario: 70B Model, 10,000 Concurrent Users, Fixed BudgetExpert2h
  4. Expert

    Design platforms and read the frontier.

    34 topics · ~39h 30m

    1. 12 Inference Platform EngineeringGoal: build the system that lets many teams serve many models on shared hardware, safely and efficiently. 11 topics · ~13h
      1. 00Module overview
      2. 01Platform Overview: Control Plane vs Data PlaneAdvanced1h
      3. 02Model RegistryIntermediate1h
      4. 03Kubernetes for GPUsAdvanced1h 30m
      5. 04GPU Scheduling and Model PlacementAdvanced1h 30m
      6. 05Multi-Tenancy and QuotasAdvanced1h 15m
      7. 06Routing StrategiesAdvanced1h 15m
      8. 07Prefix-Aware RoutingAdvanced1h
      9. 08Inference GatewaysAdvanced1h 15m
      10. 09Cascades and FallbacksAdvanced1h
      11. 10Cluster and Capacity ManagementAdvanced1h
      12. 11Building the Platform: A RoadmapAdvanced1h 15m
    2. 13 Advanced LLM InferenceGoal: the techniques that define the current frontier of production inference — and an honest assessment of which are ready. 12 topics · ~15h 15m
      1. 00Module overview
      2. 01MQA, GQA, and MLAAdvanced1h 15m
      3. 02Mixture of ExpertsAdvanced1h 30m
      4. 03MoE Serving and Expert ParallelismAdvanced1h 15m
      5. 04Long-Context InferenceAdvanced1h 30m
      6. 05Chunked PrefillAdvanced1h
      7. 06Disaggregated Prefill/DecodeAdvanced1h 30m
      8. 07KV Offloading and TransferAdvanced1h 15m
      9. 08Request Scheduling and PrioritiesAdvanced1h
      10. 09Speculative Decoding VariantsAdvanced1h 15m
      11. 10Low-Bit Inference: FP4 and BeyondAdvanced1h
      12. 11Triton KernelsAdvanced1h 30m
      13. 12Compilers: torch.compile and BeyondAdvanced1h 15m
    3. 14 Research and FrontierGoal: be able to read the field, judge what matters, and know where it's heading. 11 topics · ~11h 15m
      1. 00Module overview
      2. 01How to Read an Inference PaperAdvanced1h 15m
      3. 02Attention ResearchAdvanced1h
      4. 03KV Cache ResearchAdvanced1h
      5. 04Quantization ResearchAdvanced1h
      6. 05Decoding ResearchAdvanced1h
      7. 06MoE ResearchAdvanced45 min
      8. 07Long-Context ResearchAdvanced1h
      9. 08Scheduling and Systems ResearchAdvanced1h
      10. 09NVIDIA Architecture TrendsAdvanced1h
      11. 10Alternative AcceleratorsAdvanced1h 15m
      12. 11Memory-Centric and Disaggregated FuturesAdvanced1h

Build

Reference

About

From “what is a matrix multiply” to “design the serving stack for a 70B model with 10,000 concurrent users on a fixed GPU budget.”

This repository is a textbook + lab + roadmap for inference engineering: the discipline of running trained machine-learning models — especially large language models — as fast, cheap, and reliable production systems.


What this repository teaches#

Inference engineering sits at the intersection of four fields that are usually taught separately:

   Machine learning            Computer architecture
   (what the model is)         (what the hardware does)
            \                        /
             \                      /
              +--------------------+
              |  INFERENCE         |
              |  ENGINEERING       |
              +--------------------+
             /                      \
            /                        \
   Distributed systems          Production operations
   (many machines)              (SLOs, cost, reliability)

Most people learn one corner and stall. A model researcher can explain attention but cannot explain why their server falls over at 40 concurrent users. A backend engineer can build a bulletproof API but cannot explain why a 7B model at batch size 1 leaves 97% of the GPU idle. This curriculum builds all four corners in the order in which they actually depend on each other.

By the end you should be able to:

  • Read a model config and predict its memory footprint, KV cache growth, and peak throughput before deploying it.
  • Look at a slow inference service and systematically determine whether it is memory-bandwidth bound, compute bound, queueing bound, or bound by something silly like tokenizer overhead.
  • Explain — and implement — continuous batching, paged attention, prefix caching, speculative decoding, tensor parallelism, and quantization.
  • Write a CUDA kernel and a Triton kernel, and profile both.
  • Design a multi-tenant inference platform with routing, autoscaling, quotas, and cost attribution.
  • Read a frontier inference paper and judge whether it is production-relevant or a lab curiosity.

Who this is for#

You areThis will
A backend / platform / DevOps engineer moving into AI infraGive you the ML and GPU foundations you’re missing, using systems vocabulary you already have
An ML engineer who trains modelsGive you the systems, GPU, and production layers that training rarely teaches
A student targeting inference / performance rolesGive you a complete, sequential path with projects and interview questions
An SRE inheriting a GPU fleetGive you the mental models to reason about capacity, cost, and failure

Assumed: you can program in Go, you understand basic software engineering, you can use a terminal, and you know what an HTTP request is.

Not assumed: any ML, linear algebra beyond high school, GPU knowledge, CUDA, or distributed systems.


Prerequisites#

Practical, not theoretical:

  • Go 1.22+. Every runnable example is a single-file Go program using only the standard library: save it, go run it.
  • Python 3.10+ is useful later, but only to run existing engines (vLLM, PyTorch) that your Go code talks to. You will not need to write it.
  • Comfort with Linux/macOS command line.
  • Optional but strongly recommended: access to any NVIDIA GPU. A free Colab T4, a rented A10G/L4, or a consumer RTX card is enough for 90% of the exercises. Sections VI, IX, and several projects have CPU-only fallbacks marked [No GPU? Do this instead].
  • git, docker for the later sections.

You do not need an H100. You need curiosity and the willingness to measure things.


Why Go for inference engineering#

The arithmetic of a model runs in CUDA kernels driven by C++ and Python engines. That is not going to change, and this course teaches you how those engines work from the inside.

But most inference engineering is everything around the kernel: schedulers, batchers, KV-cache managers, routers, gateways, rate limiters, autoscalers, load generators. That is concurrent systems programming, and it is what Go is for. So in this course:

  • Mechanisms are built in Go. A transformer forward pass, a KV cache, paged attention, continuous batching, speculative decoding, quantization, a fair scheduler, a consistent-hash router — each as a small program you can read in one sitting and run on a laptop.
  • Real engines are measured from Go. Clients that drive any OpenAI-compatible server and measure time-to-first-token, inter-token latency and goodput correctly.
  • Framework internals stay in their own language. Where a lesson is specifically about a PyTorch, CUDA or Triton API (mostly module VI and parts of X and XIII), the snippet is shown as that API really looks. Translating it would teach you something that does not exist.

Every complete Go program in the lessons is compiled and vetted by tools/gocheck, so what you copy is known to build.

For the hardware underneath all of this, see the sister course GPU Engineering.


Curriculum map#

flowchart LR
  subgraph F["Foundations"]
    direction TB
    I["I Fundamentals"]
    II["II Computer Systems"]
    III["III ML Fundamentals"]
    IV["IV NN Inference"]
  end
  subgraph C["Core"]
    direction TB
    V["V LLM Inference"]
    VI["VI GPU Computing"]
  end
  subgraph B["Build"]
    direction TB
    VII["VII Optimization"]
    VIII["VIII Serving Systems"]
  end
  subgraph S["Scale"]
    direction TB
    IX["IX Distributed"]
    X["X Memory & Performance"]
  end
  subgraph O["Operate"]
    direction TB
    XI["XI Production"]
    XII["XII Platform"]
  end
  subgraph R["Frontier"]
    direction TB
    XIII["XIII Advanced LLM"]
    XIV["XIV Research"]
  end
  F --> C --> B --> S --> O --> R

  class I,II,III,IV neutral
  class V,VI memory
  class VII,VIII compute
  class IX,X queue
  class XI,XII io
  class XIII,XIV warn
I    Fundamentals ............... the vocabulary and the physics
II   Computer Systems .......... CPU, memory, OS, network, PCIe
III  ML Fundamentals ........... linear algebra -> transformers -> number formats
IV   NN Inference .............. the forward pass as an engineering artifact
V    LLM Inference ............. prefill/decode, KV cache, batching, PagedAttention  [CORE]
VI   GPU Computing ............. architecture, CUDA, memory hierarchy, profiling      [CORE]
VII  Inference Optimization .... quantization, fusion, FlashAttention, spec decoding
VIII Serving Systems ........... vLLM/SGLang/TGI/Triton architectures, APIs, queues
IX   Distributed Inference ..... TP/PP/EP, NCCL, NVLink, when more GPUs hurt
X    Memory & Performance ...... roofline, profiling, "why is my system slow?"
XI   Production Inference ...... SLOs, capacity, cost/token, observability, security
XII  Platform Engineering ...... registry, K8s GPU scheduling, routing, gateways
XIII Advanced LLM Inference .... MoE, MLA, disaggregation, chunked prefill, kernels
XIV  Research & Frontier ....... how to read the field and judge what matters

Each section has its own README.md with a per-file index, difficulty labels, and estimated time.


The main path (sequential)#

Read I → II → III → IV → V → VI → VII → VIII → IX → X → XI → XII → XIII → XIV.

This is the intended order and every file assumes only what came before it. See ROADMAP.md for the full dependency graph, time estimates, and checkpoints.

If you are in a hurry (“I need to be useful in 3 weeks”)#

I (all)  ->  III.11-12  ->  V (all)  ->  VI.01-05  ->  VII.01-02, 10
         ->  VIII.01-04, 11  ->  X.01-03  ->  XI.01-03

This is the “operate an LLM serving stack competently” path. It skips the deep hardware and distributed material, which you can return to.

If you already know ML#

Skim III, read IV.03/IV.07/IV.09 closely, then go straight to V and VI. Do not skip V — it is where most ML people have the largest gaps.

If you already know systems/GPUs#

Skim II and VI, read III fully (do not skip: the attention math matters), then V, VII, IX.


How to use this repository#

  1. Read sequentially, but do the exercises. Every concept file ends with a hands-on exercise. The exercises are where the learning actually happens. Reading alone produces the illusion of understanding — inference engineering punishes that illusion very quickly, because the hardware does not care what you believe.

  2. Keep a measurement journal. Every time a file asks you to measure something, record the number and your machine. Six months later, “an H100 does ~3.3 TB/s of HBM bandwidth” will be a fact you own rather than one you looked up.

  3. Do the projects in order. projects/ contains fifteen builds that mirror the sections. Project N is designed to be doable after section N-ish. Each project README lists its exact prerequisites.

  4. Track progress in PROGRESS.md. Check items off. It is a real checklist of every file and project.

  5. When you meet an unfamiliar term, check GLOSSARY.md first. It is alphabetical and deliberately blunt.

  6. Read the “Common mistakes” section of every file twice. Those sections encode the mistakes that cost real teams real money.

File format#

Every major concept file follows the same twelve-part structure:

1.  What is it?              7.  Performance implications
2.  Why does it exist?       8.  Production implications
3.  Simple analogy           9.  Common mistakes
4.  Tiny example             10. Hands-on exercise
5.  Technical explanation    11. Interview questions
6.  Under the hood           12. Further reading

Shorter connective files (indexes, methodology notes) use a lighter structure. Every file carries a header:


Project progression#

#ProjectAfter sectionWhat it proves
01NumPy inference engineIIIYou understand a forward pass end to end
02CPU matmul benchmarkII, IIIYou can measure FLOPs and bandwidth
03First GPU kernelVIYou can write and launch CUDA
04Tiny transformer engineIV, VYou understand transformer inference mechanically
05KV cacheVYou understand the single most important optimization
06LLM inference serverVIIIYou can wrap a model in a real API
07Dynamic batchingVIIIYou understand throughput/latency tradeoffs
08Continuous batchingV, VIIIYou understand modern LLM serving
09KV cache managerV, XYou understand paged memory management
10Quantized inferenceVIIYou can trade accuracy for speed deliberately
11GPU benchmark suiteVI, XYou can characterize hardware
12Inference gatewayVIII, XIYou can build the production front door
13Multi-GPU inferenceIXYou understand tensor parallelism concretely
14Distributed inferenceIXYou understand multi-node and collectives
15Mini inference platformXIIYou can build the whole thing

A note on honesty#

This curriculum distinguishes three tiers of knowledge and labels them:

  • [FUNDAMENTAL] — physics and math. True in 1995, true in 2035. Memory bandwidth, arithmetic intensity, Amdahl’s law.
  • [ESTABLISHED] — production-proven engineering. Continuous batching, PagedAttention, FP8/INT8 quantization, tensor parallelism. Deployed at scale by many organizations.
  • [EMERGING] — promising but not universally production-ready. Some speculative decoding variants, aggressive low-bit formats, disaggregated serving in its newer forms.

Numbers in this curriculum (bandwidths, prices, model sizes) are illustrative and drift. The methods for computing them do not. Always re-derive with current hardware specs.


Repository files#

  • ROADMAP.md — the full I→XIV path with dependencies and checkpoints
  • GLOSSARY.md — every term, defined bluntly
  • GPU-HARDWARE.md — GPU memory types, spec sheet, memory budgets, nvidia-smi cheat sheet
  • RESOURCES.md — papers, docs, blogs, talks, organized by section
  • PROGRESS.md — your checklist

Currency#

The mechanisms in this curriculum are stable; product facts are not. Hardware tables, engine recommendations and project statuses were last reviewed on 3 October 2026. What that review changed:

  • Hardware. Rubin (HBM4, ~22 TB/s) has shipped since August 2026 and Blackwell Ultra (B300) since 2025; both are in GPU-HARDWARE.md and XIV.09.
  • Engines. Hugging Face archived TGI in March 2026; VIII.13 and VIII.14 now treat it as a migration source. vLLM and SGLang remain the reference engines.
  • Platform. Kubernetes DRA and the Gateway API Inference Extension moved from “emerging” to stable APIs; llm-d and NVIDIA Dynamo are the open-source layers above the engine.
  • Newer papers and feeds are under “2025-2026 additions” in RESOURCES.md.

Sister paths#

  • Golang Engineering — the language this curriculum’s code is written in: memory, the runtime, concurrency, and AI building blocks in Go.
  • GPU Engineering — the device underneath all of this.
  • Observability Engineering — measuring it in production: SLOs, GPU and engine telemetry, tracing, cost per token.

The one idea to carry through everything#

If you remember nothing else from this repository, remember this:

Modern LLM inference is not limited by how fast your GPU can multiply. It is limited by how fast your GPU can move bytes.

Almost every technique in sections V, VII, IX, X, and XIII is, at bottom, a scheme to move fewer bytes or to do more arithmetic per byte moved. Once you see that, the field stops being a list of tricks and becomes a single coherent story.

Start with I-fundamentals/.

↑↓ navigate↵ openesc close