PidokuInfra
01Fundamentals
On this page

Fundamentals

FoundationsModule 0111 topics~11h 30m

Topics, in order

01 What Is Inference?Inference is running a trained model to get an answer. Beginner 45 min 02 Training vs InferenceTraining and inference are the two phases of a model's life. Beginner 45 min 03 Model Anatomy: Parameters, Weights, ActivationsOpen a model file and you find three categories of numbers. Two live on disk; one exists only while the model runs. Beginner 1h 04 Tensors and the Forward PassA tensor is an n-dimensional array of numbers. That is the whole definition. Physicists mean something more specific; in ML, "tensor" means "array with a shape." Beginner 1h 05 Latency, Throughput, and the Metrics That MatterLatency is how long one thing takes. Throughput is how many things you finish per unit time. They are different, they are often in conflict, and for LLMs you need at least four numbers, not … Beginner 1h 15m 06 Batching: The Central TradeoffBatching means processing several requests in one pass through the model, so the weights are read from memory once and used many times. Beginner 1h 07 Compute vs MemoryEvery computation has two costs: doing the arithmetic and moving the data to where the arithmetic happens. One of them takes longer. Whichever it is, that is your bottleneck, and optimizing … Intermediate 1h 15m 08 FLOPs, Bandwidth, and Arithmetic IntensityThis file is the quantitative heart of Section I. Everything here is arithmetic you should be able to do on a whiteboard. Do not skim it. Intermediate 1h 30m 09 CPU vs GPUTwo fundamentally different bets about what a processor is for. Beginner 1h 10 Why Inference Is Not Normal Backend EngineeringEvery instinct a good backend engineer has — stateless services, horizontal scaling, uniform request cost, retries are cheap, autoscale on CPU — is either wrong or dangerously incomplete for … Intermediate 1h 11 Why Inference Gets Expensive at ScaleThe economics of inference: where the money goes, why costs grow the way they do, and which engineering levers actually move the number. Intermediate 1h

About this module

Goal of this section: give you the vocabulary and the physics of inference. By the end you will be able to read any inference discussion without getting lost, and — more importantly — you will be able to estimate whether a proposed system can possibly work, using nothing but arithmetic.

This section is deliberately light on ML. We are not yet asking “what is a transformer.” We are asking “what does it mean to run any trained model, and what makes that hard?”

Files#

#FileLevelTime
01What is inference?Beginner45 min
02Training vs inferenceBeginner45 min
03Model anatomy: parameters, weights, activationsBeginner60 min
04Tensors and the forward passBeginner60 min
05Latency, throughput, and the metrics that matterBeginner75 min
06Batching: the central tradeoffBeginner60 min
07Compute vs memoryIntermediate75 min
08FLOPs, bandwidth, arithmetic intensityIntermediate90 min
09CPU vs GPUBeginner60 min
10Why inference is not normal backend engineeringIntermediate60 min
11Why inference gets expensive at scaleIntermediate60 min

The thread running through this section#

flowchart TD
  N0["A model is a big pile of numbers<br/><b>03</b>"]
  N1["Running it means moving those numbers through arithmetic<br/><b>04</b>"]
  N2["Moving numbers costs time; arithmetic costs time<br/><b>07, 08</b>"]
  N3["Which one dominates depends on how much arithmetic per byte you do<br/><b>08</b>"]
  N4["Batching is the main lever on that ratio<br/><b>06</b>"]
  N5["Which is why every serving system is, underneath, a batching scheduler<br/><b>10, 11</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5

  class N0,N1 neutral
  class N2 io
  class N3 queue
  class N4 compute
  class N5 memory

Checkpoint A#

Before moving to Section II, you should be able to answer, without looking anything up:

  1. Latency vs throughput — define both and give an example where improving one hurts the other.
  2. A 7B model in FP16: how much memory for weights? Show the arithmetic.
  3. Why is generating token #500 no more expensive in FLOPs than token #5, but more expensive in memory traffic?
  4. What is arithmetic intensity, and why is batch size the main lever on it?
  5. Why does a GPU beat a CPU at inference — is it the FLOPs, the bandwidth, or both?

↑↓ navigate↵ openesc close