1. What is it?#
Inference is running a trained model to get an answer.
That’s it. You have a model — a large collection of numbers that somebody spent money to produce — and you feed it an input, and it produces an output. That act is inference.
input ──────► [ model ] ──────► output
"translate: billions "bonjour"
hello" of numbersThe word comes from statistics: you are inferring an answer from evidence. In engineering practice it just means “the forward direction of the model, in production, on real requests.”
Every time you type into a chat assistant, get a product recommendation, have a photo auto-tagged, or see a fraud score on a transaction — that’s inference. It is the part of machine learning that users actually touch, and it is where the vast majority of the compute spend of a deployed ML system ends up over the system’s lifetime.
Inference engineering is the discipline of making that act fast, cheap, and reliable at scale. It is the subject of this entire repository.
2. Why does it exist?#
A trained model is useless sitting in a file. The value is realized only when it answers questions. So somebody has to:
- load tens of gigabytes of numbers onto hardware that can do arithmetic quickly,
- accept requests from many users at once,
- run the arithmetic,
- return answers within a time budget users will tolerate,
- do this without spending more money than the answers are worth.
Each bullet is a whole engineering problem. Historically, “inference” was an afterthought — you trained a model and then someone wrapped it in a Flask app. That worked when models were 50 MB and a request took 5 ms.
It stopped working when models became 140 GB and a single request could occupy an $30,000 GPU for thirty seconds. At that point, the difference between a naive deployment and a good one is the difference between $2.00 and $0.05 per million tokens — a 40x cost difference on the same hardware, running the same model, producing the same answers.
That gap is why inference engineering exists as a specialty.
3. Simple analogy#
A model is a factory. Training built the factory. Inference is running the production line.
- Building the factory (training) is a one-time, enormous, capital-intensive project. You do it once, with a huge team, and you can tolerate it taking weeks.
- Running the line (inference) happens millions of times a day, forever. Every second of inefficiency is multiplied by every unit produced. A 10% improvement to the factory construction schedule saves you once. A 10% improvement to the production line saves you every day for the life of the factory.
This is why enormous engineering effort goes into inference optimizations that would look absurdly micro-optimized in ordinary software: shaving 200 microseconds off a kernel matters when that kernel runs 10 billion times a day.
A second analogy, useful later: a model is a very large lookup table that you have to compute rather than store. It “contains” answers to more questions than could ever be enumerated, and the price you pay for that compression is that retrieving any single answer requires a large amount of arithmetic.
4. Tiny example#
Here is inference, complete, with no ML libraries. This model predicts a house price from its size:
package main
import "fmt"
// The "trained model": two learned numbers.
const (
weight = 150.0 // dollars per square foot
bias = 20000.0 // base price
)
func infer(squareFeet float64) float64 {
return weight*squareFeet + bias
}
func main() {
fmt.Println(infer(1500)) // 245000
}That is a genuine, if tiny, model. It has 2 parameters. Inference is one multiply and one add.
Now scale the same idea:
| Model | Parameters | Arithmetic per token/request |
|---|---|---|
| House price above | 2 | ~2 FLOPs |
| Small image classifier | 5 million | ~10 million FLOPs |
| Llama 3 8B | 8 billion | ~16 billion FLOPs per token |
| Llama 3 70B | 70 billion | ~140 billion FLOPs per token |
Nothing conceptually changed. output = f(input, parameters). Only the size changed — by ten
orders of magnitude. Every hard problem in this curriculum comes from that scale, not from
any additional conceptual complexity. Hold onto that; it will keep you oriented when things
get dense in Section V.
A useful rule of thumb you will meet again in file 08:
FLOPs per token ≈ 2 × number of parametersThe 2 is because each parameter participates in one multiply and one add. Check it: 8B params → 16 GFLOPs per token. That matches the table.
5. Technical explanation#
Formally, inference is evaluating a function
y = f(x; θ)where:
xis the input (a prompt, an image, a feature vector),θ(theta) is the set of learned parameters — the model weights, fixed at inference time,fis the model architecture — the fixed sequence of operations,yis the output (a token, a class, a score).
Three properties make this different from evaluating a normal function:
1. θ is huge and must be resident. For a 70B model in 16-bit precision, θ occupies 140 GB.
You cannot read it from disk per request. It has to live in fast memory attached to your
compute device, which immediately constrains what hardware you can use and how many models you
can host per machine.
2. f is a fixed dataflow graph, not arbitrary control flow. There are no unbounded loops,
no I/O in the middle, no branching on data in the way normal programs branch. This is
wonderful for optimization: because the shape of the computation is known in advance, you can
plan memory, fuse operations, and compile the whole thing. Much of Sections IV and VII exploits
exactly this property.
3. For generative models, f is applied repeatedly, feeding its own output back in. This is
autoregression, and it is the source of most of the difficulty in Section V. One user request
is not one evaluation of f — it is hundreds or thousands of sequential evaluations, each of
which must wait for the previous one.
That third point deserves emphasis now, because it is the fork in the road for the whole field:
Non-generative inference (image classifier, ranker, embedder):
request ──► one forward pass ──► answer [ a function call ]
Generative inference (LLM):
request ──► forward ──► token 1
──► forward ──► token 2
──► forward ──► token 3
... [ a loop, with state ]
──► forward ──► token 500 ──► answerThe left-hand case is a normal, if expensive, RPC. The right-hand case is a stateful, long-lived, variable-duration session masquerading as an RPC. Almost everything strange about LLM serving follows from that mismatch.
6. Under the hood#
Follow a single request through a real system. Details will be filled in over the next several sections; for now, just see the shape:
1. HTTP request arrives at a gateway
"prompt": "Explain gravity"
2. Auth, rate limit, routing decisions
3. Tokenization: text → integers
"Explain gravity" → [50, 8161, 15883]
4. Request enters a queue
5. Scheduler picks it up, groups it with other waiting requests
6. Tokens are copied CPU → GPU over PCIe
7. PREFILL: the whole prompt runs through all layers at once
→ produces the first output token
→ and fills the KV cache (Section V)
8. DECODE loop, once per output token:
- run all layers for ONE new token
- append to KV cache
- sample the next token
- stream it to the client
- repeat until stop condition
9. Detokenize, finish the HTTP stream, free the KV cache memorySteps 7 and 8 are where 95%+ of the compute goes, and they have completely different performance characteristics — one is compute-bound, one is memory-bound. That split (file V.03) is the central fact of LLM serving.
What is physically happening in steps 7-8: billions of numbers are being read out of GPU memory, multiplied and added in enormous parallel arrays of arithmetic units, and the results written back. The GPU is not “thinking.” It is doing linear algebra at a rate of ~10^15 operations per second, and the interesting engineering question is almost always whether the arithmetic units are actually busy or are sitting idle waiting for data.
7. Performance implications#
Four numbers characterize any inference workload. Learn to ask for all four:
| Metric | Meaning | Typical target |
|---|---|---|
| TTFT — time to first token | How long until the user sees anything | 200 ms – 1 s |
| ITL — inter-token latency | Gap between successive tokens | 10 – 50 ms |
| Throughput | Tokens/sec across all concurrent users | as high as possible |
| Cost | $ per million tokens | as low as possible |
These are in tension. Section I.05 and I.06 dissect the tensions. The short version:
- Bigger batches → more throughput, lower cost, worse latency for the individual user.
- More GPUs per model → lower latency, worse efficiency (communication overhead).
- Lower precision → faster and cheaper, possible quality loss.
There is no configuration that wins on all axes. Inference engineering is the practice of choosing your losses deliberately rather than accidentally.
8. Production implications#
Some consequences that surprise people arriving from ordinary backend work:
Requests are not fungible. In a normal service, requests take roughly similar time. In LLM serving, one request might generate 5 tokens and another 4,000 — an 800x difference in cost, with no way to know in advance. Every queueing, load-balancing, and autoscaling assumption you have needs re-examination.
Capacity is memory-shaped, not CPU-shaped. You do not run out of “CPU”; you run out of GPU memory for KV cache, usually abruptly, usually under load. Section V.06 shows the arithmetic.
Cold start is measured in minutes. Loading 140 GB of weights takes real time. You cannot scale out in response to a traffic spike the way you can with stateless web servers. This single fact reshapes capacity planning (Section XI).
Utilization is the whole game economically. A GPU costs the same whether it is 5% busy or 95% busy. The entire economics of an inference platform is “keep the expensive silicon busy with useful work.”
9. Common mistakes#
“Inference is just training’s forward pass.” It shares the math but almost nothing else. Training does forward + backward, runs on fixed-size batches you control, is throughput-only, and tolerates minutes of latency. Inference does forward only, on adversarially variable inputs you don’t control, under latency SLOs, with a live user waiting. See file 02.
“We’ll optimize inference later.” Inference cost is the dominant lifetime cost of a deployed model, and some decisions (attention variant, hidden size, vocabulary size, context length) are baked in at training time and cannot be optimized away later. Inference constraints should influence model design. Section XIV covers hardware-aware model design.
“The GPU is at 100% so we’re done.” nvidia-smi’s utilization number means “a kernel was
resident,” not “the chip was doing useful work.” A kernel stalled on memory shows 100%. See
Section X.04 — this one misconception has probably wasted more GPU-hours than any other.
Optimizing throughput when the user cares about latency (or vice versa). Decide which one your product needs before you tune anything. A batch-embedding job and an interactive chat assistant want opposite configurations of the same engine.
Confusing FLOPs with FLOP/s. FLOPs = a count of operations. FLOP/s = a rate. “This model is 16 GFLOPs per token” and “this GPU is 1000 TFLOP/s” are different kinds of number, and dividing them gives you a time. Getting this confused makes all your estimates nonsense.
10. Hands-on exercise#
Part A — build the world’s smallest inference server.
// infer.go — run with: go run infer.go
package main
import (
"fmt"
"time"
)
const weight, bias = 150.0, 20000.0
func infer(sqft float64) float64 { return weight*sqft + bias }
func main() {
const n = 100_000_000
var sink float64 // use every result so the compiler cannot delete the loop
t0 := time.Now()
for i := 0; i < n; i++ {
sink += infer(float64(i % 3000))
}
dt := time.Since(t0)
fmt.Printf("%.0f inferences/sec, %.2f ns each (checksum %.0f)\n",
n/dt.Seconds(), float64(dt.Nanoseconds())/n, sink)
}Run it. A compiled language gets close to a billion inferences/sec on one core.
Now answer:
- This model does 2 FLOPs per inference. What is your measured FLOP/s?
- Your CPU is capable of roughly 10^11 FLOP/s. Why are you nowhere near it?
- Change
inferto take a[]float64of 1,000,000 inputs and fill a[]float64of outputs in one call. Then split that slice acrossruntime.NumCPU()goroutines. Measure both. How much faster, and why?
(The answer to #3 is a preview of the entire curriculum: you removed per-item overhead and let the hardware do many operations at once. Batching and vectorization are the same idea appearing at different scales.)
Part B — estimation practice. Without a computer:
- A 13B parameter model in FP16. How many GB just for weights?
- If your GPU has 80 GB, how many copies of that model fit? What else needs to fit?
- At 2 FLOPs per parameter per token, how many FLOPs to generate one token?
- If your GPU sustains 300 TFLOP/s, what is the theoretical minimum time per token?
- Real systems get nowhere near that at batch size 1. Guess why. (Then read file 08.)
Part C — journal. Start a numbers.md file. Record: your CPU model, the ns/inference from
Part A, and your batched and multi-goroutine speedups. You will keep adding to this file for the rest of the
curriculum.
11. Interview questions#
- Explain inference to a backend engineer who has never touched ML, in 60 seconds.
- What are the four metrics you would put on a dashboard for an LLM inference service, and why those four?
- Why is inference cost usually larger than training cost over a model’s lifetime? When is that not true?
- A colleague says “we’ll just autoscale like any other service.” What do you tell them?
- Given a model with P parameters, estimate FLOPs per token and memory for weights. Now do it for P = 70B at FP16 and at INT4.
- Name three things that make a generative model harder to serve than an image classifier.
12. Further reading#
- [FUNDAMENTAL] Kaplan et al., “Scaling Laws for Neural Language Models” (2020) — read the appendix on FLOPs accounting, ignore the rest for now.
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022) — you are not ready for all of it yet; read section 2 and come back after Section V.
- [REFERENCE] vLLM docs, “Architecture Overview” — skim to see the shape of a real system.
- Next: 02 — Training vs inference