This file teaches the mechanism. Section VII covers the specific production methods (GPTQ, AWQ, SmoothQuant, FP8) and when to use each.
1. What is it?#
Representing a tensor with fewer bits by mapping its range onto a small set of levels.
Original (FP16): [-0.83, 0.12, 1.47, -1.92, 0.55, ...]
Quantized (INT8): [-55, 8, 98, -128, 37, ...] + scale = 0.015
Reconstructed: [-0.825, 0.120, 1.470, -1.920, 0.555, ...]You store the small integers and one (or a few) scaling factors. The values you get back are approximations — the quantization error is the price.
2. Why does it exist?#
Because decode is memory-bandwidth-bound (Section I.07), and bytes moved is the binding constraint. Halving the bytes halves the time. Everything else about quantization is a consequence of pursuing that.
Secondary benefits: the model fits in less memory (more KV cache, cheaper GPUs), and on hardware with low-precision tensor cores, compute throughput increases too.
3. Simple analogy#
Rounding prices to the nearest dollar.
$4.37, $12.81, $0.99 → $4, $13, $1. You lose cents but the list is shorter to write and faster to add up.
The engineering questions are exactly the ones quantization asks:
- What’s the range? If prices go from $0.99 to $10,000,000, rounding to dollars is fine for the big ones and destroys the small ones.
- Should I use one rounding scheme for everything, or per category? (per-tensor vs per-channel)
- What about the one item priced at $50 million? (outliers — the central problem)
4. Tiny example#
Symmetric INT8 quantization, worked completely:
package main
import (
"fmt"
"math"
)
func main() {
w := []float64{-0.83, 0.12, 1.47, -1.92, 0.55, 0.03, -0.41, 0.88}
// 1. Find the scale
var absmax float64
for _, v := range w {
absmax = math.Max(absmax, math.Abs(v)) // 1.92
}
scale := absmax / 127 // 0.01512
fmt.Printf("scale: %.5f\n", scale)
// 2. Quantize
q := make([]int8, len(w))
for i, v := range w {
q[i] = int8(math.Round(v / scale))
}
fmt.Println("quantized:", q) // [-55 8 97 -127 36 2 -27 58]
// 3. Dequantize, 4. Error
var maxErr, maxRel float64
fmt.Print("recovered:")
for i, v := range w {
wHat := float64(q[i]) * scale
fmt.Printf(" %.4f", wHat) // -0.8315 0.1209 1.4665 -1.9200 0.5443 ...
err := math.Abs(v - wHat)
maxErr, maxRel = math.Max(maxErr, err), math.Max(maxRel, err/math.Abs(v))
}
fmt.Printf("\nmax error: %.4f (scale/2 = %.4f)\nrelative: %.3f\n", maxErr, scale/2, maxRel)
}Maximum error is always scale/2 for round-to-nearest. That’s the key fact: absolute error
is uniform, so relative error is terrible for small values and fine for large ones.
Now the outlier problem:
w2 := []float64{-0.83, 0.12, 1.47, -1.92, 0.55, 0.03, -0.41, 45.0} // one outlier
scale2 := 45.0 / 127 // 0.354 ← 23x larger!
for _, v := range w2 {
q := math.Round(v / scale2)
fmt.Printf("%4.0f -> %6.3f\n", q, q*scale2)
}
// q: -2 0 4 -5 2 0 -1 127
// recovered: -0.709 0 1.417 -1.772 0.709 0 -0.354 45.0Look at what happened. 0.12 became 0.0. 0.03 became 0.0. 0.55 became 0.708. One
outlier destroyed the precision of every other value. Outlier handling is the central problem
of quantization, and every named method (LLM.int8, SmoothQuant, AWQ, GPTQ) is a different
answer to it.
5. Technical explanation#
Symmetric vs asymmetric#
Symmetric: x ≈ q · s zero maps to zero
s = absmax / (2^(b-1) - 1)
simpler kernels, wastes half the range for one-sided distributions
Asymmetric: x ≈ (q - z) · s z = zero-point offset
s = (max - min) / (2^b - 1)
z = round(-min / s)
better for one-sided data (e.g. post-ReLU activations)
extra add in the kernelWeights are roughly zero-centered, so symmetric is standard for weights. Activations after ReLU/SiLU can be one-sided, so asymmetric is sometimes used there.
Granularity — the most important knob#
Per-tensor: one scale for the whole matrix smallest overhead, worst accuracy
Per-channel: one scale per output row standard for weights
Per-group: one scale per group of 32/64/128 values standard for INT4
Per-token: one scale per token (activations) standard for dynamic activation quantThe overhead:
INT4 weights with group size 128, FP16 scales:
4 bits per weight + 16 bits per 128 weights = 4 + 0.125 = 4.125 bits/weight
→ effective 4.125/16 = 25.8% of FP16. The scales cost ~3%.Finer granularity → better accuracy → slightly more memory and slightly more kernel work. Group size 128 is the near-universal default for INT4 because it captures most of the accuracy benefit at small cost.
Weight-only vs weight-and-activation#
WEIGHT-ONLY (W4A16, W8A16):
store weights in INT4/INT8
dequantize to FP16 in the kernel, compute in FP16
✓ big memory/bandwidth saving → helps DECODE a lot
✗ no compute saving → helps PREFILL almost none
✓ easy: no activation calibration needed
→ GPTQ, AWQ
WEIGHT AND ACTIVATION (W8A8, FP8):
both quantized; compute in INT8/FP8 tensor cores
✓ memory saving AND ~2x compute
✗ harder: activations have per-input outliers
✗ needs calibration
→ SmoothQuant, FP8 with per-tensor scalesChoose by phase: decode-heavy workloads → weight-only INT4 is excellent. Prefill-heavy (RAG, long documents) → you need W8A8/FP8 to get compute benefits.
Static vs dynamic activation quantization#
Static: scales computed offline from a calibration set. Fastest; risky if
production data has a different distribution.
Dynamic: scales computed per batch at runtime (a max-reduction over the tensor).
Costs a pass over the data; robust. Standard for per-token quantization.Where the error comes from#
For a matmul Y = X W:
Ŷ = X̂ Ŵ = (X + εx)(W + εw) ≈ XW + X·εw + εx·W
error ≈ ||X||·||εw|| + ||εx||·||W||Two consequences:
- Error is proportional to the magnitude of the other operand. Quantizing a weight column that multiplies large activations hurts more. This is AWQ’s entire insight — protect the weight channels that see large activations.
- Errors accumulate through layers, and through generated tokens.
What to quantize and what to leave alone#
Quantize aggressively: FFN weights (65-70% of parameters)
Quantize carefully: attention projections
Usually leave in FP16: embeddings, LM head, norm weights, biases
Quantize separately: KV cache (Section VII.12) — different tradeoffsThe LM head is often left alone: it’s a single large matrix whose errors directly perturb the sampling distribution.
6. Under the hood#
A W4A16 GEMM kernel:
for each tile:
load INT4 weights from HBM ← 4x less traffic than FP16
load FP16 scales for the groups
UNPACK: two 4-bit values per byte → separate values
DEQUANTIZE: w_fp16 = (int4_val - zero) * scale
load FP16 activations
tensor core MMA in FP16
accumulate FP32The unpack+dequantize is extra work, done in registers. It costs FLOPs — but decode has FLOPs to spare (Section I.07), so trading compute for bandwidth is exactly the right trade. In prefill, where you’re compute-bound, that same extra work is a pure cost, which is the mechanistic reason W4A16 doesn’t help prefill.
7. Performance implications#
Measured, typical, for a 70B model on H100:
| Scheme | Memory | Decode speedup | Prefill speedup | Quality (typical) |
|---|---|---|---|---|
| BF16 | 140 GB | 1.0x | 1.0x | reference |
| FP8 (W8A8) | 70 GB | 1.8-2.0x | 1.6-1.9x | -0.1 to -0.5% |
| INT8 (W8A8) | 70 GB | 1.7-1.9x | 1.5-1.8x | -0.2 to -1% |
| INT8 weight-only | 70 GB | 1.7-1.9x | ~1.0x | -0.1 to -0.5% |
| INT4 (W4A16, GPTQ/AWQ) | 35 GB | 2.5-3.5x | 0.9-1.1x | -1 to -3% |
| INT4 group=32 | 37 GB | 2.4-3.3x | 0.9-1.1x | -0.5 to -2% |
Note INT4 can be slower than FP16 at prefill: the dequantization overhead with no compute benefit. Some engines automatically use different paths for prefill and decode for this reason.
8. Production implications#
- Quantize offline, ship the quantized checkpoint. Never quantize at startup (Section II.07).
- Evaluate on your own task, with long generations. Public benchmark deltas do not transfer.
- Record the exact quantization config in your registry: method, bits, group size, symmetric/asymmetric, which layers excluded, calibration dataset.
- Prefer FP8 on Hopper/Blackwell if quality permits — it’s simpler than INT8 (no zero points, wide dynamic range) and gets compute benefits.
- Watch the prefill regression. If your workload is RAG-shaped, W4A16 may make things worse.
- A/B in production. Quality regressions from quantization are exactly the kind that offline evals miss.
9. Common mistakes#
Assuming quantization is free. It costs quality. Measure it.
Using per-tensor scales for INT4. Accuracy collapses. Use group-wise.
Quantizing embeddings and the LM head with everything else. Disproportionate quality cost for little memory benefit.
Expecting prefill speedup from weight-only quantization. It doesn’t work that way.
Evaluating with perplexity only. Perplexity is insensitive to the failure modes users notice (reasoning, instruction following, long-form coherence).
Ignoring calibration data distribution. Calibrating on Wikipedia and serving code produces worse results than calibrating on code.
Forgetting that KV cache quantization is a separate decision with its own tradeoffs.
10. Hands-on exercise#
A. Implement it. Write symmetric INT8 quantize/dequantize with per-tensor, per-channel, and per-group (128) granularity. For a real weight matrix from a model, measure mean and max relative error for each. Plot error vs group size.
B. The outlier experiment. Take a real activation tensor from a model (hook a forward pass). Plot its distribution. Find the outliers. Quantize with and without clipping the top 0.1% and compare error on the non-outlier values.
C. End-to-end quality. Quantize a small model to INT8 and INT4 (use bitsandbytes,
gptq, or autoawq). Measure perplexity on a held-out set AND generate 20 long responses,
comparing them qualitatively to FP16.
D. Phase asymmetry. Measure prefill and decode throughput for FP16 and INT4 weight-only on the same model. Confirm decode improves and prefill doesn’t. Explain using arithmetic intensity.
E. Scale overhead. Compute the exact bits-per-weight for INT4 with group sizes 32, 64, 128, and per-tensor. Plot bits/weight vs quality error from exercise A.
11. Interview questions#
- Explain symmetric INT8 quantization, including how the scale is computed.
- What is the outlier problem and why does it make activation quantization harder than weight quantization?
- Compare per-tensor, per-channel, and per-group quantization.
- Why does W4A16 speed up decode but not prefill? Give the mechanism.
- What is the effective bits-per-weight for INT4 with group size 128?
- Which parts of a model would you leave unquantized, and why?
- How would you validate a quantized model before deploying it?
12. Further reading#
- [ESTABLISHED] Dettmers et al., “LLM.int8()” (2022) — the outlier discovery
- [ESTABLISHED] Frantar et al., “GPTQ” (2022)
- [ESTABLISHED] Lin et al., “AWQ” (2023)
- [ESTABLISHED] Xiao et al., “SmoothQuant” (2022)
- [REFERENCE]
bitsandbytes,autoawq,llm-compressordocumentation - Next: Section IV — Neural Network Inference