1. Problem → Why → Optimization#
PROBLEM
Decode reads every weight from HBM for every token. A 70B FP16 model
reads 140 GB per step → 42 ms floor on an H100 → 24 tokens/sec.
WHY IT HAPPENS
Arithmetic intensity of GEMV is ~1 FLOP/byte against a ridge point of 296.
The GPU is 99.7% idle, waiting for bytes.
OPTIMIZATION
Store weights (and/or activations, and/or KV cache) in fewer bits.
Fewer bytes → proportionally faster decode.
HOW IT WORKS
Map a tensor's value range onto a small set of levels, store the level
indices plus a scale, reconstruct on the fly in the kernel.
TRADE-OFFS
↓ memory, ↓ bandwidth, sometimes ↑ compute throughput
↑ quantization error → quality loss
↑ implementation complexity, calibration requirement
WHEN TO USE
Almost always for production LLM serving. FP8/INT8 is near-free.
INT4 when decode-bound and quality budget allows.
WHEN NOT TO USE
- Prefill-dominated workloads with weight-only quantization (no benefit)
- When you haven't validated quality on YOUR task
- Models already at the edge of acceptable quality
- When the engineering time is better spent on scheduling (see 01)Diagram — Which kind of quantization?#
flowchart TB
S{"What limits you?"}
S -->|"decode speed or memory at small batch"| W["Weight-only quantization"]
S -->|"prefill or high-batch throughput"| A["Weights + activations<br/>FP8 or INT8"]
S -->|"concurrency - KV memory"| K["KV cache quantization"]
W --> W8{"How aggressive?"}
W8 -->|"8-bit"| I8["INT8 / FP8<br/>near-lossless, round-to-nearest is fine"]
W8 -->|"4-bit"| I4["INT4 with GPTQ / AWQ<br/>calibration required"]
I8 --> E["Evaluate on YOUR task,<br/>not perplexity alone"]
I4 --> E
A --> E
K --> E
class S,W8 queue
class W,K,I8,I4 memory
class A compute
class E neutral2. The decision guide#
This is the practical core of the file.
START: what is your workload shape?
├─ DECODE-HEAVY (chat, creative, agents; output >> input)
│ ├─ Hopper/Blackwell GPU?
│ │ ├─ YES → FP8 (W8A8). 1.8-2x, minimal quality loss. ← START HERE
│ │ └─ NO → INT8 weight-only, or INT4 (AWQ/GPTQ) if quality allows
│ └─ Need more? → INT4 W4A16 (2.5-3.5x decode) + validate carefully
│
├─ PREFILL-HEAVY (RAG, summarization, long documents)
│ ├─ Weight-only quantization gives ~NOTHING. Do not bother.
│ ├─ Hopper+ → FP8 W8A8 (compute AND memory benefit)
│ └─ Otherwise → INT8 W8A8 with SmoothQuant
│
├─ BALANCED
│ └─ FP8 if available, else INT8 W8A8
│
└─ MEMORY-CONSTRAINED (model doesn't fit)
└─ INT4 weight-only. Fitting at all beats everything.
THEN, separately:
├─ Long context (>16k) and high batch?
│ └─ Also quantize the KV cache to FP8 (file 12). Often a bigger win
│ than further weight quantization.3. Simple analogy#
Shipping by weight. Your freight cost is per kilogram, and your cargo is mostly packaging. Quantization is repacking into smaller boxes: same goods, less weight, lower cost — up to the point where the goods get damaged.
The interesting part is that different goods tolerate different amounts of compression. Attention projections are fragile; FFN weights are robust; embeddings and the LM head are somewhere in between. Good quantization exploits this.
4. Tiny example — the four decisions#
Any quantization scheme is defined by four choices:
1. WHAT weights only? + activations? + KV cache?
2. HOW MANY BITS 8, 6, 4, 3, 2?
3. GRANULARITY per-tensor, per-channel, per-group(32/64/128)?
4. STATIC OR DYNAMIC scales computed offline or at runtime?The named methods are just points in this space:
Method What Bits Granularity Scales
────────────────────────────────────────────────────────────────
GPTQ W 4/3 group 128 static, Hessian-optimized
AWQ W 4 group 128 static, activation-aware
SmoothQuant W + A 8 per-channel W, static/dynamic
per-token A
FP8 (TE) W + A 8 per-tensor static or dynamic
LLM.int8() W + A 8 per-channel dynamic, outliers in FP16
bitsandbytes W 4/8 group 64 dynamic (NF4)
KV quant KV cache 8/4 per-token/head dynamicOnce you see the four axes, the zoo of method names stops being intimidating.
5. Technical explanation#
The gain, by phase and scheme#
Decode gain Prefill gain Memory Quality Δ
BF16 (baseline) 1.00x 1.00x 1.00x —
FP8 W8A8 (Hopper) 1.8-2.0x 1.6-1.9x 0.50x -0.1 to -0.5%
INT8 W8A8 (SmoothQuant) 1.7-1.9x 1.5-1.8x 0.50x -0.2 to -1.0%
INT8 weight-only 1.7-1.9x ~1.0x 0.50x -0.1 to -0.5%
INT4 W4A16 (AWQ/GPTQ, g128) 2.5-3.5x 0.9-1.1x 0.28x -1 to -3%
INT4 g32 2.4-3.3x 0.9-1.1x 0.30x -0.5 to -2%
INT3 3.0-4.0x 0.9x 0.22x -3 to -8%
INT2 — — 0.16x usually unusable
FP8 KV cache varies* — KV 0.5x -0.1 to -0.5%*KV quantization’s gain depends entirely on the KV/weight byte ratio (Section V.06). At long context and high batch it can exceed the weight-quantization gain.
Why the prefill column looks like that#
Weight-only (W4A16):
Weights are dequantized to FP16 inside the kernel; the MATH is still FP16.
Prefill is compute-bound → no gain from fewer weight bytes.
In fact, the dequantization adds work → can be slightly SLOWER.
W8A8 (FP8/INT8):
Both operands are low precision → tensor cores run at 2x.
Prefill IS faster.This is the single most important thing to understand about quantization, and the most common misconception. It’s also why the decision tree branches on workload shape first.
What to quantize and what to leave#
QUANTIZE AGGRESSIVELY
FFN gate/up/down 65-70% of parameters, robust to quantization
QUANTIZE WITH CARE
Attention q/k/v/o more sensitive; errors propagate through attention
USUALLY LEAVE IN HIGHER PRECISION
Embeddings gather is cheap anyway; quality cost disproportionate
LM head errors directly perturb the sampling distribution
Norm weights tiny, and scale-critical
Router weights (MoE) a wrong routing decision is catastrophic; these are tiny
QUANTIZE SEPARATELY (different tradeoffs)
KV cache file 12The first and last layers are also often kept at higher precision — they’re disproportionately sensitive in most models.
Validating a quantized model#
Repeat from Section IV.12, because it matters:
L1 NUMERICAL max/mean logit error, KL divergence, top-1 agreement
target: KL < 0.01, top-1 > 99%
L2 TASK perplexity + benchmarks ON YOUR DOMAIN
L3 BEHAVIORAL 500+ token generations, instruction following, structured
output validity, refusal behavior
L4 PRODUCTION A/B test with human or LLM-judge evaluationL3 is where INT4 usually shows its cost and where teams that stopped at L2 get surprised.
6. Under the hood: where the speedup comes from#
W4A16 decode GEMM, per weight tile:
HBM → registers: 4 bits/weight ← 4x less traffic than FP16
registers: unpack 2 weights per byte
dequantize: w = (q - z) * s (FP16 arithmetic)
tensor core: FP16 MMA
The unpack+dequant costs ~10-20 extra FLOPs per weight. Decode has FLOPs
to spare (intensity 1 vs ridge 296), so this is free.
In PREFILL, the same extra FLOPs are NOT free — you're already compute-bound.Production W4A16 kernels (Marlin, Machete, exllamav2) go further: they pre-permute weights at load time so the unpacking is branch-free and coalesced, and they fuse the dequantization into the tensor-core feeding path. These achieve 3-4x over FP16 at small batch, versus ~1.5x for a naive dequantize-then-GEMM.
7-9. Performance, production, mistakes#
Performance: the table in section 5. Note that the gains are not multiplicative with batching gains — quantization reduces the weight-read term, which batching was already amortizing. At batch 256, weights are a smaller share of bytes, so quantization helps less.
Production:
- Quantize offline, ship the quantized checkpoint. Never at startup.
- Record the exact config in your registry: method, bits, group size, excluded layers, calibration dataset, library version.
- Use calibration data resembling your traffic. Calibrating on Wikipedia and serving code is measurably worse.
- A/B in production. Offline evals miss things.
- Have a rollback plan. Keep the FP16 checkpoint deployable.
- Check that your engine has an optimized kernel for the scheme you chose. A W4A16 model on a naive dequant path can be slower than FP16.
Mistakes:
- Expecting prefill gains from weight-only quantization.
- Per-tensor scales for INT4. Accuracy collapse.
- Quantizing the LM head and embeddings with everything else.
- Validating with perplexity only.
- Assuming published benchmark deltas transfer to your task.
- Quantizing before fixing scheduling. A 2x on a system running at batch 4 is worth less than fixing the batch.
10. Hands-on exercise#
A. Run the decision tree. For three real workloads you know (or invent plausible ones), walk the decision tree and justify the recommendation.
B. Measure the phase asymmetry. Take a model, quantize it to INT4 weight-only. Measure prefill throughput and decode throughput before and after. Confirm decode improves ~2.5x and prefill doesn’t. Explain using arithmetic intensity.
C. Layer sensitivity. Quantize a model with (i) everything at INT4, (ii) everything except LM head and embeddings, (iii) FFN only. Measure perplexity and memory for each. Which gives the best quality-per-byte?
D. Calibration matters. Quantize the same model with calibration sets drawn from Wikipedia and from code. Evaluate both on code and on prose. Quantify the mismatch penalty.
E. Full validation. Take one quantized model through all four validation levels. Write up the results as you would for a deployment review.
11. Interview questions#
- Why does quantization speed up decode? Give the mechanism.
- Why does weight-only quantization not speed up prefill?
- What are the four axes that define a quantization scheme?
- Which layers would you leave unquantized and why?
- How would you choose between FP8 and INT4 for a given workload?
- How do you validate a quantized model before deployment?
- When does KV cache quantization matter more than weight quantization?
12. Further reading#
- [ESTABLISHED] Dettmers et al., “LLM.int8()” (2022)
- [ESTABLISHED] Frantar et al., “GPTQ” (2022); Lin et al., “AWQ” (2023)
- [ESTABLISHED] Xiao et al., “SmoothQuant” (2022)
- [REFERENCE]
llm-compressor,autoawq,bitsandbytesdocumentation - Next: 03 — GPTQ and AWQ