PidokuInfra

Quantization Overview and Decision Guide

Intermediate 1h 30m Difficulty 3/5 Topic 02 of 14

Prerequisites III.11, III.12


1. Problem → Why → Optimization#

PROBLEM
  Decode reads every weight from HBM for every token. A 70B FP16 model
  reads 140 GB per step → 42 ms floor on an H100 → 24 tokens/sec.

WHY IT HAPPENS
  Arithmetic intensity of GEMV is ~1 FLOP/byte against a ridge point of 296.
  The GPU is 99.7% idle, waiting for bytes.

OPTIMIZATION
  Store weights (and/or activations, and/or KV cache) in fewer bits.
  Fewer bytes → proportionally faster decode.

HOW IT WORKS
  Map a tensor's value range onto a small set of levels, store the level
  indices plus a scale, reconstruct on the fly in the kernel.

TRADE-OFFS
  ↓ memory, ↓ bandwidth, sometimes ↑ compute throughput
  ↑ quantization error → quality loss
  ↑ implementation complexity, calibration requirement

WHEN TO USE
  Almost always for production LLM serving. FP8/INT8 is near-free.
  INT4 when decode-bound and quality budget allows.

WHEN NOT TO USE
  - Prefill-dominated workloads with weight-only quantization (no benefit)
  - When you haven't validated quality on YOUR task
  - Models already at the edge of acceptable quality
  - When the engineering time is better spent on scheduling (see 01)

Diagram — Which kind of quantization?#

flowchart TB
  S{"What limits you?"}
  S -->|"decode speed or memory at small batch"| W["Weight-only quantization"]
  S -->|"prefill or high-batch throughput"| A["Weights + activations<br/>FP8 or INT8"]
  S -->|"concurrency - KV memory"| K["KV cache quantization"]
  W --> W8{"How aggressive?"}
  W8 -->|"8-bit"| I8["INT8 / FP8<br/>near-lossless, round-to-nearest is fine"]
  W8 -->|"4-bit"| I4["INT4 with GPTQ / AWQ<br/>calibration required"]
  I8 --> E["Evaluate on YOUR task,<br/>not perplexity alone"]
  I4 --> E
  A --> E
  K --> E

  class S,W8 queue
  class W,K,I8,I4 memory
  class A compute
  class E neutral

2. The decision guide#

This is the practical core of the file.

START: what is your workload shape?

├─ DECODE-HEAVY (chat, creative, agents; output >> input)
│   ├─ Hopper/Blackwell GPU?
│   │   ├─ YES → FP8 (W8A8).  1.8-2x, minimal quality loss.  ← START HERE
│   │   └─ NO  → INT8 weight-only, or INT4 (AWQ/GPTQ) if quality allows
│   └─ Need more? → INT4 W4A16 (2.5-3.5x decode) + validate carefully
│
├─ PREFILL-HEAVY (RAG, summarization, long documents)
│   ├─ Weight-only quantization gives ~NOTHING. Do not bother.
│   ├─ Hopper+ → FP8 W8A8 (compute AND memory benefit)
│   └─ Otherwise → INT8 W8A8 with SmoothQuant
│
├─ BALANCED
│   └─ FP8 if available, else INT8 W8A8
│
└─ MEMORY-CONSTRAINED (model doesn't fit)
    └─ INT4 weight-only. Fitting at all beats everything.

THEN, separately:
├─ Long context (>16k) and high batch?
│   └─ Also quantize the KV cache to FP8 (file 12). Often a bigger win
│      than further weight quantization.

3. Simple analogy#

Shipping by weight. Your freight cost is per kilogram, and your cargo is mostly packaging. Quantization is repacking into smaller boxes: same goods, less weight, lower cost — up to the point where the goods get damaged.

The interesting part is that different goods tolerate different amounts of compression. Attention projections are fragile; FFN weights are robust; embeddings and the LM head are somewhere in between. Good quantization exploits this.


4. Tiny example — the four decisions#

Any quantization scheme is defined by four choices:

1. WHAT      weights only? + activations? + KV cache?
2. HOW MANY BITS   8, 6, 4, 3, 2?
3. GRANULARITY     per-tensor, per-channel, per-group(32/64/128)?
4. STATIC OR DYNAMIC   scales computed offline or at runtime?

The named methods are just points in this space:

Method        What        Bits   Granularity      Scales
────────────────────────────────────────────────────────────────
GPTQ          W           4/3    group 128        static, Hessian-optimized
AWQ           W           4      group 128        static, activation-aware
SmoothQuant   W + A       8      per-channel W,   static/dynamic
                                 per-token A
FP8 (TE)      W + A       8      per-tensor       static or dynamic
LLM.int8()    W + A       8      per-channel      dynamic, outliers in FP16
bitsandbytes  W           4/8    group 64         dynamic (NF4)
KV quant      KV cache    8/4    per-token/head   dynamic

Once you see the four axes, the zoo of method names stops being intimidating.


5. Technical explanation#

The gain, by phase and scheme#

                          Decode gain    Prefill gain   Memory   Quality Δ
BF16 (baseline)              1.00x          1.00x        1.00x      —
FP8 W8A8 (Hopper)            1.8-2.0x       1.6-1.9x     0.50x   -0.1 to -0.5%
INT8 W8A8 (SmoothQuant)      1.7-1.9x       1.5-1.8x     0.50x   -0.2 to -1.0%
INT8 weight-only             1.7-1.9x       ~1.0x        0.50x   -0.1 to -0.5%
INT4 W4A16 (AWQ/GPTQ, g128)  2.5-3.5x       0.9-1.1x     0.28x   -1 to -3%
INT4 g32                     2.4-3.3x       0.9-1.1x     0.30x   -0.5 to -2%
INT3                         3.0-4.0x       0.9x         0.22x   -3 to -8%
INT2                         —              —            0.16x   usually unusable
FP8 KV cache                 varies*        —            KV 0.5x  -0.1 to -0.5%

*KV quantization’s gain depends entirely on the KV/weight byte ratio (Section V.06). At long context and high batch it can exceed the weight-quantization gain.

Why the prefill column looks like that#

Weight-only (W4A16):
  Weights are dequantized to FP16 inside the kernel; the MATH is still FP16.
  Prefill is compute-bound → no gain from fewer weight bytes.
  In fact, the dequantization adds work → can be slightly SLOWER.

W8A8 (FP8/INT8):
  Both operands are low precision → tensor cores run at 2x.
  Prefill IS faster.

This is the single most important thing to understand about quantization, and the most common misconception. It’s also why the decision tree branches on workload shape first.

What to quantize and what to leave#

QUANTIZE AGGRESSIVELY
  FFN gate/up/down       65-70% of parameters, robust to quantization
  
QUANTIZE WITH CARE
  Attention q/k/v/o      more sensitive; errors propagate through attention

USUALLY LEAVE IN HIGHER PRECISION
  Embeddings             gather is cheap anyway; quality cost disproportionate
  LM head                errors directly perturb the sampling distribution
  Norm weights           tiny, and scale-critical
  Router weights (MoE)   a wrong routing decision is catastrophic; these are tiny

QUANTIZE SEPARATELY (different tradeoffs)
  KV cache               file 12

The first and last layers are also often kept at higher precision — they’re disproportionately sensitive in most models.

Validating a quantized model#

Repeat from Section IV.12, because it matters:

L1 NUMERICAL   max/mean logit error, KL divergence, top-1 agreement
               target: KL < 0.01, top-1 > 99%
L2 TASK        perplexity + benchmarks ON YOUR DOMAIN
L3 BEHAVIORAL  500+ token generations, instruction following, structured
               output validity, refusal behavior
L4 PRODUCTION  A/B test with human or LLM-judge evaluation

L3 is where INT4 usually shows its cost and where teams that stopped at L2 get surprised.


6. Under the hood: where the speedup comes from#

W4A16 decode GEMM, per weight tile:

  HBM → registers:  4 bits/weight   ← 4x less traffic than FP16
  registers:        unpack 2 weights per byte
                    dequantize: w = (q - z) * s    (FP16 arithmetic)
  tensor core:      FP16 MMA
  
The unpack+dequant costs ~10-20 extra FLOPs per weight. Decode has FLOPs
to spare (intensity 1 vs ridge 296), so this is free.

In PREFILL, the same extra FLOPs are NOT free — you're already compute-bound.

Production W4A16 kernels (Marlin, Machete, exllamav2) go further: they pre-permute weights at load time so the unpacking is branch-free and coalesced, and they fuse the dequantization into the tensor-core feeding path. These achieve 3-4x over FP16 at small batch, versus ~1.5x for a naive dequantize-then-GEMM.


7-9. Performance, production, mistakes#

Performance: the table in section 5. Note that the gains are not multiplicative with batching gains — quantization reduces the weight-read term, which batching was already amortizing. At batch 256, weights are a smaller share of bytes, so quantization helps less.

Production:

  • Quantize offline, ship the quantized checkpoint. Never at startup.
  • Record the exact config in your registry: method, bits, group size, excluded layers, calibration dataset, library version.
  • Use calibration data resembling your traffic. Calibrating on Wikipedia and serving code is measurably worse.
  • A/B in production. Offline evals miss things.
  • Have a rollback plan. Keep the FP16 checkpoint deployable.
  • Check that your engine has an optimized kernel for the scheme you chose. A W4A16 model on a naive dequant path can be slower than FP16.

Mistakes:

  • Expecting prefill gains from weight-only quantization.
  • Per-tensor scales for INT4. Accuracy collapse.
  • Quantizing the LM head and embeddings with everything else.
  • Validating with perplexity only.
  • Assuming published benchmark deltas transfer to your task.
  • Quantizing before fixing scheduling. A 2x on a system running at batch 4 is worth less than fixing the batch.

10. Hands-on exercise#

A. Run the decision tree. For three real workloads you know (or invent plausible ones), walk the decision tree and justify the recommendation.

B. Measure the phase asymmetry. Take a model, quantize it to INT4 weight-only. Measure prefill throughput and decode throughput before and after. Confirm decode improves ~2.5x and prefill doesn’t. Explain using arithmetic intensity.

C. Layer sensitivity. Quantize a model with (i) everything at INT4, (ii) everything except LM head and embeddings, (iii) FFN only. Measure perplexity and memory for each. Which gives the best quality-per-byte?

D. Calibration matters. Quantize the same model with calibration sets drawn from Wikipedia and from code. Evaluate both on code and on prose. Quantify the mismatch penalty.

E. Full validation. Take one quantized model through all four validation levels. Write up the results as you would for a deployment review.


11. Interview questions#

  1. Why does quantization speed up decode? Give the mechanism.
  2. Why does weight-only quantization not speed up prefill?
  3. What are the four axes that define a quantization scheme?
  4. Which layers would you leave unquantized and why?
  5. How would you choose between FP8 and INT4 for a given workload?
  6. How do you validate a quantized model before deployment?
  7. When does KV cache quantization matter more than weight quantization?

12. Further reading#

  • [ESTABLISHED] Dettmers et al., “LLM.int8()” (2022)
  • [ESTABLISHED] Frantar et al., “GPTQ” (2022); Lin et al., “AWQ” (2023)
  • [ESTABLISHED] Xiao et al., “SmoothQuant” (2022)
  • [REFERENCE] llm-compressor, autoawq, bitsandbytes documentation
  • Next: 03 — GPTQ and AWQ

↑↓ navigate↵ openesc close