PidokuInfra

Activation Quantization and SmoothQuant

Intermediate Advanced 1h 15m Difficulty 4/5 Topic 04 of 14

Prerequisites 02, 03


1. Problem → Why → Optimization#

PROBLEM   Weight-only quantization doesn't speed up prefill, because the
          compute is still FP16. To get compute benefits you must quantize
          activations too — but naive activation quantization destroys quality.

WHY       Activations contain systematic OUTLIERS: a small number of channels
          have magnitudes 20-100x larger than the rest, consistently, in
          every transformer above ~6B parameters.

OPTIMIZE  SmoothQuant: migrate the outlier magnitude from activations into
          weights via a per-channel scaling that cancels mathematically.

2. Why outliers exist#

Dettmers et al. (LLM.int8()) discovered this empirically: above roughly 6.7B parameters, transformers develop emergent outlier features — specific hidden dimensions where activation magnitudes are enormous and which appear in essentially every token.

Typical activation distribution in a large transformer's FFN input:
  99.9% of values:  |x| < 3
  ~0.1% of values (specific channels):  |x| up to 70-100

Quantizing per-tensor to INT8:
  scale = 100/127 = 0.787
  a value of 2.0 → round(2.0/0.787) = 3 → 2.36     18% error!
  Most of the tensor is destroyed to accommodate a handful of channels.

These outliers are functionally important — removing them degrades the model badly — so you can’t simply clip them.

Weights, by contrast, are well-behaved: roughly Gaussian, few outliers. This asymmetry is why weight-only quantization is easy and activation quantization is hard.


3. Simple analogy#

A photograph with one very bright light source.

Auto-exposure for the whole frame: the bright light forces a short exposure, and everything else is a black silhouette. You’ve lost all the detail that matters to accommodate one pixel.

Options:

  • Per-region exposure (per-channel quantization) — good, but requires per-region processing.
  • Neutral density filter over the bright area, brightened afterward (SmoothQuant) — reduce the dynamic range you must capture, then compensate.
  • Photograph the bright part separately in a different mode (LLM.int8’s mixed precision).

All three are real quantization strategies.


4. Tiny example#

Go
package main

import (
	"fmt"
	"math"
)

func quantInt8(v []float64) []float64 {
	var absmax float64
	for _, x := range v {
		absmax = math.Max(absmax, math.Abs(x))
	}
	s := absmax / 127
	out := make([]float64, len(v))
	for i, x := range v {
		out[i] = math.Round(x/s) * s
	}
	return out
}

func dot(a, b []float64) (s float64) {
	for i := range a {
		s += a[i] * b[i]
	}
	return s
}

func spread(v []float64) float64 { // largest / smallest magnitude
	lo, hi := math.Inf(1), 0.0
	for _, x := range v {
		lo, hi = math.Min(lo, math.Abs(x)), math.Max(hi, math.Abs(x))
	}
	return hi / lo
}

func main() {
	x := []float64{1.0, 2.0, 1.5, 60.0, 0.5} // channel 3 is an outlier
	w := []float64{0.5, 0.3, 0.4, 0.2, 0.6}
	fmt.Printf("exact : %.4f\n", dot(x, w)) // 0.5+0.6+0.6+12.0+0.3 = 14.0

	// --- Naive per-tensor INT8 on both ---
	fmt.Printf("naive : %.4f\n", dot(quantInt8(x), quantInt8(w)))

	// --- SmoothQuant: migrate magnitude from x to w ---
	const alpha = 0.5
	s, mean := make([]float64, len(x)), 0.0
	for i := range x {
		s[i] = math.Pow(math.Abs(x[i]), alpha) / math.Pow(math.Abs(w[i]), 1-alpha)
		mean += s[i] / float64(len(x))
	}
	xs, ws := make([]float64, len(x)), make([]float64, len(x))
	for i := range x {
		s[i] /= mean                        // normalize
		xs[i], ws[i] = x[i]/s[i], w[i]*s[i] // xs·ws == x·w exactly
	}
	fmt.Printf("smooth: %.4f\n", dot(quantInt8(xs), quantInt8(ws)))
	fmt.Printf("ranges — x: %.0fx   x smoothed: %.1fx\n", spread(x), spread(xs))
}

Output:

exact : 14.0000
naive : 13.8027      (1.4% error)
smooth: 13.9671      (0.2% error)
ranges — x: 120x   x smoothed: 6.3x

The dynamic range dropped from 120x to 6.3x, and the quantization error dropped 6x. The weights absorbed the difficulty — and weights are easier to quantize because they’re static and you can use per-channel scales.


5. Technical explanation#

SmoothQuant, formally#

For Y = X W with X: (T, C_in) and W: (C_in, C_out):

Y = X W = (X diag(s)⁻¹)(diag(s) W) = X̂ Ŵ

Choose s_j = max(|X_j|)^α / max(|W_j|)^(1-α)      per input channel j

α = 0    → all difficulty in the activations (no change)
α = 0.5  → balanced (default; works for most models)
α = 0.8  → more difficulty pushed to weights (for models with severe outliers)

The diag(s)⁻¹ is folded into the preceding operation — the LayerNorm’s weight, or the previous linear layer’s output scaling — so there is zero runtime cost.

LayerNorm(x) · γ  →  LayerNorm(x) · (γ / s)      ← fold here

This folding is what makes SmoothQuant free. It works because transformers have a LayerNorm or a linear layer immediately before every quantized matmul.

Granularity for activations#

Per-tensor:   one scale for the whole activation tensor.
              Fastest. Requires SmoothQuant or you lose too much.
Per-token:    one scale per token (row). Computed dynamically — a max-reduction
              over the hidden dimension. Cheap and much more robust.
              THE PRACTICAL DEFAULT.
Per-channel:  one scale per hidden dimension. Best accuracy, but INCOMPATIBLE
              with standard INT8 GEMM (the scale must be applied along the
              reduction dimension, which the tensor core can't do).

Per-token dynamic + per-channel static weights is the standard combination, because it’s both accurate and kernel-friendly:

Y[t,o] = Σ_i X̂[t,i] Ŵ[i,o] × s_x[t] × s_w[o]
                                ^^^^^^^^^^^^^^ both outside the reduction ✓

Per-channel activation scales would put s_x[i] inside the sum, which the tensor core can’t handle. This constraint drives the whole design.

LLM.int8() — the mixed-precision alternative#

1. Identify outlier channels at runtime (|x| > threshold, typically 6.0).
2. Split the matmul:
     outlier columns  → FP16 matmul  (~0.1% of columns)
     regular columns  → INT8 matmul
3. Sum the results.

Preserves quality very well but is slower than pure INT8 because of the split and the extra FP16 path. It was the first demonstration that outliers were the problem; SmoothQuant is the faster successor.

Static vs dynamic activation scales#

STATIC:   scales from calibration, fixed at runtime.
          ✓ zero runtime cost
          ✗ fails if production activations exceed the calibration range
            (a longer prompt, a different language, an adversarial input)
            → clipping → quality collapse on those inputs

DYNAMIC:  scale = max(|x|) computed per token, at runtime.
          ✓ robust to any input
          ✗ a reduction over the hidden dim per token (cheap: ~1% of the GEMM)

Use dynamic per-token. The robustness is worth the 1%. Static per-tensor activation scaling is a known source of rare, hard-to-reproduce quality failures.


6-9. Under the hood, performance, production, mistakes#

Under the hood — the INT8 GEMM:

INT8 × INT8 → INT32 accumulator (hardware requirement)
Then dequantize: Y_fp16[t,o] = Y_int32[t,o] × s_x[t] × s_w[o]

The dequantization is fused into the GEMM epilogue — free.

Performance:

70B, Hopper:
                        Prefill      Decode      Memory
BF16                     1.00x        1.00x       140 GB
INT8 W8A8 (SmoothQuant)  1.5-1.8x     1.7-1.9x     70 GB
FP8 W8A8                 1.6-1.9x     1.8-2.0x     70 GB
INT8 weight-only         ~1.0x        1.7-1.9x     70 GB

Note that W8A8 and weight-only give the same decode gain (both halve the weight bytes) but only W8A8 gives a prefill gain. If your workload has any meaningful prefill, W8A8 is strictly better — the only cost is the calibration complexity.

Production:

  • Prefer FP8 over INT8 on Hopper+. FP8’s wide dynamic range makes outliers far less problematic — often you don’t need SmoothQuant at all. Simpler and equally fast.
  • Use dynamic per-token activation scales.
  • Calibrate on representative data, including your longest and most unusual inputs.
  • Test with adversarial inputs: very long prompts, unusual languages, repeated tokens. Static-scale failures show up there.
  • Monitor for clipping if you use static scales — count how often activations exceed the calibrated range.

Mistakes:

  • Naive per-tensor activation quantization without SmoothQuant. Quality collapse.
  • Static scales with unrepresentative calibration. Rare, severe, hard-to-debug failures.
  • Trying per-channel activation scales with a standard INT8 GEMM. Doesn’t work.
  • Assuming FP16 outlier behavior transfers to all model sizes. Outliers emerge above ~6B.
  • Not testing on out-of-distribution inputs.

10. Hands-on exercise#

A. Find the outliers. Hook a real model’s forward pass and collect activation statistics per channel for the FFN input of a middle layer. Plot the per-channel max. How many channels are outliers? What’s the ratio to the median channel?

B. Reproduce the failure. Quantize those activations per-tensor to INT8 and measure the relative error on non-outlier values. Then apply SmoothQuant with α ∈ {0, 0.25, 0.5, 0.75, 1.0} and plot error vs α. Where’s the minimum?

C. Per-token vs per-tensor. Compare quantization error for per-tensor and per-token activation scales on real activations. Quantify the improvement.

D. Static scale failure. Calibrate static scales on short English prompts. Then run a very long prompt and a non-English prompt. Measure how often activations exceed the calibrated range and what it does to output quality.

E. FP8 vs INT8. If you have Hopper hardware, quantize the same model to INT8 (with SmoothQuant) and FP8 (without). Compare quality and speed. Was SmoothQuant necessary for FP8?


11. Interview questions#

  1. Why is activation quantization harder than weight quantization?
  2. What are emergent outlier features and when do they appear?
  3. Explain SmoothQuant’s mechanism, including why it’s free at runtime.
  4. Why can’t you use per-channel activation scales with a standard INT8 GEMM?
  5. Static vs dynamic activation scales — which and why?
  6. Why does W8A8 beat weight-only for prefill-heavy workloads?
  7. Why does FP8 often not need SmoothQuant?

12. Further reading#

  • [ESTABLISHED] Dettmers et al., “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale” (2022) — the outlier discovery
  • [ESTABLISHED] Xiao et al., “SmoothQuant” (2022)
  • [REFERENCE] NVIDIA TensorRT-LLM quantization documentation
  • Next: 05 — FP8 and INT8 inference

↑↓ navigate↵ openesc close