PidokuInfra

Precision and Quantization on Hardware

Advanced Intermediate 1h Difficulty 4/5 Topic 02 of 04

Prerequisites II.04, IV.01

The idea in one minute#

Quantization stores a model’s numbers in fewer bits. From the hardware’s point of view this is the best trade available, because it improves three limits at once: the model takes less memory, each pass moves fewer bytes (so memory-bound work speeds up), and low-bit arithmetic runs faster on tensor cores (so compute-bound work speeds up).

The price is accuracy, and the craft is choosing where the lost precision does no harm.

A picture#

flowchart LR
  FP["Weights in FP16<br/>16 bits each"] --> Q["Quantize<br/>scale and round"]
  Q --> I8["Weights in 8 or 4 bits<br/>plus a scale per group"]
  I8 --> M["Capacity<br/>2x to 4x more fits"]
  I8 --> BW["Bandwidth<br/>2x to 4x fewer bytes per token"]
  I8 --> TC["Arithmetic<br/>faster low-bit tensor cores"]
  Q -.->|"rounding error"| ACC["Accuracy cost<br/>must be measured"]
  class FP,I8 memory
  class Q neutral
  class M,BW,TC compute
  class ACC warn

How it really works#

The basic mechanism#

To store a group of weights in 8-bit integers:

scale = max(|w|) ÷ 127
q     = round(w ÷ scale)          // an integer in [-127, 127]
w'    = q × scale                 // what the model actually uses

The error per weight is at most scale ÷ 2. Using one scale per small group (per row, or per block of 32–128 weights) keeps the scale — and so the error — small where weights are small.

What each format buys on an H100-class GPU#

FormatBits8B-parameter modelMemory-bound speed vs FP16Tensor-core path
FP323232 GB0.5xNo (slow path)
FP16 / BF161616 GB1xYes
FP8 / INT888 GB~2xYes
INT444 GBup to ~4xUsually unpacked to a wider format first
FP444 GBup to ~4xYes, on the newest generation

Two cautions:

  • Speedup needs hardware support. Bytes saved always help capacity. They only help speed if the kernel can consume the format efficiently. Formats the tensor cores do not natively handle must be unpacked on the fly, which costs arithmetic and can cancel the gain at large batch sizes.
  • Weights and activations are different questions. Quantizing weights shrinks storage and traffic. Quantizing activations (the values flowing through the model) is what lets the arithmetic itself run in low precision — and it is harder, because activations vary with the input.

Why networks survive it#

A weight’s exact value matters less than the combined effect of thousands of them, and rounding errors that are not systematically biased largely cancel in the sum. In practice 8-bit is close to lossless for most models, 4-bit is usually acceptable with careful methods, and below that quality falls quickly.

“Usually acceptable” is an empirical claim about your model and your task. Always evaluate.

The parts that stay in high precision#

As in II.04: normalization layers, softmax, and long accumulations are sensitive. Quantized models keep these in FP16 or FP32. The first and last layers are also often left at higher precision because errors there are not averaged away.

Code#

Quantize a matrix to 8 and 4 bits, measure the error in a matrix-vector product, and see why per-group scales matter.

Go
// quant.go — quantize weights, then measure what it does to the output.
package main

import (
	"fmt"
	"math"
	"math/rand"
)

// quantize rounds each group of `group` weights to `bits`-bit signed integers with its own scale,
// and returns the de-quantized values the model would actually compute with.
func quantize(w []float64, bits, group int) []float64 {
	levels := math.Pow(2, float64(bits-1)) - 1 // 127 for 8 bits, 7 for 4 bits
	out := make([]float64, len(w))
	for g := 0; g < len(w); g += group {
		end := min(g+group, len(w))
		var maxAbs float64
		for _, v := range w[g:end] {
			maxAbs = max(maxAbs, math.Abs(v))
		}
		scale := maxAbs / levels
		for i := g; i < end; i++ {
			out[i] = math.Round(w[i]/scale) * scale
		}
	}
	return out
}

func matVec(w, x []float64, rows, cols int) []float64 {
	y := make([]float64, rows)
	for i := 0; i < rows; i++ {
		for j := 0; j < cols; j++ {
			y[i] += w[i*cols+j] * x[j]
		}
	}
	return y
}

func relErr(a, b []float64) float64 {
	var num, den float64
	for i := range a {
		num += (a[i] - b[i]) * (a[i] - b[i])
		den += a[i] * a[i]
	}
	return math.Sqrt(num / den)
}

func main() {
	const rows, cols = 512, 512
	rng := rand.New(rand.NewSource(1))
	w, x := make([]float64, rows*cols), make([]float64, cols)
	for i := range w {
		w[i] = rng.NormFloat64() * 0.02
		if rng.Intn(1000) == 0 {
			w[i] *= 20 // a few large outliers, as real weight matrices have
		}
	}
	for i := range x {
		x[i] = rng.NormFloat64()
	}
	exact := matVec(w, x, rows, cols)

	fmt.Println("bits  group      output error   size vs FP16")
	for _, c := range []struct{ bits, group int }{
		{8, rows * cols}, {8, 64}, {4, rows * cols}, {4, 64}, {2, 64},
	} {
		got := matVec(quantize(w, c.bits, c.group), x, rows, cols)
		name := fmt.Sprint(c.group)
		if c.group == rows*cols {
			name = "whole"
		}
		fmt.Printf("%4d  %-8s  %11.2f%%   %9.0f%%\n", c.bits, name, 100*relErr(exact, got), 100*float64(c.bits)/16)
	}
}

Look at the two 4-bit rows. With one scale for the whole matrix, the rare large weights force a coarse scale on everything. With a scale per 64 weights, the error drops sharply. That is the single most important idea in practical quantization.

Even the grouped 4-bit error here is large, because this program rounds each weight to the nearest level without looking at anything else. Production methods (GPTQ, AWQ and their descendants) choose the rounding that minimises the error in the layer’s output on sample data, and that is what makes 4-bit usable in practice.

Remember this#

  • Fewer bits improve capacity, bandwidth-bound speed and (with hardware support) arithmetic speed at once.
  • Speed gains require a format the kernels consume natively.
  • Small quantization groups keep the scale tight and the error low.
  • 8-bit is nearly free; 4-bit needs care; always measure accuracy on your own task.

Try it#

  1. Run quant.go. Remove the outliers. How do the “whole” rows change, and why?
  2. Add group sizes 16 and 256 at 4 bits. Each group stores one 16-bit scale: compute the true bits per weight for each group size.
  3. An H100 holds a 70B model in which format(s)? Use II.04’s table.

Check yourself#

  1. Which three hardware limits does quantization relax?
  2. When does a smaller format not make inference faster?
  3. Why do per-group scales reduce error?

↑↓ navigate↵ openesc close