PidokuInfra

Number Formats: FP32, FP16, BF16, FP8, INT8, INT4

Foundations Intermediate 1h 30m Difficulty 3/5 Topic 11 of 12

Prerequisites 07


1. What is it?#

How a number is encoded in bits. The choice determines memory footprint, memory bandwidth, arithmetic throughput, and how much precision you lose.

FP32:  [S][ 8 exp ][      23 mantissa      ]   4 bytes
FP16:  [S][5 exp][  10 mantissa  ]             2 bytes
BF16:  [S][ 8 exp ][ 7 mantissa ]              2 bytes
FP8 E4M3: [S][4exp][3 man]                     1 byte
FP8 E5M2: [S][ 5exp][2man]                     1 byte
INT8:  [S][    7 bits    ]                     1 byte  (+ a scale per group)
INT4:  [S][3 bits]                             0.5 byte (+ scale)

Since decode is memory-bound, bytes per parameter is directly proportional to decode latency. Halving the bits roughly halves the time. That is the entire commercial motivation.


2. Why does it exist?#

Floating point trades a fixed bit budget between range (how big/small numbers can be, set by the exponent) and precision (how finely you can distinguish nearby numbers, set by the mantissa).

Different workloads want different splits. Neural networks turn out to need surprisingly little precision and a fair amount of range — which is why BF16 (range of FP32, precision of almost nothing) works so well.


3. Simple analogy#

Scientific notation with a fixed number of digits.

6.02 × 10²³ — the exponent (23) gives range, the mantissa (6.02) gives precision.

With a fixed budget of characters you choose: more exponent digits lets you write both 10⁻³⁰ and 10³⁰, but then you can only write 6 × 10²³, not 6.022 × 10²³. More mantissa digits gives 6.02214 × 10²³ but you can’t represent very large or very small values at all.

FP16 chose precision. BF16 chose range. For neural networks, range turned out to matter more.


4. Tiny example#

Go
// formats.go — FP16 and BF16 built from FP32 bits, so you can see what each one keeps.
package main

import (
	"fmt"
	"math"
)

// toBF16 keeps the top 16 bits of an FP32 (sign, 8 exponent bits, 7 fraction bits).
func toBF16(f float32) float32 {
	b := math.Float32bits(f)
	b += 0x7fff + (b>>16)&1 // round to nearest, ties to even
	return math.Float32frombits(b &^ 0xffff)
}

// toFP16Bits converts to IEEE half precision: sign, 5 exponent bits, 10 fraction bits.
func toFP16Bits(f float32) uint16 {
	b := math.Float32bits(f)
	sign := uint16(b>>16) & 0x8000
	exp := int(b>>23&0xff) - 127 + 15
	frac := b & 0x7fffff
	switch {
	case exp >= 31: // too large: infinity
		return sign | 0x7c00
	case exp <= 0: // too small: subnormal or zero
		if exp < -10 {
			return sign
		}
		frac = (frac | 0x800000) >> uint(1-exp)
		return sign | uint16((frac+0x1000)>>13)
	}
	h := sign | uint16(exp)<<10 | uint16(frac>>13)
	if rem := frac & 0x1fff; rem > 0x1000 || (rem == 0x1000 && h&1 == 1) {
		h++ // round to nearest, ties to even
	}
	return h
}

func fromFP16Bits(h uint16) float32 {
	sign, exp, frac := uint32(h&0x8000)<<16, uint32(h>>10)&0x1f, uint32(h&0x3ff)
	switch exp {
	case 0:
		return float32(math.Copysign(float64(frac)*math.Pow(2, -24), -float64(sign>>31)+0.5))
	case 31:
		return math.Float32frombits(sign | 0x7f800000 | frac<<13)
	}
	return math.Float32frombits(sign | (exp+112)<<23 | frac<<13)
}

func toFP16(f float32) float32 { return fromFP16Bits(toFP16Bits(f)) }

func main() {
	x := float32(0.1)
	fmt.Printf("FP32: %.20f\n", x)         // 0.10000000149011611938
	fmt.Printf("FP16: %.20f\n", toFP16(x)) // 0.09997558593750000000
	fmt.Printf("BF16: %.20f\n", toBF16(x)) // 0.10009765625000000000

	fmt.Println(toFP16(1e30), toBF16(1e30))   // +Inf 1.0004e+30    FP16 overflows at 65504
	fmt.Println(toFP16(1e-30), toBF16(1e-30)) // 0 1.0002e-30       FP16 underflows

	fmt.Printf("%d = %#04x = %016b\n", toFP16Bits(x), toFP16Bits(x), toFP16Bits(x)) // 11878 = 0x2e66
}

BF16 is less precise than FP16 for this value (7 vs 10 mantissa bits). But:

Go
// From the program above:
fmt.Println(toFP16(1e30))  // +Inf          ← FP16 overflows at 65504
fmt.Println(toBF16(1e30))  // ≈1e+30        ← fine: BF16 has FP32's exponent range

fmt.Println(toFP16(1e-30)) // 0             ← underflows
fmt.Println(toBF16(1e-30)) // ≈1e-30        ← fine

BF16 never overflows where FP32 wouldn’t. That single property eliminates an entire class of production bugs (NaN in attention, inf in SwiGLU) at the cost of precision that neural networks mostly don’t need.


5. Technical explanation#

The table#

FormatBitsExpMantMaxMin normalRel. precisionTensor core?
FP32328233.4e381.2e-38~1e-7(TF32 on Ampere+)
TF3219*8103.4e381.2e-38~1e-3yes
FP1616510655046.1e-5~1e-3yes
BF1616873.4e381.2e-38~1e-2yes
FP8 E4M38434482^-6~6e-2Hopper+
FP8 E5M2852573442^-14~1e-1Hopper+
FP6 / FP46/4variesvariessmall—very coarseBlackwell
INT88——127 × scale—scale/127yes
INT44——7 × scale—scale/7via dequant

*TF32 is stored in 32 bits but computes with 10 mantissa bits.

Throughput on NVIDIA hardware#

H100 SXM (dense, no sparsity):
  FP64        34 TFLOP/s
  FP32        67 TFLOP/s
  TF32       495 TFLOP/s
  BF16/FP16  990 TFLOP/s
  FP8       1979 TFLOP/s
  INT8      1979 TOP/s

Each halving of precision doubles tensor-core throughput. Combined with halving memory traffic, low precision is a double win — which is why the industry moves down the ladder relentlessly.

What each is used for#

FP32   master weights in training; accumulation; softmax; a few sensitive ops
TF32   drop-in faster FP32 matmul (Ampere+). Enabled by default in recent PyTorch.
BF16   the default for modern training and much inference. Safe range.
FP16   still common; requires care with range (loss scaling in training, FP32 softmax)
FP8    production inference on Hopper/Blackwell. E4M3 for weights/activations.
INT8   production inference everywhere; needs calibration
INT4   weight-only quantization for memory-bound decode; the sweet spot for many

FP8’s two flavors#

E4M3: 4 exponent, 3 mantissa. Max 448. More precision, less range.
      → weights and activations (forward pass)
E5M2: 5 exponent, 2 mantissa. Max 57344. More range, less precision.
      → gradients (training), where range matters more

For inference you almost always want E4M3, usually with a per-tensor or per-channel scaling factor to bring values into its narrow range.

Accumulation precision#

Critical and often overlooked: the format of the inputs and the format of the accumulator are different.

FP16 inputs × FP16 inputs → FP32 accumulator → FP16 output
FP8 inputs  × FP8 inputs  → FP32 accumulator → FP16/FP8 output
INT8 × INT8 → INT32 accumulator → dequantize

Accumulating a 4096-length dot product in FP16 would lose catastrophic precision (each addition rounds; errors accumulate as ~sqrt(K)). Tensor cores always accumulate in higher precision. When you write custom kernels, preserve this.

Denormals and zeros#

Very small values below the “normal” range are represented as denormals (gradual underflow). They’re slow on some hardware and often flushed to zero (FTZ). For inference this is almost always fine and sometimes faster.


6. Under the hood#

Bit-level: FP16 0.1:

0.1 in binary = 0.0001100110011001100... (repeating)
normalized:     1.100110011001100... × 2⁻⁴

sign:      0
exponent:  -4 + 15 (bias) = 11 = 01011
mantissa:  1001100110  (10 bits, rounded)

bits: 0 01011 1001100110  =  0x2E66  →  0.0999755859375

The error is 2.4e-5, about 1e-4 relative — consistent with 10 mantissa bits (2⁻¹⁰ ≈ 1e-3 worst case, better on average).

You can inspect this directly:

Go
// Using toFP16Bits from the program in section 4:
fmt.Println(toFP16Bits(0.1))           // 11878 = 0x2E66
fmt.Printf("%016b\n", toFP16Bits(0.1)) // 0 01011 1001100110

Doing this once for a few values makes floating point stop being magic.


7. Performance implications#

For a 70B model, decode, single H100:

PrecisionWeight bytesMin ms/tokenMax tok/sQuality
FP32280 GBdoesn’t fit—reference
BF16140 GB41.823.9reference
FP870 GB20.947.8~0-1% loss
INT870 GB20.947.8~0-1% loss
INT435 GB10.495.7~1-3% loss

Straight-line scaling, because decode is bandwidth-bound and bandwidth is what you’re saving.

For prefill (compute-bound) the picture differs:

BF16 → FP8:  2x, because tensor-core throughput doubles
BF16 → INT4 weight-only:  ~1x, because compute is still BF16 after dequantization

Same quantization, completely different benefit by phase. This is the single most misunderstood point about quantization and it comes up in every interview.


8. Production implications#

  • Default to BF16 for any model you serve unquantized. The range safety is worth the precision.
  • Enable TF32 for any residual FP32 matmuls (torch.backends.cuda.matmul.allow_tf32 = True).
  • FP8 on Hopper/Blackwell is production-ready and gives ~2x with minimal quality loss for most models. It requires per-tensor (or finer) scaling factors, computed by calibration.
  • Verify accumulation precision in any custom kernel. FP16 accumulation over long reductions is a correctness bug waiting to happen.
  • Always evaluate quality after a precision change, on your own task, with long generations.
  • Record the serving precision in your model registry. “Llama-3-70B” is not a sufficient description of what you deployed.

9. Common mistakes#

Using FP16 where BF16 is available. Range problems for no benefit.

Accumulating in low precision. Silent accuracy loss.

Assuming quantization helps prefill as much as decode. It usually doesn’t.

Comparing quantized models only on short benchmarks. Degradation compounds over long generations.

Forgetting the scale factors. INT8/INT4 store integers plus a scale per tensor/channel/ group. The scales are part of your memory budget (small but not zero) and part of your kernel’s work.

Mixing dtypes accidentally. One FP32 tensor in an FP16 graph forces upcasts and disables tensor cores for that op. Check with a profiler.


10. Hands-on exercise#

A. Bit inspection. Write a function that prints the sign/exponent/mantissa bits for a value in FP32, FP16, and BF16. Run it on 0.1, 1.0, 65504, 1e30, 1e-30. Explain each result.

B. Find the limits. Empirically determine, for each format: max value, min normal value, and the smallest ε such that 1 + ε != 1. Compare to the table.

C. Precision loss in a dot product. Compute a length-4096 dot product of random values in FP32, in FP16 with FP16 accumulation, and in FP16 with FP32 accumulation. Measure relative error against a FP64 reference. Explain the difference.

D. Throughput. Benchmark matmul on your GPU in FP32, TF32, BF16, FP16, and (if Hopper+) FP8. Report achieved TFLOP/s for each. Compare to spec. Record in numbers.md.

E. Range failure. Construct a realistic attention computation that overflows in FP16 and works in BF16. What context length or head dimension triggers it?


11. Interview questions#

  1. Compare FP16 and BF16. Which would you choose for inference and why?
  2. What is the max value in FP16, and why does that matter for attention?
  3. What are E4M3 and E5M2 and when do you use each?
  4. Why do tensor cores accumulate in FP32?
  5. Why does INT4 weight-only quantization help decode much more than prefill?
  6. What is TF32 and when is it used?
  7. A model produces NaN in FP16 but works in BF16. Explain the likely cause.

12. Further reading#

  • [FUNDAMENTAL] Goldberg, “What Every Computer Scientist Should Know About Floating-Point Arithmetic” (1991)
  • [ESTABLISHED] Micikevicius et al., “FP8 Formats for Deep Learning” (2022)
  • [ESTABLISHED] Micikevicius et al., “Mixed Precision Training” (2018)
  • [REFERENCE] NVIDIA Transformer Engine docs; PyTorch dtype documentation
  • Next: 12 — Quantization fundamentals

↑↓ navigate↵ openesc close