1. Problem → Why → Optimization#
PROBLEM BF16 inference uses 2 bytes per value and runs tensor cores at
half their potential rate on Hopper+.
WHY The hardware supports 8-bit tensor core operations at 2x the
throughput, and 8-bit values halve memory traffic.
OPTIMIZE Serve in FP8 (E4M3) or INT8, quantizing both weights and activations.This is the production default for modern inference on Hopper and Blackwell. It is [ESTABLISHED], not experimental.
2. FP8 vs INT8 — the comparison#
FP8 E4M3 INT8
Representation 1 sign, 4 exp, 3 mant 1 sign, 7 magnitude
Max value 448 127
Dynamic range ~2^-9 to 448 (~10⁵) -127 to 127 (~10²)
Precision ~6% relative, uniform absolute, uniform
across magnitudes (relative error bad for small values)
Zero point not needed sometimes needed (asymmetric)
Outlier tolerance GOOD (exponent absorbs) POOR (needs SmoothQuant)
Hardware Hopper+ Turing+
Throughput 2x BF16 2x BF16
Typical quality Δ -0.1 to -0.5% -0.2 to -1.0%
Calibration simple (per-tensor amax) needs care (SmoothQuant)The key difference: FP8 has an exponent. That means its relative precision is roughly constant across magnitudes, so a channel with values around 60 and a channel with values around 0.5 are both represented with ~6% relative error. INT8’s uniform absolute spacing means the 0.5-magnitude channel is destroyed if the tensor’s max is 60.
This is why FP8 usually doesn’t need SmoothQuant and INT8 does. It is the single most practical thing to know about the two formats.
3. Simple analogy#
Measuring with a ruler versus with scientific notation.
INT8 is a ruler with 254 evenly spaced marks. If your longest object is 60 metres, each mark is 24 cm — and you cannot measure a 5 cm object at all.
FP8 is scientific notation with 3 significant figures. It measures 60 metres as 60.0 m and 5 cm as 5.00 cm, both to 3 figures. The relative precision is the same at every scale.
For data with a wide dynamic range — which activations have — scientific notation wins.
4. Tiny example#
package main
import (
"fmt"
"math"
)
// int8PerTensor: one scale for the whole tensor, 255 evenly spaced levels.
func int8PerTensor(x []float64) []float64 {
var absmax float64
for _, v := range x {
absmax = math.Max(absmax, math.Abs(v))
}
s := absmax / 127
out := make([]float64, len(x))
for i, v := range x {
out[i] = math.Round(v/s) * s
}
return out
}
// fp8E4M3: 1 sign bit, 4 exponent bits, 3 mantissa bits. Every power of two gets
// 8 evenly spaced values, so the RELATIVE error is about the same at every magnitude.
func fp8E4M3(v float64) float64 {
if v == 0 {
return 0
}
e := math.Max(math.Floor(math.Log2(math.Abs(v))), -6) // smallest normal exponent
step := math.Pow(2, e-3) // 3 mantissa bits
q := math.Round(v/step) * step
return math.Max(-448, math.Min(448, q)) // largest representable value
}
func main() {
x := []float64{60.0, 0.5, 2.0, 0.05, 30.0}
i8 := int8PerTensor(x)
fmt.Printf("INT8 : %.4f\nrel err:", i8)
for i := range x {
fmt.Printf(" %.3f", math.Abs((i8[i]-x[i])/x[i]))
}
fmt.Print("\nFP8 : [")
for _, v := range x {
fmt.Printf("%.4f ", fp8E4M3(v))
}
fmt.Print("]\nrel err:")
for _, v := range x {
fmt.Printf(" %.3f", math.Abs((fp8E4M3(v)-v)/v))
}
fmt.Println()
}Output:
INT8 : [60.0000 0.4724 1.8898 0.0000 30.2362]
rel err: 0.000 0.055 0.055 1.000 0.008 ← 0.05 became ZERO
FP8 : [60.0000 0.5000 2.0000 0.0508 30.0000 ]
rel err: 0.000 0.000 0.000 0.016 0.000 ← every value within a few percentINT8 annihilated the 0.05 value. FP8 kept it to within 2%. With outliers present, that’s the difference between a working and a broken quantization.
5. Technical explanation#
FP8 in practice#
Format: E4M3 for weights and forward activations
E5M2 for gradients (training only)
Scaling: even with FP8's range, per-tensor scaling helps:
x_fp8 = clamp(x / s, -448, 448)
s chosen so max|x|/s ≈ 448
This is much simpler than INT8's requirements.
Granularity: per-tensor is usually sufficient (contrast with INT8)
per-channel weights + per-token activations for extra quality
(some implementations use per-block, e.g. DeepSeek's 128×128 blocks)
Accumulation: FP32, alwaysNVIDIA’s Transformer Engine handles scale management, including “delayed scaling” (using a history of recent amax values to set the scale, avoiding a synchronization).
INT8 in practice#
Requires the full apparatus from file 04:
Weights: per-channel symmetric, static (from the checkpoint)
Activations: per-token symmetric, dynamic (computed at runtime)
Plus: SmoothQuant preprocessing to make it work at all
Accumulation: INT32
Dequant: Y_fp16 = Y_int32 × s_x[token] × s_w[channel] (fused in the epilogue)Choosing between them#
Hopper/Blackwell available?
├─ YES → FP8. Simpler, better quality, same speed. Done.
└─ NO (Ampere, Turing, or non-NVIDIA)
└─ INT8 with SmoothQuant. Validate carefully.
Special cases:
- AMD MI300: FP8 supported. Use it.
- Older hardware: INT8 only.
- Extreme memory constraints: neither; use INT4 weight-only.What DeepSeek did (worth knowing)#
DeepSeek-V3 trained and serves in FP8 with fine-grained (per-128×128-block for weights, per-128-element-group for activations) scaling and FP32 accumulation at intervals. This is the most aggressive production FP8 deployment publicly documented, and their technical report is worth reading for the numerical details.
Takeaway: FP8 is not just a serving optimization; it’s becoming the native precision.
The blockwise/fine-grained trend#
Per-tensor: 1 scale simplest, most quality loss
Per-channel: C scales standard for weights
Per-token: T scales standard for activations
Per-block: (T/128)×(C/128) scales ← DeepSeek-style, best qualityFiner granularity costs a little memory and a little kernel complexity, and buys quality. The trend is toward finer.
6-9. Under the hood, performance, production, mistakes#
Under the hood — a FP8 GEMM epilogue:
tensor core: FP8 × FP8 → FP32 accumulator
epilogue (in registers):
acc *= (s_a × s_b) dequantize
acc += bias
acc = activation(acc)
if next layer is FP8:
acc = clamp(acc / s_out, -448, 448) requantize
store as FP8 ← never touches HBM in FP16!
else:
store as BF16That last branch matters: with FP8 throughout, activations move between layers as 1 byte, halving activation traffic too. Mixed FP8/BF16 pipelines lose that.
Performance (70B, H100 node, TP=8):
Prefill tok/s Decode tok/s (b=64) Memory Max batch
BF16 28,000 1,530 141 GB ~400
FP8 50,000 2,900 71 GB ~800
INT8 (SmoothQ) 46,000 2,750 71 GB ~800Both roughly 1.8-1.9x. FP8 slightly ahead and much simpler to produce.
Production:
- Default to FP8 on Hopper+. It is the current production standard.
- Use per-tensor or finer scaling; use delayed scaling to avoid syncs.
- Validate at all four levels (Section IV.12). FP8’s quality loss is small but nonzero.
- Watch for NaN/inf. FP8’s max is 448; an unscaled activation exceeding it becomes inf. Scale management is where FP8 bugs live.
- Ship the quantized checkpoint with its scales.
- Note that FP8 KV cache is a separate decision (file 12) and often an additional 1.3-1.8x.
Mistakes:
- Using INT8 on Hopper when FP8 is available. More work, worse quality.
- Forgetting scale management. FP8 overflow → inf → NaN.
- Per-tensor INT8 activations without SmoothQuant. Quality collapse.
- Assuming FP8 is lossless. It isn’t; validate.
- Mixing FP8 and BF16 unnecessarily, losing the activation-traffic benefit.
10. Hands-on exercise#
A. Compare the formats. Run the section 4 example. Extend it to a real activation tensor from a model. Plot per-value relative error for INT8 and FP8. Where does each fail?
B. Quantize and measure. Take a model, produce FP8 and INT8 versions. Measure: prefill throughput, decode throughput, memory, perplexity, and one downstream benchmark. Build the comparison table.
C. Scale management. Deliberately set an FP8 scale too small and observe the overflow to inf. Then implement amax-based scaling and confirm it’s fixed.
D. Activation traffic. Compare a fully-FP8 pipeline to one that dequantizes to BF16 between layers. Measure the activation memory traffic difference in a profile.
E. Granularity. If your tooling supports it, compare per-tensor, per-channel, and per-block FP8 scaling on quality. Is the finer granularity worth it for your model?
11. Interview questions#
- Compare FP8 E4M3 and INT8. Which handles outliers better and why?
- Why does FP8 often not need SmoothQuant?
- What is delayed scaling and what problem does it solve?
- What is FP8’s maximum value and what happens when you exceed it?
- Why does a fully-FP8 pipeline beat a mixed FP8/BF16 one?
- When would you use INT8 instead of FP8?
- What is fine-grained (block) scaling and what does it buy?
12. Further reading#
- [ESTABLISHED] Micikevicius et al., “FP8 Formats for Deep Learning” (2022)
- [REFERENCE] NVIDIA Transformer Engine documentation
- [ESTABLISHED] DeepSeek-V3 technical report — FP8 training and serving at scale
- [ESTABLISHED] Xiao et al., “SmoothQuant” (for the INT8 path)
- Next: 06 — INT4 and low-bit inference