1. Where the field is#
FORMAT BITS STATUS NOTES
FP32 32 legacy only for accumulation and a few sensitive ops
BF16 16 ESTABLISHED the safe default
FP16 16 ESTABLISHED range issues at depth (Section III.11)
FP8 8 ESTABLISHED production standard on Hopper+
INT8 8 ESTABLISHED production standard elsewhere
INT4 4 ESTABLISHED weight-only; production-common
FP6 6 EMERGING Blackwell hardware support
FP4 4 EMERGING Blackwell (MXFP4/NVFP4); training and inference
INT3 3 RESEARCH significant quality loss
INT2 2 RESEARCH needs learned codebooks
1.58-bit — RESEARCH BitNet: trained low-bit, not quantizedThe frontier is at 4 bits with hardware support. Below that, the techniques change character: you stop quantizing a pretrained model and start training for low precision.
2. What Blackwell changes#
Blackwell adds NATIVE tensor core support for FP4 and FP6, with
microscaling (MX) formats.
MXFP4: 4-bit elements + a shared 8-bit exponent scale per block of 32
→ effective bits: 4 + 8/32 = 4.25 bits/value
→ the block-level scale is what makes 4 bits viable
NVFP4: NVIDIA's variant, block size 16 with an FP8 scale,
plus a per-tensor FP32 scale
→ effective bits: 4 + 8/16 = 4.5
→ finer blocks → better accuracy
THROUGHPUT (approximate, per NVIDIA's Blackwell figures)
BF16: 1.0x
FP8: 2.0x
FP4: 4.0x
MEMORY
FP4 weights are 1/4 of BF16
→ a 405B model at FP4 is ~203 GB (fits on 3 GPUs instead of 11)The critical difference from INT4 weight-only (Section VII.06): FP4 has tensor core support, so it accelerates COMPUTE as well as memory traffic. That means it helps prefill too — which INT4 weight-only doesn’t.
Decode Prefill Memory
INT4 W4A16 2.5-3.5x ~1.0x 0.28x no compute benefit
FP4 (W4A4) 3.0-4.0x ~3.5x 0.28x compute AND memory3. Why block scaling makes 4 bits work#
Recall Section III.12: quantization error is bounded by the SCALE,
and the scale is determined by the RANGE of the group.
per-tensor scale, 4 bits (16 levels), range determined by the
tensor's max → catastrophic for small values
per-block scale (32 values), 4 bits → the range is only that
block's range, which is much narrower
→ error relative to the block's own magnitude, not the tensor'sEFFECTIVE BITS COMPARISON
INT4, per-tensor: 4.0 bits unusable
INT4, group 128, FP16 scale: 4.125 workable (Section VII.06)
MXFP4, block 32, E8 scale: 4.25 better
NVFP4, block 16, FP8 scale: 4.5 best of these
The extra 0.25-0.5 bits buys most of the accuracy.This is the same lesson as Section VII.06’s group-size discussion, taken to its conclusion: fine-grained scaling is what makes aggressive quantization viable, and the hardware now supports it natively.
4. What’s needed to use FP4 well#
1. QUANTIZATION-AWARE PREPARATION
Naive round-to-nearest at 4 bits loses too much.
Need: GPTQ/AWQ-style calibration (Section VII.03), or
quantization-aware fine-tuning, or
training in low precision from the start
2. SELECTIVE PRECISION
Not everything should be FP4:
✓ FFN weights (the bulk)
~ attention projections (more sensitive)
✗ embeddings, LM head, norms, router weights (MoE)
✗ accumulators (FP32, always)
3. KERNEL SUPPORT
Blackwell tensor cores + CUTLASS/cuBLAS support for MXFP4/NVFP4
Engine support is arriving
4. VALIDATION
The full Level 1-4 (Section IV.12), with extra attention to
long generations and reasoning tasks — where low-bit hurts most
(Section VII.06)5. The honest assessment#
WHAT'S SOLID
✓ FP8 is production-standard and well-validated
✓ INT4 weight-only is production-common for decode-bound workloads
✓ block/microscaling is the right mechanism
✓ Blackwell's hardware support is real
WHAT'S NOT YET SETTLED
? quality at FP4 for demanding tasks (reasoning, code, long-form)
? whether post-training FP4 is sufficient or QAT is required
? engine support maturity
? which format (MXFP4 vs NVFP4 vs others) becomes standard
? whether FP6 is the practical sweet spot rather than FP4
WHAT TO DO NOW
→ deploy FP8 today; it's mature
→ INT4 weight-only where decode-bound and validated
→ watch FP4; re-evaluate when you have Blackwell and engine support
→ don't build architecture assuming FP4 works for your task until
you've measured itThe pattern from FP8’s adoption is instructive: hardware support arrived, then a year of validation work, then it became standard. FP4 is roughly at the start of that curve.
6. Sub-4-bit: a different problem#
Below 4 bits, post-training quantization stops working well and the
approaches change:
VECTOR / CODEBOOK QUANTIZATION [EMERGING→RESEARCH]
AQLM, QuIP#: quantize GROUPS of weights jointly to entries in a
learned codebook.
✓ excellent quality at 2-3 bits
✗ the codebook lookup is not GEMM-friendly → slow kernels
→ good compression, poor speed. Useful for fitting, not for going fast.
TRAINED LOW-BIT [RESEARCH]
BitNet b1.58: train the model with ternary weights {-1, 0, 1}
✓ no quantization error — the model IS low-bit
✓ multiplication becomes addition
✗ requires training from scratch
✗ quality at scale is unproven relative to full-precision models
→ genuinely interesting; not something you deploy today
MIXED-BIT [EMERGING]
different bits per layer/channel based on measured sensitivity
✓ best quality per average bit
✗ complex kernels; irregular memory layoutThe honest summary: 4 bits is the practical floor for post-training quantization of a pretrained model. Below that you’re either accepting significant quality loss, using slow kernels, or training differently.
7. Accumulation, always#
Whatever the input precision, accumulate high:
FP4 × FP4 → FP32 accumulator
FP8 × FP8 → FP32
INT4 × INT4 → INT32
WHY: a K=8192 reduction accumulated in 4 bits is meaningless.
Error grows as ~sqrt(K) × quantization_step.
The hardware enforces this for tensor cores. If you write custom
kernels, preserve it. (Section III.11.)Also: some operations should never go below FP16/BF16 regardless of what the weights are:
softmax (attention and output)
layer norm / RMSNorm reductions
residual stream accumulation across many layers
RoPE sin/cos computation
sampling8. Production implications#
- Deploy FP8 now. It’s mature, well-supported, and gives ~2x.
- INT4 weight-only where you’re decode-bound, with validation (Section VII.06).
- FP4: evaluate when you have the hardware and engine support. Don’t commit architecture to it yet.
- Selective precision. Embeddings, LM head, norms, and MoE routers stay higher.
- Validate at Level 3 (long generations, reasoning) for any sub-8-bit format.
- Accumulate in FP32.
- Track the field. This is the fastest-moving area in inference, and the hardware roadmap drives it.
9. Common mistakes#
Assuming FP4 ≈ INT4. FP4 has tensor core support and helps compute; INT4 weight-only doesn’t.
Per-tensor scaling at 4 bits. Unusable. Block scaling is what makes it work.
Quantizing embeddings, the LM head, or MoE routers to 4 bits.
Validating with perplexity only. Low-bit’s cost shows in reasoning and long-form.
Expecting sub-4-bit codebook methods to be fast. They compress well and run slowly.
Committing to FP4 in a design before measuring it on your task.
Accumulating in low precision in a custom kernel.
10. Hands-on exercise#
A. The quality curve. Quantize a model to 8, 6, 5, 4, and 3 bits (using whatever methods your tooling supports). Plot perplexity and a reasoning benchmark vs bits. Where’s the cliff for each metric?
B. Block size effect. At 4 bits, vary the block/group size over {16, 32, 64, 128, per-tensor}. Plot quality vs effective bits-per-weight. Quantify what fine-grained scaling buys.
C. Selective precision. Quantize everything to 4 bits, then progressively exclude (embeddings, LM head, norms, attention). Measure quality at each step. Which exclusions matter most per byte?
D. Compute vs memory benefit. If you have Blackwell hardware, measure prefill and decode throughput at BF16, FP8, and FP4. Confirm FP4 helps both phases, unlike INT4 weight-only.
E. Accumulation. Compute a K=8192 dot product with FP4 inputs and FP4, FP16, and FP32 accumulation. Measure relative error against an FP64 reference.
F. Task sensitivity. For a 4-bit model, evaluate: knowledge recall, arithmetic, code generation, and 1000-token creative writing. Rank the degradation. Does it match the pattern from Section VII.06?
11. Interview questions#
- What’s the current production frontier for quantization precision?
- How does FP4 differ from INT4 weight-only in what it accelerates?
- What is block/microscaling and why does it make 4 bits viable?
- What would you exclude from 4-bit quantization, and why?
- Why does post-training quantization stop working below 4 bits?
- What is BitNet’s approach and how does it differ from quantization?
- Why must accumulation always be higher precision?
12. Further reading#
- [ESTABLISHED] Micikevicius et al., “FP8 Formats for Deep Learning” (2022)
- [EMERGING] Open Compute Project MX (microscaling) format specification
- [REFERENCE] NVIDIA Blackwell architecture documentation
- [EMERGING] Egiazarian et al., “AQLM” (2024); Tseng et al., “QuIP#” (2024)
- [RESEARCH] Ma et al., “The Era of 1-bit LLMs: BitNet b1.58” (2024)
- Next: 11 — Triton kernels