PidokuInfra

Low-Bit Inference: FP4 and Beyond

Expert Advanced 1h Difficulty 4/5 Topic 10 of 12

Prerequisites VII.05, VII.06


1. Where the field is#

FORMAT   BITS   STATUS           NOTES
FP32     32     legacy           only for accumulation and a few sensitive ops
BF16     16     ESTABLISHED      the safe default
FP16     16     ESTABLISHED      range issues at depth (Section III.11)
FP8      8      ESTABLISHED      production standard on Hopper+
INT8     8      ESTABLISHED      production standard elsewhere
INT4     4      ESTABLISHED      weight-only; production-common
FP6      6      EMERGING         Blackwell hardware support
FP4      4      EMERGING         Blackwell (MXFP4/NVFP4); training and inference
INT3     3      RESEARCH         significant quality loss
INT2     2      RESEARCH         needs learned codebooks
1.58-bit —      RESEARCH         BitNet: trained low-bit, not quantized

The frontier is at 4 bits with hardware support. Below that, the techniques change character: you stop quantizing a pretrained model and start training for low precision.


2. What Blackwell changes#

Blackwell adds NATIVE tensor core support for FP4 and FP6, with
microscaling (MX) formats.

MXFP4: 4-bit elements + a shared 8-bit exponent scale per block of 32
  → effective bits: 4 + 8/32 = 4.25 bits/value
  → the block-level scale is what makes 4 bits viable

NVFP4: NVIDIA's variant, block size 16 with an FP8 scale,
       plus a per-tensor FP32 scale
  → effective bits: 4 + 8/16 = 4.5
  → finer blocks → better accuracy

THROUGHPUT (approximate, per NVIDIA's Blackwell figures)
  BF16:  1.0x
  FP8:   2.0x
  FP4:   4.0x

MEMORY
  FP4 weights are 1/4 of BF16
  → a 405B model at FP4 is ~203 GB (fits on 3 GPUs instead of 11)

The critical difference from INT4 weight-only (Section VII.06): FP4 has tensor core support, so it accelerates COMPUTE as well as memory traffic. That means it helps prefill too — which INT4 weight-only doesn’t.

              Decode    Prefill    Memory
INT4 W4A16    2.5-3.5x  ~1.0x      0.28x     no compute benefit
FP4 (W4A4)    3.0-4.0x  ~3.5x      0.28x     compute AND memory

3. Why block scaling makes 4 bits work#

Recall Section III.12: quantization error is bounded by the SCALE,
and the scale is determined by the RANGE of the group.

  per-tensor scale, 4 bits (16 levels), range determined by the
  tensor's max → catastrophic for small values
  
  per-block scale (32 values), 4 bits → the range is only that
  block's range, which is much narrower
  → error relative to the block's own magnitude, not the tensor's
EFFECTIVE BITS COMPARISON
  INT4, per-tensor:           4.0 bits    unusable
  INT4, group 128, FP16 scale: 4.125      workable (Section VII.06)
  MXFP4, block 32, E8 scale:   4.25       better
  NVFP4, block 16, FP8 scale:  4.5        best of these

The extra 0.25-0.5 bits buys most of the accuracy.

This is the same lesson as Section VII.06’s group-size discussion, taken to its conclusion: fine-grained scaling is what makes aggressive quantization viable, and the hardware now supports it natively.


4. What’s needed to use FP4 well#

1. QUANTIZATION-AWARE PREPARATION
   Naive round-to-nearest at 4 bits loses too much.
   Need: GPTQ/AWQ-style calibration (Section VII.03), or
         quantization-aware fine-tuning, or
         training in low precision from the start

2. SELECTIVE PRECISION
   Not everything should be FP4:
     ✓ FFN weights (the bulk)
     ~ attention projections (more sensitive)
     ✗ embeddings, LM head, norms, router weights (MoE)
     ✗ accumulators (FP32, always)

3. KERNEL SUPPORT
   Blackwell tensor cores + CUTLASS/cuBLAS support for MXFP4/NVFP4
   Engine support is arriving

4. VALIDATION
   The full Level 1-4 (Section IV.12), with extra attention to
   long generations and reasoning tasks — where low-bit hurts most
   (Section VII.06)

5. The honest assessment#

WHAT'S SOLID
  ✓ FP8 is production-standard and well-validated
  ✓ INT4 weight-only is production-common for decode-bound workloads
  ✓ block/microscaling is the right mechanism
  ✓ Blackwell's hardware support is real

WHAT'S NOT YET SETTLED
  ? quality at FP4 for demanding tasks (reasoning, code, long-form)
  ? whether post-training FP4 is sufficient or QAT is required
  ? engine support maturity
  ? which format (MXFP4 vs NVFP4 vs others) becomes standard
  ? whether FP6 is the practical sweet spot rather than FP4

WHAT TO DO NOW
  → deploy FP8 today; it's mature
  → INT4 weight-only where decode-bound and validated
  → watch FP4; re-evaluate when you have Blackwell and engine support
  → don't build architecture assuming FP4 works for your task until
    you've measured it

The pattern from FP8’s adoption is instructive: hardware support arrived, then a year of validation work, then it became standard. FP4 is roughly at the start of that curve.


6. Sub-4-bit: a different problem#

Below 4 bits, post-training quantization stops working well and the
approaches change:

VECTOR / CODEBOOK QUANTIZATION  [EMERGING→RESEARCH]
  AQLM, QuIP#: quantize GROUPS of weights jointly to entries in a
  learned codebook.
  ✓ excellent quality at 2-3 bits
  ✗ the codebook lookup is not GEMM-friendly → slow kernels
  → good compression, poor speed. Useful for fitting, not for going fast.

TRAINED LOW-BIT  [RESEARCH]
  BitNet b1.58: train the model with ternary weights {-1, 0, 1}
  ✓ no quantization error — the model IS low-bit
  ✓ multiplication becomes addition
  ✗ requires training from scratch
  ✗ quality at scale is unproven relative to full-precision models
  → genuinely interesting; not something you deploy today

MIXED-BIT  [EMERGING]
  different bits per layer/channel based on measured sensitivity
  ✓ best quality per average bit
  ✗ complex kernels; irregular memory layout

The honest summary: 4 bits is the practical floor for post-training quantization of a pretrained model. Below that you’re either accepting significant quality loss, using slow kernels, or training differently.


7. Accumulation, always#

Whatever the input precision, accumulate high:

  FP4 × FP4 → FP32 accumulator
  FP8 × FP8 → FP32
  INT4 × INT4 → INT32

WHY: a K=8192 reduction accumulated in 4 bits is meaningless.
     Error grows as ~sqrt(K) × quantization_step.

The hardware enforces this for tensor cores. If you write custom
kernels, preserve it. (Section III.11.)

Also: some operations should never go below FP16/BF16 regardless of what the weights are:

  softmax (attention and output)
  layer norm / RMSNorm reductions
  residual stream accumulation across many layers
  RoPE sin/cos computation
  sampling

8. Production implications#

  • Deploy FP8 now. It’s mature, well-supported, and gives ~2x.
  • INT4 weight-only where you’re decode-bound, with validation (Section VII.06).
  • FP4: evaluate when you have the hardware and engine support. Don’t commit architecture to it yet.
  • Selective precision. Embeddings, LM head, norms, and MoE routers stay higher.
  • Validate at Level 3 (long generations, reasoning) for any sub-8-bit format.
  • Accumulate in FP32.
  • Track the field. This is the fastest-moving area in inference, and the hardware roadmap drives it.

9. Common mistakes#

Assuming FP4 ≈ INT4. FP4 has tensor core support and helps compute; INT4 weight-only doesn’t.

Per-tensor scaling at 4 bits. Unusable. Block scaling is what makes it work.

Quantizing embeddings, the LM head, or MoE routers to 4 bits.

Validating with perplexity only. Low-bit’s cost shows in reasoning and long-form.

Expecting sub-4-bit codebook methods to be fast. They compress well and run slowly.

Committing to FP4 in a design before measuring it on your task.

Accumulating in low precision in a custom kernel.


10. Hands-on exercise#

A. The quality curve. Quantize a model to 8, 6, 5, 4, and 3 bits (using whatever methods your tooling supports). Plot perplexity and a reasoning benchmark vs bits. Where’s the cliff for each metric?

B. Block size effect. At 4 bits, vary the block/group size over {16, 32, 64, 128, per-tensor}. Plot quality vs effective bits-per-weight. Quantify what fine-grained scaling buys.

C. Selective precision. Quantize everything to 4 bits, then progressively exclude (embeddings, LM head, norms, attention). Measure quality at each step. Which exclusions matter most per byte?

D. Compute vs memory benefit. If you have Blackwell hardware, measure prefill and decode throughput at BF16, FP8, and FP4. Confirm FP4 helps both phases, unlike INT4 weight-only.

E. Accumulation. Compute a K=8192 dot product with FP4 inputs and FP4, FP16, and FP32 accumulation. Measure relative error against an FP64 reference.

F. Task sensitivity. For a 4-bit model, evaluate: knowledge recall, arithmetic, code generation, and 1000-token creative writing. Rank the degradation. Does it match the pattern from Section VII.06?


11. Interview questions#

  1. What’s the current production frontier for quantization precision?
  2. How does FP4 differ from INT4 weight-only in what it accelerates?
  3. What is block/microscaling and why does it make 4 bits viable?
  4. What would you exclude from 4-bit quantization, and why?
  5. Why does post-training quantization stop working below 4 bits?
  6. What is BitNet’s approach and how does it differ from quantization?
  7. Why must accumulation always be higher precision?

12. Further reading#

  • [ESTABLISHED] Micikevicius et al., “FP8 Formats for Deep Learning” (2022)
  • [EMERGING] Open Compute Project MX (microscaling) format specification
  • [REFERENCE] NVIDIA Blackwell architecture documentation
  • [EMERGING] Egiazarian et al., “AQLM” (2024); Tseng et al., “QuIP#” (2024)
  • [RESEARCH] Ma et al., “The Era of 1-bit LLMs: BitNet b1.58” (2024)
  • Next: 11 — Triton kernels

↑↓ navigate↵ openesc close