PidokuInfra

Tensor Cores and Number Formats

Basic Intermediate 55 min Difficulty 3/5 Topic 04 of 05

Prerequisites 01, I.04

The idea in one minute#

A tensor core is a circuit inside each SM that does one thing: multiply two small matrices and add the result to a third, in a single operation. Because it is built for that one job and works on small number formats (16, 8 or even 4 bits per number), it does 10–30x more arithmetic per second than the GPU’s general-purpose units.

Nearly all of a modern GPU’s headline FLOP figure comes from tensor cores, and you only get it if your numbers are in a format the tensor cores accept.

An analogy#

A general arithmetic unit is a skilled worker with a calculator: any operation, one at a time.

A tensor core is a stamping press. It produces one specific part — a small matrix product — in a single stroke. It cannot do anything else, but nothing matches it at that part. And the thinner the sheet metal (fewer bits per number), the more parts per stroke.

A picture#

flowchart LR
  subgraph IN["Inputs"]
    A["Matrix A<br/>small tile"]
    B["Matrix B<br/>small tile"]
    C["Matrix C<br/>accumulator"]
  end
  TC["Tensor core<br/>D = A x B + C<br/>one operation"]
  D["Matrix D"]
  A --> TC
  B --> TC
  C --> TC
  TC --> D
  class A,B,C,D memory
  class TC compute

How it really works#

Number formats#

A floating-point number stores a sign, an exponent and a fraction. Fewer bits means less precision and a smaller range — and half the memory, half the bandwidth, and simpler circuits.

FormatBitsRough precisionTypical use
FP646415–16 digitsScience, finance
FP32327 digitsClassic default
FP16163–4 digits, max ≈ 65,504AI inference and training
BF16162–3 digits, same range as FP32AI training (safer range)
FP881–2 digitsRecent AI inference/training
INT88256 levelsQuantized inference
FP4 / INT4416 levelsAggressively quantized inference

Neural networks tolerate low precision remarkably well: their weights are noisy estimates to begin with, so three digits are usually enough.

What the tensor core does to the FLOP count#

H100, vendor figures:

PathFormatPeak
General unitsFP64~34 TFLOP/s
General unitsFP32~67 TFLOP/s
Tensor coresFP16 / BF16~990 TFLOP/s
Tensor coresFP8 / INT8~1,980 TFLOP/s

Same chip. A 30x spread depending only on the number format and whether the work is expressed as matrix multiplication.

The conditions for using them#

  1. The operation must be a matrix multiply (or a convolution, which libraries turn into one). An elementwise a + b never touches a tensor core.
  2. The data must be in a supported format. FP32 inputs on most GPUs fall back to the slow path (some GPUs offer a reduced-precision “TF32” mode as a middle ground).
  3. Dimensions should be friendly, typically multiples of 8, so tiles fill completely.

You rarely call a tensor core yourself. Libraries (cuBLAS, cuDNN) and frameworks choose them automatically when the conditions hold. Your job is to make the conditions hold.

Mixed precision#

Low precision is safe for the bulk multiplies but risky for a few steps, such as summing thousands of small numbers or computing exp() of large values. Mixed precision runs the heavy matrix work in FP16/BF16/FP8 and keeps the fragile steps in FP32. It gets nearly all the speed and nearly all the accuracy.

Code#

Go has no 16-bit float type, which makes it a good place to see what the format does. This program converts FP32 values to FP16 bits and back, by hand, and shows the two ways it bites: rounding and overflow.

Go
// half.go — what FP16 keeps and what it loses.
package main

import (
	"fmt"
	"math"
)

// roundTrip converts a value to IEEE 754 half precision (round to nearest) and back.
func roundTrip(f float32) float32 {
	if math.IsNaN(float64(f)) {
		return f
	}
	const maxHalf = 65504
	if f > maxHalf {
		return float32(math.Inf(1))
	}
	if f < -maxHalf {
		return float32(math.Inf(-1))
	}
	if f == 0 {
		return 0
	}
	// FP16 has an 11-bit significand: keep 11 significant bits of the value.
	exp := math.Floor(math.Log2(math.Abs(float64(f))))
	exp = math.Max(exp, -14) // below this, half precision loses bits (subnormals)
	step := math.Pow(2, exp-10)
	return float32(math.Round(float64(f)/step) * step)
}

func main() {
	for _, v := range []float32{1.0, 0.1, 3.14159265, 1000.123, 65504, 70000, 1e-7} {
		h := roundTrip(v)
		fmt.Printf("fp32 %-12g -> fp16 %-12g  error %.2e\n", v, h, math.Abs(float64(v-h)))
	}

	// The classic trap: adding a small number to a big one.
	sum32, sum16 := float32(2048), float32(2048)
	for i := 0; i < 1000; i++ {
		sum32 += 0.5
		sum16 = roundTrip(sum16 + 0.5)
	}
	fmt.Printf("\n2048 + 1000 x 0.5: fp32 = %g, fp16 = %g\n", sum32, sum16)
}

The last line is the important one. In FP16, numbers near 2048 are spaced 2 apart, so adding 0.5 does nothing — a thousand times. This is exactly why accumulations stay in FP32 under mixed precision.

Remember this#

  • Tensor cores do D = A × B + C on small matrix tiles in one operation.
  • They provide most of a GPU’s peak arithmetic — but only for matmul, in low-precision formats.
  • Fewer bits = less memory, less bandwidth, more FLOP/s, less accuracy.
  • Mixed precision: heavy work in low precision, fragile steps in FP32.

Try it#

  1. Run half.go. Which input overflows? Which loses the most relative accuracy?
  2. Change the accumulation to start at 1.0 instead of 2048. Does FP16 now keep up? Why?
  3. An 8-billion-parameter model: how many GB in FP32, FP16, FP8 and INT4?

Check yourself#

  1. What single operation does a tensor core perform?
  2. Give two conditions for work to run on tensor cores.
  3. Why is summation kept in FP32 under mixed precision?

↑↓ navigate↵ openesc close