PidokuInfra

Reading a Spec Sheet

Foundations Beginner 45 min Difficulty 2/5 Topic 04 of 04

Prerequisites 01, 02

The idea in one minute#

A GPU data sheet lists dozens of numbers. For compute work, five decide almost everything: memory capacity, memory bandwidth, peak arithmetic rate (for the number format you will use), interconnect speed, and power. Learn to find these five and you can compare any two GPUs in a minute.

A picture#

flowchart TB
  Q1{"Does my data<br/>fit in memory?"} -->|"no"| X1["Need more VRAM<br/>or more GPUs"]
  Q1 -->|"yes"| Q2{"Is the work mostly<br/>reading big arrays?"}
  Q2 -->|"yes"| A1["Speed is set by<br/>memory bandwidth"]
  Q2 -->|"no, dense arithmetic"| A2["Speed is set by<br/>peak FLOP/s"]
  X1 --> A3["Then interconnect<br/>speed matters too"]
  class Q1,Q2 queue
  class A1 memory
  class A2 compute
  class X1,A3 io

How it really works#

The five numbers#

NumberUnitWhat it limits
Memory capacityGBWhether a workload fits at all
Memory bandwidthGB/s or TB/sSpeed of work that mostly reads/writes data (most LLM serving)
Peak arithmeticTFLOP/s, per number formatSpeed of dense arithmetic (training, large batches)
InterconnectGB/s (PCIe, NVLink)How fast data moves host↔GPU and GPU↔GPU
Power (TDP)WElectricity, cooling, how many fit in a rack

A FLOP is one floating-point operation (one add or one multiply). TFLOP/s is a trillion per second.

Worked example: three generations#

A100 (2020)H100 (2022)B200 (2024–25)
Memory40 or 80 GB HBM2e80 GB HBM3~180–192 GB HBM3e
Memory bandwidth~2.0 TB/s~3.35 TB/s~8 TB/s
FP16 tensor arithmetic312 TFLOP/s~990 TFLOP/s~2,250 TFLOP/s
Lowest fast formatFP16 / INT8FP8FP4
GPU-to-GPU linkNVLink 3, 600 GB/sNVLink 4, 900 GB/sNVLink 5, 1,800 GB/s
Power400 W700 W~1,000 W

Numbers are vendor figures, rounded. Use them to compare shapes, not to settle a purchase. The table stops at B200 because these three are well documented; the newer B300 and Rubin (288 GB each, Rubin with HBM4 at roughly 22 TB/s) are compared in VII.01.

Notice what grew and by how much from A100 to H100: arithmetic by about 3x, bandwidth by only 1.7x. Arithmetic has been growing faster than memory speed for years. That widening gap is why so much of modern GPU engineering is about memory.

Traps on the data sheet#

  • The headline FLOP number uses the most favourable format. A “1,979 TFLOPS” claim for the H100 is for FP8 (or FP16 with a sparsity feature). In FP32 the same chip does about 67. Always ask “in which format?”.
  • Peak is not sustained. Real kernels reach 50–80% of peak arithmetic on a good day.
  • Capacity is not all yours. The driver, the framework and fragmentation take a few GB.
  • Consumer vs data-center cards. An RTX 4090 has excellent arithmetic but 24 GB of memory, no NVLink, and a cooler designed for a desktop case, not a rack.

Product families (NVIDIA)#

FamilyExamplesBuilt for
GeForce RTX4090, 5090Gaming, single-workstation development
RTX workstationRTX 6000 AdaProfessional desktops, more memory
Data center, inferenceT4, L4, L40SEfficient serving, video
Data center, flagshipA100, H100, H200, B200, B300, RubinTraining and large-scale inference
EmbeddedJetsonRobots, edge devices

AMD (Instinct MI300 and MI400 series), Intel (Gaudi), Apple (M-series) and cloud vendors’ own chips are covered in module VII.

Code#

A spec sheet becomes useful the moment you compute with it. This program answers the two first-order questions for running a language model: does it fit, and how fast can one user possibly go?

Go
// specsheet.go — two questions you can answer from a data sheet alone.
package main

import "fmt"

type GPU struct {
	Name        string
	MemGB       float64
	BandwidthGB float64 // GB/s
}

func main() {
	gpus := []GPU{
		{"T4", 16, 320}, {"L4", 24, 300}, {"RTX 4090", 24, 1008},
		{"A100 80GB", 80, 2039}, {"H100", 80, 3350}, {"B200", 180, 8000},
	}
	const params, bytesPerParam = 8e9, 2.0 // an 8-billion-parameter model in FP16
	weightsGB := params * bytesPerParam / 1e9

	fmt.Printf("model weights: %.0f GB\n\n", weightsGB)
	fmt.Println("GPU          fits?   max tokens/s for one user")
	for _, g := range gpus {
		if weightsGB > g.MemGB*0.9 { // leave ~10% for everything else
			fmt.Printf("%-11s  no\n", g.Name)
			continue
		}
		// To produce one token the GPU must read every weight once.
		fmt.Printf("%-11s  yes     %.0f\n", g.Name, g.BandwidthGB/weightsGB)
	}
}

The second column uses only memory bandwidth — arithmetic rate does not appear. For one user at a time, an LLM is limited by how fast the GPU can read its own weights. Module V returns to this in detail.

Remember this#

  • Five numbers: capacity, bandwidth, arithmetic (per format), interconnect, power.
  • Capacity decides if it runs; bandwidth and arithmetic decide how fast.
  • Always ask which number format a FLOP figure refers to.
  • Arithmetic grows faster than bandwidth each generation, so memory is increasingly the limit.

Try it#

  1. Run specsheet.go. Change the model to 70 billion parameters. Which single GPUs can hold it? What if you store weights in 4 bits each (bytesPerParam = 0.5)?
  2. Look up the data sheet of any GPU you can access. Fill in the five numbers.
  3. Compute bandwidth ÷ arithmetic (bytes per FLOP) for the three GPUs in the table. What is the trend?

Check yourself#

  1. Which spec decides whether a model can run on a GPU at all?
  2. Why can two FLOP figures for the same GPU differ by 30x?
  3. Why does memory bandwidth matter more than FLOPs when serving one LLM user?

↑↓ navigate↵ openesc close