PidokuInfra

GPU Generations

Expert Advanced 45 min Difficulty 3/5 Topic 01 of 04

Prerequisites I.04, II.03, II.04

The idea in one minute#

NVIDIA now releases a new data-center GPU generation roughly every year. Each one moves the same handful of levers: more memory capacity, more memory bandwidth, a smaller number format on the tensor cores, a faster interconnect, and more power. If you track those five levers you can read any future announcement in a minute — including ones that do not exist yet.

A picture#

flowchart LR
  V["Volta 2017<br/>first tensor cores"] --> A["Ampere 2020<br/>A100, MIG"]
  A --> H["Hopper 2022<br/>H100, FP8"]
  H --> B["Blackwell 2024-25<br/>B200, B300, FP4, liquid racks"]
  B --> R["Rubin 2026<br/>HBM4, shipping"]
  R --> N["Rubin Ultra 2027, Feynman 2028<br/>announced"]
  class V,A neutral
  class H,B,R compute
  class N io

How it really works#

The five levers across generations#

A100 (Ampere)H100 (Hopper)B200 (Blackwell)B300 (Blackwell Ultra)Rubin
First shipments202020222024–252025August 2026
Memory80 GB HBM2e80 GB HBM3~180–192 GB HBM3e288 GB HBM3e288 GB HBM4
Bandwidth~2 TB/s~3.35 TB/s~8 TB/s~8 TB/s~22 TB/s
Smallest fast formatFP16 / INT8FP8FP4FP4FP4
Peak tensor arithmetic (smallest format)~0.6 PFLOP/s~2 PFLOP/s~9–10 PFLOP/s~15 PFLOP/s~50 PFLOP/s (NVIDIA’s inference figure)
NVLink per GPU600 GB/s900 GB/s1,800 GB/s1,800 GB/s3,600 GB/s
Power per GPU400 W700 W~1,000 W~1,400 W~1,800–2,300 W (reported)

Figures are vendor numbers, rounded. Rubin began production shipments in August 2026 and independent measurements are still scarce, so treat its column as claims to be verified. Variants exist within each generation: H200 is an H100 with more and faster memory, and Blackwell Ultra is a mid-cycle part with more memory and more FP4 arithmetic at the same bandwidth — a capacity upgrade, not a decode-speed upgrade.

Where the line stands (checked 3 October 2026)#

  • Rubin is shipping, only as part of liquid-cooled racks: the Vera Rubin NVL72 pairs 72 Rubin GPUs with 36 of NVIDIA’s own Vera CPUs. There is no air-cooled Rubin.
  • Blackwell and Blackwell Ultra are what most new capacity still runs on; Hopper (H100, H200) remains the largest installed base and the best-documented target.
  • Announced, not shipping: Rubin Ultra (second half of 2027: four dies per package, HBM4e, racks of roughly 600 kW) and Feynman (2028). Announcements slip; plan on what you can rent.
  • An inference-specific sibling. At GTC 2026 NVIDIA added a non-GPU rack, the Groq 3 LPX, built on technology licensed from Groq, to sit beside Rubin racks for low-latency token generation (lesson 02). The inference-oriented “Rubin CPX” GPU announced in 2025 is no longer on the public roadmap.

Reading the table#

  • Arithmetic grows by shrinking the format. Much of each generation’s headline FLOP gain comes from supporting a smaller number format, not from doing the same arithmetic faster. The gain is only real for you if your model runs acceptably at that precision (V.02).
  • Bandwidth is catching up. For years arithmetic outgrew bandwidth, pushing more workloads into the memory-bound regime. The jumps to HBM3e and HBM4 are large specifically because LLM inference made bandwidth the bottleneck people pay to remove.
  • Capacity grows in steps. It follows HBM stack sizes, and it is the lever that decides whether a model needs one GPU or two.
  • Power goes up every time. Performance per watt improves, but absolute watts per GPU rise, which is why cooling and electricity now shape data-center design (lesson 03).

The unit of sale is changing#

Until recently you bought GPUs and built servers. Flagship parts are now designed, sold and installed as whole racks: dozens of GPUs, CPUs, NVLink switches, liquid cooling and power delivery engineered as one machine. The “GPU” a large customer buys is increasingly a rack.

How to evaluate a new generation for your workload#

  1. Compute your workload’s regime on the current hardware (IV.01).
  2. Memory-bound? Your speedup tracks the bandwidth ratio, not the FLOP ratio.
  3. Compute-bound? It tracks the FLOP ratio in the format you can actually use.
  4. Capacity-bound? It tracks the memory ratio — and may let you drop a multi-GPU split.
  5. Divide by the price ratio. A generation that is 2.5x faster at 2x the price is a modest improvement in cost per token, whatever the launch slides say.

Older generations do not become useless. Fully depreciated A100s serving small models can beat new flagships on cost per token.

Code#

Go
// upgrade.go — what a new generation is worth depends on your regime.
package main

import "fmt"

type GPU struct {
	Name                       string
	MemGB, BandwidthTB, PFLOPs float64
	PricePerHour               float64
}

func main() {
	old := GPU{"H100", 80, 3.35, 2.0, 2.50}
	gens := []GPU{
		{"B200", 180, 8.0, 9.0, 5.00},
		{"B300", 288, 8.0, 15.0, 6.00},
		{"Rubin (vendor figures)", 288, 22, 50, 10.00},
	}
	fmt.Println("vs H100               memory-bound  compute-bound  capacity   price  $/work mem-bound  $/work compute-bound")
	for _, g := range gens {
		bw, fl, mem, price := g.BandwidthTB/old.BandwidthTB, g.PFLOPs/old.PFLOPs, g.MemGB/old.MemGB, g.PricePerHour/old.PricePerHour
		fmt.Printf("%-22s %10.1fx  %12.1fx  %7.1fx  %5.1fx  %15.2fx  %19.2fx\n",
			g.Name, bw, fl, mem, price, price/bw, price/fl)
	}
	fmt.Println("\n$/work below 1.00x means cheaper per unit of work than the H100.")
}

Prices are placeholders, and the Rubin row uses vendor figures. Replace them with real quotes and measurements before deciding anything — the structure of the calculation is what to keep. Notice the B300 row: more memory and more arithmetic than a B200, and no faster at memory-bound work.

Remember this#

  • Five levers per generation: capacity, bandwidth, smallest format, interconnect, power.
  • Headline FLOP gains often come from a smaller format; they are yours only if you can use it.
  • Your speedup follows the lever that matches your regime.
  • Always convert to cost per unit of work.

Try it#

  1. Run upgrade.go with current cloud prices for two GPUs you can actually rent.
  2. Your workload is batch-1 decoding of a 16 GB model. Which column predicts your speedup?
  3. A 100 GB model needs two H100s with tensor parallelism. What changes on a 180 GB GPU, beyond raw speed? (Use VI.02.)

Check yourself#

  1. Name the five levers.
  2. Why can a “5x more FLOPs” generation give you only 2x?
  3. Why might an old GPU have the best cost per token?

↑↓ navigate↵ openesc close