The idea in one minute#
NVIDIA GPUs are the default for AI, but not the only option. Alternatives fall into a few families: other GPUs (AMD, Intel), cloud vendors’ own chips (Google TPU, AWS, Microsoft, Meta), inference-specialized designs (Groq, Cerebras, SambaNova), and unified-memory systems-on-chip (Apple M-series). Each makes a different bet about which constraint matters most.
You can evaluate any of them with the tools you already have: capacity, bandwidth, arithmetic, interconnect, power — plus one new factor that often decides the matter: software.
A picture#
flowchart TB
Q{"What limits you?"}
Q -->|"Software risk, flexibility"| NV["NVIDIA GPU<br/>largest ecosystem"]
Q -->|"Memory capacity per dollar"| AMD["AMD Instinct<br/>large HBM, ROCm"]
Q -->|"Cost at cloud scale"| ASIC["Cloud ASICs<br/>TPU, Trainium, Maia"]
Q -->|"Tokens per second per user"| SPEC["Specialized inference<br/>Groq, Cerebras"]
Q -->|"Local, private, low power"| APL["Unified memory SoC<br/>Apple M-series"]
class Q queue
class NV compute
class AMD,ASIC io
class SPEC memory
class APL neutralHow it really works#
The families#
| Family | Examples | The bet | Trade-off |
|---|---|---|---|
| General GPUs | NVIDIA; AMD Instinct (MI300/MI350 series; MI455X in “Helios” racks); Intel | Flexibility: any model, training and inference | Not optimal for any single workload |
| Cloud ASICs | Google TPU (Ironwood); AWS Trainium3 / Inferentia; Microsoft Maia 200; Meta MTIA | Own the design, cut cost at huge scale for known workloads | Only in that cloud; software tied to it |
| Wafer-scale | Cerebras | One enormous chip: massive bandwidth, no inter-chip communication | Exotic system; model support is narrower |
| Deterministic inference chips | Groq LPU — now also sold by NVIDIA as Groq 3 LPX racks | Very high tokens/s per user via on-chip memory and a static schedule | Limited memory per chip; many chips per model |
| Reconfigurable dataflow | SambaNova | Hardware shaped to the model’s data flow | Niche tooling |
| Unified-memory SoCs | Apple M-series; some laptop/edge chips | CPU and GPU share one large memory pool: no PCIe copies | Lower bandwidth and FLOPs than data-center parts |
The field in October 2026#
Details date quickly; the pattern they show is the point.
- AMD went rack-scale. The Helios rack (72 Instinct MI455X GPUs, each with a reported 432 GB of HBM4) began shipping to large customers in the second half of 2026 — AMD’s first answer to NVIDIA’s rack systems, and still a bet on memory per device.
- Cloud ASICs are specializing by workload. Google’s seventh-generation TPU (Ironwood) is generally available and aimed at inference, and its announced eighth generation splits into separate training and inference chips. AWS’s Trainium3 is generally available; Microsoft’s Maia 200 (January 2026) is described as an inference accelerator.
- The specialist was absorbed by the incumbent. NVIDIA licensed Groq’s inference technology in December 2025 (a reported $20 billion) and announced the Groq 3 LPX rack at GTC 2026 as a low-latency companion to its GPU racks. The lever Groq attacked — per-user token speed — turned out to matter enough for the GPU vendor to sell it.
- Wafer-scale kept going. Cerebras runs production inference for large customers and presented its next systems at Hot Chips 2026.
Every one of these is a vendor claim until you run your model on it.
Why “FLOPs per dollar” is not enough#
A faster, cheaper chip is worthless if your model does not run on it, or runs through an immature kernel that reaches 20% of the hardware’s potential. CUDA’s advantage is fifteen years of tuned libraries (III.05), and every framework, optimization and new research technique lands on CUDA first.
Questions to ask of any alternative:
- Does my exact model and serving stack run on it today, at full speed?
- Are the optimizations I rely on (quantization formats, attention kernels, batching server) supported?
- How do I debug and profile it?
- Can I get capacity, and from more than one supplier?
For large, stable workloads the answers can justify the switch — that is why the largest operators build their own chips. For a small team iterating quickly, software maturity usually outweighs hardware advantages.
Where alternatives genuinely win#
- AMD GPUs frequently offer more memory per dollar. For capacity-bound inference (V.03) that is exactly the right lever, and the major inference servers support them.
- Specialized inference chips deliver several times the per-user token rate of a GPU. They attack the batch-1 memory-bound limit (II.03) directly, by putting weights in much faster memory. Valuable when latency per user is the product.
- Cloud ASICs win on cost for the cloud’s own high-volume models, and are offered to customers at attractive prices for supported frameworks.
- Apple silicon makes large models usable on a laptop because the GPU can address all of system memory. Bandwidth (roughly 100–800 GB/s depending on the chip) limits speed, but a 64 GB model simply runs, with no PCIe in the way.
CPUs still count#
A modern server CPU with wide vector instructions and many memory channels runs small models, embeddings and classical ML perfectly well, with no accelerator to schedule. From IV.01: if the workload is memory-bound and small, the CPU’s bandwidth may be enough.
Code#
A comparison harness using the course’s own model: fit, memory-bound speed, and cost.
// compare.go — evaluate any accelerator with the same five questions.
package main
import "fmt"
type Accel struct {
Name string
MemGB, BandwidthGBs float64
DollarsPerHour float64
Mature bool // does your stack run on it today, unmodified?
}
func main() {
const modelGB = 40.0 // e.g. a 70B model at ~4.5 bits per weight
candidates := []Accel{
{"Data-center GPU A", 80, 3350, 2.50, true},
{"Data-center GPU B (more memory)", 192, 5300, 2.20, true},
{"Laptop SoC, unified memory", 128, 546, 0.10, true},
{"Server CPU, 12 memory channels", 512, 300, 0.80, true},
{"Specialized inference system", 64, 20000, 6.00, false},
}
fmt.Println("accelerator fits tok/s (1 user) $ per M tokens (1 user) software")
for _, a := range candidates {
if modelGB > a.MemGB*0.9 {
fmt.Printf("%-34s no\n", a.Name)
continue
}
tps := a.BandwidthGBs / modelGB
cost := a.DollarsPerHour / (tps * 3600) * 1e6
sw := "ready"
if !a.Mature {
sw = "verify first"
}
fmt.Printf("%-34s yes %12.0f %22.2f %s\n", a.Name, tps, cost, sw)
}
fmt.Println("\nSingle-user figures. Batching changes the cost column by 10-100x on GPUs.")
}All rows are illustrative. Note the last line of output: single-user cost flatters hardware that cannot batch and penalizes hardware that can. Decide which case is yours before comparing.
Remember this#
- Alternatives differ in which constraint they attack: capacity, bandwidth, per-user speed, cost at scale, or locality.
- Software maturity is frequently the deciding factor.
- Evaluate with capacity, bandwidth, arithmetic, interconnect, power — then cost per unit of your work.
- CPUs and unified-memory laptops are legitimate inference hardware for the right model size.
Try it#
- Run
compare.go. Add aBatchfield and a compute limit, and recompute cost per million tokens at batch 32. - Pick one alternative and find out whether a model you use runs on it with your serving stack. How long did it take you to find out? That time is part of the cost.
- Why does unified memory remove a whole class of problems from module III?
Check yourself#
- Name three families of non-NVIDIA accelerators and each one’s bet.
- Why can a slower chip with mature software beat a faster one without?
- Which hardware lever do specialized inference chips attack?