1. Why look beyond NVIDIA#
REASONS
cost NVIDIA's margins are high; alternatives are cheaper per FLOP
availability supply constraints are real
lock-in a single-vendor dependency is a business risk
specialization some designs are genuinely better for inference specifically
REASONS NOT TO
software CUDA's ecosystem is a decade ahead
maturity kernels, engines, and tooling lag
effort porting and maintaining a second stack is expensiveThe honest position: NVIDIA is the default for good reasons, and the alternatives are viable for specific situations. This file is about identifying those situations.
2. The landscape#
ACCELERATOR TYPE MATURITY BEST FOR
AMD Instinct GPU good large-memory inference; ROCm
(MI300X-MI355X; MI455X in Helios racks)
Intel Gaudi 2/3 ASIC moderate cost-sensitive training+inference
Google TPU ASIC mature* if you're on GCP / using JAX
AWS Trainium / ASIC moderate if you're on AWS, cost-sensitive
Inferentia2
Microsoft Maia ASIC internal Azure's own inference workloads
Cerebras WSE wafer-scale niche very low-latency inference
Groq LPU ASIC niche** extremely low-latency decode
SambaNova RDU niche
Apple M-series integrated good on-device, unified memory
Qualcomm, ARM mobile NPU good on-device
* mature within Google's ecosystem; less so outside it
** licensed by NVIDIA in December 2025; also sold as NVIDIA's Groq 3 LPX rackSTATE OF THE FIELD — checked 3 October 2026 (vendor claims unless measured)
AMD Helios: first rack-scale system, 72 x MI455X, a reported
432 GB HBM4 per GPU; shipping to large customers from H2 2026
Google 7th-gen TPU (Ironwood) generally available, positioned for
inference; 8th generation announced as separate training
and inference chips
AWS Trainium3 generally available; Trainium4 announced
Microsoft Maia 200 (January 2026), described as an inference chip
NVIDIA licensed Groq's LPU technology; Groq 3 LPX rack announced
at GTC 2026 alongside Rubin
Cerebras in production for large customers; next systems shown at
Hot Chips 2026
THE PATTERN: everyone moved to rack-scale systems, and nearly
everyone now builds something inference-specific. The evaluation
method in this file has not changed.3. AMD MI300X — the closest substitute#
SPECS (MI300X)
memory: 192 GB HBM3 ← the headline advantage
bandwidth: 5.3 TB/s ← higher than H100's 3.35
FP16: ~1,300 TF
FP8: ~2,600 TF
Infinity Fabric between GPUs
RIDGE POINT: 1300/5.3 = 245 (vs H100's 296)
→ slightly less memory-bound than H100 for the same workload
WHY IT'S INTERESTING FOR INFERENCE
✓ 192 GB means a 70B model at FP16 fits on ONE GPU (140 GB + KV)
→ no tensor parallelism needed → no AllReduce overhead
✓ higher bandwidth than H100 → faster decode
✓ generally lower cost per unit
THE SOFTWARE QUESTION
ROCm has matured substantially. vLLM, SGLang, and PyTorch all
support it. FlashAttention has ROCm implementations.
Triton compiles to AMD (Section XIII.11) — a portability advantage.
✗ but: kernel coverage lags, new model architectures are supported
later, and debugging tools are less mature
✗ FP8 support and its ecosystem trail Hopper'sMI300X’s 192 GB is a genuine architectural advantage for inference. Fitting a 70B model on one GPU eliminates tensor parallelism entirely, which removes 160 AllReduces per token (Section IX.03).
When to evaluate it: if you’re serving models in the 30-100B range, cost matters, and you can tolerate a software maturity gap.
4. Google TPU#
DESIGN PHILOSOPHY
systolic array for matmul, large on-chip memory, high-bandwidth
interconnect (ICI) designed for large-scale training
TPU v5e / v5p / v6 (Trillium) / v7 (Ironwood)
memory: 16-95 GB HBM per chip through v5p; more on newer
generations — check the current GCP documentation
interconnect: very high-bandwidth mesh — better multi-chip scaling
than PCIe-based systems
FOR INFERENCE
✓ excellent price/performance within GCP
✓ the interconnect makes large-model serving straightforward
✓ JAX/XLA is a mature, well-designed stack
✗ requires the JAX/XLA programming model (not PyTorch-native,
though PyTorch/XLA exists)
✗ static shapes are required → dynamic-shape LLM serving needs
bucketing
✗ available only on GCP
WHEN
→ you're already on GCP
→ you use JAX
→ you can accept the static-shape constraintThe static-shape requirement is the practical friction for LLM serving: Section IV.11’s dynamic shape problem is harder when the compiler requires static shapes. Bucketing works but costs padding.
5. AWS Inferentia2 / Trainium#
DESIGN
purpose-built for inference (Inferentia) and training (Trainium)
NeuronCore-v2, 32 GB HBM per chip, high inter-chip bandwidth
FOR INFERENCE
✓ significantly cheaper per token than GPU instances for supported
models, within AWS
✓ Neuron SDK integrates with PyTorch; vLLM has a Neuron backend
✗ model support is narrower — check YOUR model
✗ compilation is required (AOT, like TensorRT — Section VII.09)
✗ AWS-only
WHEN
→ you're on AWS, cost-sensitive, and your model is supported
→ the model set is stable (compilation is AOT)6. Intel Gaudi#
DESIGN
Gaudi 2: 96 GB HBM2e, 2.45 TB/s; Gaudi 3: 128 GB HBM2e, 3.7 TB/s
integrated RoCE ports on-chip → scale-out without separate NICs
FOR INFERENCE
✓ competitive price/performance
✓ integrated networking simplifies multi-node
✓ vLLM and PyTorch support exists
✗ ecosystem maturity trails
✗ smaller installed base → fewer people have solved your problem
WHEN
→ cost is the dominant factor and you can invest in the software7. The low-latency specialists#
GROQ LPU
DESIGN: deterministic execution, all weights in on-chip SRAM
(no HBM), software-scheduled
→ extremely low decode latency (hundreds of tokens/sec per user)
✓ latency that GPU architectures cannot match at batch 1
✗ SRAM capacity is small → a large model needs MANY chips
✗ the economics depend entirely on utilization across those chips
→ a fundamentally different point in the design space
CEREBRAS WSE
DESIGN: wafer-scale chip, enormous on-chip memory
→ very low latency, high single-model throughput
✗ availability, cost, and ecosystem are niche
THE COMMON THREAD
both attack the memory-bandwidth problem by putting weights in
SRAM instead of HBM.
→ SRAM bandwidth is ~10-100x HBM
→ but SRAM capacity is ~1000x smaller
→ so you need many chips, and the economics turn on whether you
can keep them all busyThese are interesting because they attack the actual bottleneck differently. Section I.07’s memory wall is a consequence of using DRAM; using SRAM instead removes it and creates a capacity problem instead.
Whether that trade wins depends on utilization — which is a business question, not a technical one.
8. On-device#
APPLE M-SERIES
unified memory: 400-800 GB/s CPU-GPU-NPU shared
→ no PCIe transfer; the whole memory is "GPU memory"
→ a 70B model at INT4 (35 GB) runs on a 64 GB Mac
✓ genuinely capable for local inference
✗ bandwidth is far below datacenter GPUs
QUALCOMM / ARM NPUs
→ phone-scale models (1-8B quantized)
→ INT4/INT8, very power-constrained
WHY IT MATTERS FOR SERVER-SIDE ENGINEERS
→ on-device handles some of your traffic if you build for it
→ the techniques are the same (quantization, KV management),
with tighter constraints
→ hybrid architectures (small model on-device, escalate to server)
are a real cost lever (Section XII.09's cascade, across the
network boundary)9. How to evaluate an alternative#
1. DOES IT RUN YOUR MODEL?
→ check the specific architecture, not "supports LLMs"
→ check the specific features (GQA? MoE? long context?)
2. WHAT'S THE COST PER MILLION TOKENS?
→ Section XI.03's methodology, on THEIR hardware
→ not $/FLOP or $/GB — cost per unit of YOUR work
3. WHAT'S THE SOFTWARE EFFORT?
→ porting, debugging, maintaining a second stack
→ this is usually the dominant cost, and it's ongoing
4. WHAT'S THE ECOSYSTEM RISK?
→ will your next model be supported? How quickly?
→ how many people have solved the problems you'll hit?
5. WHAT'S THE EXIT?
→ if it doesn't work out, how hard is it to move back?
→ keep your stack portable (OpenAI API, standard checkpoints)
THE PRAGMATIC ANSWER FOR MOST TEAMS
evaluate alternatives when: cost is a first-order concern AND
your model set is stable AND you have engineering capacity.
Otherwise the software cost exceeds the hardware saving.10. The portability layer#
WHAT MAKES A MULTI-VENDOR STACK POSSIBLE
✓ OpenAI-compatible API (clients don't change)
✓ standard checkpoints (safetensors, not vendor formats)
✓ PyTorch as the model definition (most backends support it)
✓ Triton for custom kernels (compiles to CUDA and ROCm)
✓ engine abstraction (vLLM supports CUDA, ROCm, Neuron, TPU)
✓ vendor-neutral metrics (normalize at collection)
✗ CUDA-specific kernels
✗ TensorRT engines
✗ vendor-specific quantization formatsDesigning for portability costs little upfront and preserves optionality. The main discipline: keep vendor-specific optimization in a replaceable layer.
11. Hands-on exercise#
A. Cost comparison. For a model you serve, compute cost per million tokens on: your current hardware, and (using published specs and prices) two alternatives. Include an estimate of the software effort.
B. Ridge point comparison. Compute the ridge point for MI300X, H100, H200, and a TPU generation. Which is least memory-bound for decode?
C. The 192 GB question. For a 70B model, compare: MI300X single-GPU (no TP) versus H100 TP=2. Estimate the AllReduce overhead avoided. Is the single-GPU option faster?
D. Portability audit. For your current stack, list every vendor-specific dependency. How hard would a port be? What would you change now to make it easier?
E. On-device cascade. For a task you serve, estimate what fraction could run on-device (small quantized model) versus needing the server. What’s the cost saving?
F. If you have access: run the same model on two vendors’ hardware with the same engine (vLLM supports several). Compare throughput, latency, and the effort to get it working.
12. Interview questions#
- When would you evaluate a non-NVIDIA accelerator?
- What’s the architectural advantage of MI300X’s 192 GB for inference?
- Why does TPU’s static-shape requirement complicate LLM serving?
- How do Groq and Cerebras attack the memory bandwidth problem differently?
- What’s the dominant cost of adopting an alternative accelerator?
- What would you do to keep your stack portable?
- How would you compare accelerators fairly?
13. Further reading#
- [REFERENCE] AMD ROCm documentation and MI300 specifications
- [REFERENCE] Google Cloud TPU documentation; JAX/XLA
- [REFERENCE] AWS Neuron SDK documentation
- [REFERENCE] Intel Gaudi documentation
- [REFERENCE] MLPerf Inference results — the closest thing to a fair comparison
- Next: 11 — Memory-centric and disaggregated futures