PidokuInfra

Alternative Accelerators

Expert Advanced 1h 15m Difficulty 3/5 Topic 10 of 11

Prerequisites 09, I.09


1. Why look beyond NVIDIA#

REASONS
  cost           NVIDIA's margins are high; alternatives are cheaper per FLOP
  availability   supply constraints are real
  lock-in        a single-vendor dependency is a business risk
  specialization some designs are genuinely better for inference specifically

REASONS NOT TO
  software       CUDA's ecosystem is a decade ahead
  maturity       kernels, engines, and tooling lag
  effort         porting and maintaining a second stack is expensive

The honest position: NVIDIA is the default for good reasons, and the alternatives are viable for specific situations. This file is about identifying those situations.


2. The landscape#

ACCELERATOR       TYPE          MATURITY   BEST FOR
AMD Instinct      GPU           good       large-memory inference; ROCm
 (MI300X-MI355X; MI455X in Helios racks)
Intel Gaudi 2/3   ASIC          moderate   cost-sensitive training+inference
Google TPU        ASIC          mature*    if you're on GCP / using JAX
AWS Trainium /    ASIC          moderate   if you're on AWS, cost-sensitive
 Inferentia2
Microsoft Maia    ASIC          internal   Azure's own inference workloads
Cerebras WSE      wafer-scale   niche      very low-latency inference
Groq LPU          ASIC          niche**    extremely low-latency decode
SambaNova         RDU           niche      
Apple M-series    integrated    good       on-device, unified memory
Qualcomm, ARM     mobile NPU    good       on-device

*  mature within Google's ecosystem; less so outside it
** licensed by NVIDIA in December 2025; also sold as NVIDIA's Groq 3 LPX rack
STATE OF THE FIELD — checked 3 October 2026 (vendor claims unless measured)

AMD        Helios: first rack-scale system, 72 x MI455X, a reported
           432 GB HBM4 per GPU; shipping to large customers from H2 2026
Google     7th-gen TPU (Ironwood) generally available, positioned for
           inference; 8th generation announced as separate training
           and inference chips
AWS        Trainium3 generally available; Trainium4 announced
Microsoft  Maia 200 (January 2026), described as an inference chip
NVIDIA     licensed Groq's LPU technology; Groq 3 LPX rack announced
           at GTC 2026 alongside Rubin
Cerebras   in production for large customers; next systems shown at
           Hot Chips 2026

THE PATTERN: everyone moved to rack-scale systems, and nearly
everyone now builds something inference-specific. The evaluation
method in this file has not changed.

3. AMD MI300X — the closest substitute#

SPECS (MI300X)
  memory:      192 GB HBM3      ← the headline advantage
  bandwidth:   5.3 TB/s         ← higher than H100's 3.35
  FP16:        ~1,300 TF
  FP8:         ~2,600 TF
  Infinity Fabric between GPUs

RIDGE POINT: 1300/5.3 = 245  (vs H100's 296)
  → slightly less memory-bound than H100 for the same workload

WHY IT'S INTERESTING FOR INFERENCE
  ✓ 192 GB means a 70B model at FP16 fits on ONE GPU (140 GB + KV)
    → no tensor parallelism needed → no AllReduce overhead
  ✓ higher bandwidth than H100 → faster decode
  ✓ generally lower cost per unit

THE SOFTWARE QUESTION
  ROCm has matured substantially. vLLM, SGLang, and PyTorch all
  support it. FlashAttention has ROCm implementations.
  Triton compiles to AMD (Section XIII.11) — a portability advantage.
  
  ✗ but: kernel coverage lags, new model architectures are supported
    later, and debugging tools are less mature
  ✗ FP8 support and its ecosystem trail Hopper's

MI300X’s 192 GB is a genuine architectural advantage for inference. Fitting a 70B model on one GPU eliminates tensor parallelism entirely, which removes 160 AllReduces per token (Section IX.03).

When to evaluate it: if you’re serving models in the 30-100B range, cost matters, and you can tolerate a software maturity gap.


4. Google TPU#

DESIGN PHILOSOPHY
  systolic array for matmul, large on-chip memory, high-bandwidth
  interconnect (ICI) designed for large-scale training

TPU v5e / v5p / v6 (Trillium) / v7 (Ironwood)
  memory:      16-95 GB HBM per chip through v5p; more on newer
               generations — check the current GCP documentation
  interconnect: very high-bandwidth mesh — better multi-chip scaling
                than PCIe-based systems
  
FOR INFERENCE
  ✓ excellent price/performance within GCP
  ✓ the interconnect makes large-model serving straightforward
  ✓ JAX/XLA is a mature, well-designed stack
  ✗ requires the JAX/XLA programming model (not PyTorch-native,
    though PyTorch/XLA exists)
  ✗ static shapes are required → dynamic-shape LLM serving needs
    bucketing
  ✗ available only on GCP

WHEN
  → you're already on GCP
  → you use JAX
  → you can accept the static-shape constraint

The static-shape requirement is the practical friction for LLM serving: Section IV.11’s dynamic shape problem is harder when the compiler requires static shapes. Bucketing works but costs padding.


5. AWS Inferentia2 / Trainium#

DESIGN
  purpose-built for inference (Inferentia) and training (Trainium)
  NeuronCore-v2, 32 GB HBM per chip, high inter-chip bandwidth

FOR INFERENCE
  ✓ significantly cheaper per token than GPU instances for supported
    models, within AWS
  ✓ Neuron SDK integrates with PyTorch; vLLM has a Neuron backend
  ✗ model support is narrower — check YOUR model
  ✗ compilation is required (AOT, like TensorRT — Section VII.09)
  ✗ AWS-only

WHEN
  → you're on AWS, cost-sensitive, and your model is supported
  → the model set is stable (compilation is AOT)

6. Intel Gaudi#

DESIGN
  Gaudi 2: 96 GB HBM2e, 2.45 TB/s; Gaudi 3: 128 GB HBM2e, 3.7 TB/s
  integrated RoCE ports on-chip → scale-out without separate NICs

FOR INFERENCE
  ✓ competitive price/performance
  ✓ integrated networking simplifies multi-node
  ✓ vLLM and PyTorch support exists
  ✗ ecosystem maturity trails
  ✗ smaller installed base → fewer people have solved your problem

WHEN
  → cost is the dominant factor and you can invest in the software

7. The low-latency specialists#

GROQ LPU
  DESIGN: deterministic execution, all weights in on-chip SRAM
          (no HBM), software-scheduled
  → extremely low decode latency (hundreds of tokens/sec per user)
  
  ✓ latency that GPU architectures cannot match at batch 1
  ✗ SRAM capacity is small → a large model needs MANY chips
  ✗ the economics depend entirely on utilization across those chips
  → a fundamentally different point in the design space

CEREBRAS WSE
  DESIGN: wafer-scale chip, enormous on-chip memory
  → very low latency, high single-model throughput
  ✗ availability, cost, and ecosystem are niche

THE COMMON THREAD
  both attack the memory-bandwidth problem by putting weights in
  SRAM instead of HBM.
  → SRAM bandwidth is ~10-100x HBM
  → but SRAM capacity is ~1000x smaller
  → so you need many chips, and the economics turn on whether you
    can keep them all busy

These are interesting because they attack the actual bottleneck differently. Section I.07’s memory wall is a consequence of using DRAM; using SRAM instead removes it and creates a capacity problem instead.

Whether that trade wins depends on utilization — which is a business question, not a technical one.


8. On-device#

APPLE M-SERIES
  unified memory: 400-800 GB/s CPU-GPU-NPU shared
  → no PCIe transfer; the whole memory is "GPU memory"
  → a 70B model at INT4 (35 GB) runs on a 64 GB Mac
  ✓ genuinely capable for local inference
  ✗ bandwidth is far below datacenter GPUs

QUALCOMM / ARM NPUs
  → phone-scale models (1-8B quantized)
  → INT4/INT8, very power-constrained

WHY IT MATTERS FOR SERVER-SIDE ENGINEERS
  → on-device handles some of your traffic if you build for it
  → the techniques are the same (quantization, KV management),
    with tighter constraints
  → hybrid architectures (small model on-device, escalate to server)
    are a real cost lever (Section XII.09's cascade, across the
    network boundary)

9. How to evaluate an alternative#

1. DOES IT RUN YOUR MODEL?
   → check the specific architecture, not "supports LLMs"
   → check the specific features (GQA? MoE? long context?)

2. WHAT'S THE COST PER MILLION TOKENS?
   → Section XI.03's methodology, on THEIR hardware
   → not $/FLOP or $/GB — cost per unit of YOUR work

3. WHAT'S THE SOFTWARE EFFORT?
   → porting, debugging, maintaining a second stack
   → this is usually the dominant cost, and it's ongoing

4. WHAT'S THE ECOSYSTEM RISK?
   → will your next model be supported? How quickly?
   → how many people have solved the problems you'll hit?

5. WHAT'S THE EXIT?
   → if it doesn't work out, how hard is it to move back?
   → keep your stack portable (OpenAI API, standard checkpoints)

THE PRAGMATIC ANSWER FOR MOST TEAMS
  evaluate alternatives when: cost is a first-order concern AND
  your model set is stable AND you have engineering capacity.
  Otherwise the software cost exceeds the hardware saving.

10. The portability layer#

WHAT MAKES A MULTI-VENDOR STACK POSSIBLE

  ✓ OpenAI-compatible API           (clients don't change)
  ✓ standard checkpoints            (safetensors, not vendor formats)
  ✓ PyTorch as the model definition (most backends support it)
  ✓ Triton for custom kernels       (compiles to CUDA and ROCm)
  ✓ engine abstraction              (vLLM supports CUDA, ROCm, Neuron, TPU)
  ✓ vendor-neutral metrics          (normalize at collection)

  ✗ CUDA-specific kernels
  ✗ TensorRT engines
  ✗ vendor-specific quantization formats

Designing for portability costs little upfront and preserves optionality. The main discipline: keep vendor-specific optimization in a replaceable layer.


11. Hands-on exercise#

A. Cost comparison. For a model you serve, compute cost per million tokens on: your current hardware, and (using published specs and prices) two alternatives. Include an estimate of the software effort.

B. Ridge point comparison. Compute the ridge point for MI300X, H100, H200, and a TPU generation. Which is least memory-bound for decode?

C. The 192 GB question. For a 70B model, compare: MI300X single-GPU (no TP) versus H100 TP=2. Estimate the AllReduce overhead avoided. Is the single-GPU option faster?

D. Portability audit. For your current stack, list every vendor-specific dependency. How hard would a port be? What would you change now to make it easier?

E. On-device cascade. For a task you serve, estimate what fraction could run on-device (small quantized model) versus needing the server. What’s the cost saving?

F. If you have access: run the same model on two vendors’ hardware with the same engine (vLLM supports several). Compare throughput, latency, and the effort to get it working.


12. Interview questions#

  1. When would you evaluate a non-NVIDIA accelerator?
  2. What’s the architectural advantage of MI300X’s 192 GB for inference?
  3. Why does TPU’s static-shape requirement complicate LLM serving?
  4. How do Groq and Cerebras attack the memory bandwidth problem differently?
  5. What’s the dominant cost of adopting an alternative accelerator?
  6. What would you do to keep your stack portable?
  7. How would you compare accelerators fairly?

13. Further reading#

  • [REFERENCE] AMD ROCm documentation and MI300 specifications
  • [REFERENCE] Google Cloud TPU documentation; JAX/XLA
  • [REFERENCE] AWS Neuron SDK documentation
  • [REFERENCE] Intel Gaudi documentation
  • [REFERENCE] MLPerf Inference results — the closest thing to a fair comparison
  • Next: 11 — Memory-centric and disaggregated futures

↑↓ navigate↵ openesc close