One page for the hardware facts the rest of the curriculum keeps reaching for: what GPU memory is, how much each card has, how fast it moves bytes, where the bytes go, and how to inspect all of it.
Numbers are vendor-published figures, rounded. They drift with SKUs and revisions — treat them as starting points and replace them with what you measure in Project 11. The tables were last checked on 3 October 2026; figures for parts that began shipping in 2026 (Rubin, MI455X) are vendor claims with little independent measurement yet.
1. The three numbers that decide everything#
flowchart LR CAP["CAPACITY<br/>GB of VRAM"] --> FIT["What fits<br/>weights + KV cache<br/>= model size x concurrency"] BW["BANDWIDTH<br/>GB/s of memory"] --> DEC["Decode speed<br/>memory-bound: every token<br/>re-reads every weight"] FL["COMPUTE<br/>TFLOP/s"] --> PRE["Prefill speed and<br/>large-batch throughput<br/>compute-bound"] FIT --> COST["Cost per token"] DEC --> COST PRE --> COST class CAP,BW,FIT,DEC memory class FL,PRE compute class COST neutral
For LLM serving, read a spec sheet in this order: capacity, bandwidth, then FLOPs. Most buying mistakes come from reading it in the opposite order.
2. What “GPU memory” physically is#
Two families of memory sit next to GPU dies. Which one a card uses explains most of its price and most of its inference performance.
flowchart TB
subgraph GD["GDDR card (T4, L4, A10G, L40S, RTX)"]
direction LR
D1["GPU die"] --- PCB["Memory chips soldered on the PCB<br/>around the package<br/>narrow, fast-clocked links"]
end
subgraph HB["HBM card (A100, H100, H200, B200, MI300X)"]
direction LR
D2["GPU die"] --- INT["HBM stacks on the SAME silicon interposer<br/>DRAM dies stacked vertically, wired with TSVs<br/>very wide, short links"]
end
GD --> R1["Cheaper, lower bandwidth<br/>~0.3 - 1.8 TB/s"]
HB --> R2["Expensive, supply-constrained, high bandwidth<br/>~2 - 22 TB/s"]
class PCB,INT memory
class D1,D2 compute
class R1,R2 neutral| Technology | Where it sits | Typical card bandwidth | Found in |
|---|---|---|---|
| GDDR6 | Discrete chips on the board | 0.3 - 0.9 TB/s | T4, L4, A10G, L40S |
| GDDR6X / GDDR7 | Discrete chips, faster signalling | 0.9 - 1.8 TB/s | RTX 3090/4090, RTX 5090 |
| HBM2 / HBM2e | Stacked, on interposer | 1.5 - 2.0 TB/s | A100 |
| HBM3 | Stacked, on interposer | 3.35 TB/s | H100 SXM, MI300X (5.3) |
| HBM3e | Stacked, on interposer | 4.8 - 8 TB/s | H200, B200, B300, MI325X/MI355X |
| HBM4 | Stacked, interface twice as wide | ~22 TB/s | Rubin (shipping since August 2026), MI455X |
| Unified LPDDR | Shared with the CPU | 0.1 - 0.8 TB/s | Apple M-series, DGX Spark, Jetson |
Why it matters to you:
- HBM is the scarce component. GPU supply and price track HBM and advanced-packaging capacity more than they track logic wafers.
- Bandwidth comes from width × speed. HBM gets it from thousands of short wires through the interposer; GDDR has to do it with far fewer, longer traces. That is a physical ceiling, not a tuning problem.
- ECC. Datacenter cards have error-correcting memory; consumer cards do not. A flipped bit in a weight tensor on a consumer card is silent.
Inside the GPU the hierarchy above HBM — registers, shared memory, L2 — is covered in VI.03.
3. Spec sheet: the cards you will meet#
Datacenter — NVIDIA#
| GPU | Year | Arch | VRAM | Memory type | Bandwidth | FP16/BF16 dense | Power | GPU↔GPU link |
|---|---|---|---|---|---|---|---|---|
| T4 | 2018 | Turing | 16 GB | GDDR6 | 0.32 TB/s | 65 TF | 70 W | PCIe only |
| A10G | 2021 | Ampere | 24 GB | GDDR6 | 0.60 TB/s | ~125 TF | 150 W | PCIe only |
| L4 | 2023 | Ada | 24 GB | GDDR6 | 0.30 TB/s | ~120 TF | 72 W | PCIe only |
| L40S | 2023 | Ada | 48 GB | GDDR6 | 0.86 TB/s | ~362 TF | 350 W | PCIe only |
| A100 40GB | 2020 | Ampere | 40 GB | HBM2 | 1.56 TB/s | 312 TF | 400 W | NVLink 600 GB/s |
| A100 80GB | 2020 | Ampere | 80 GB | HBM2e | 2.04 TB/s | 312 TF | 400 W | NVLink 600 GB/s |
| H100 PCIe | 2022 | Hopper | 80 GB | HBM2e | ~2.0 TB/s | ~756 TF | 350 W | NVLink bridge |
| H100 SXM | 2022 | Hopper | 80 GB | HBM3 | 3.35 TB/s | ~990 TF | 700 W | NVLink 900 GB/s |
| H200 | 2023 | Hopper | 141 GB | HBM3e | 4.8 TB/s | ~990 TF | 700 W | NVLink 900 GB/s |
| B200 | 2024 | Blackwell | ~180-192 GB | HBM3e | ~8 TB/s | ~2,250 TF | ~1,000 W | NVLink 1.8 TB/s |
| B300 | 2025 | Blackwell Ultra | 288 GB | HBM3e | ~8 TB/s | not published (~15 PF at FP4) | ~1,400 W | NVLink 1.8 TB/s |
| Rubin | 2026 | Rubin | 288 GB | HBM4 | ~22 TB/s | not published (~50 PF at FP4, NVIDIA’s inference figure) | ~1,800-2,300 W | NVLink 3.6 TB/s |
Rubin is sold only inside liquid-cooled Vera Rubin racks; you rent it, you do not slot it into a server. B300 has the B200’s bandwidth: for decode it is a capacity upgrade, not a speed one.
Datacenter — AMD#
| GPU | VRAM | Memory type | Bandwidth | Power |
|---|---|---|---|---|
| Instinct MI300X | 192 GB | HBM3 | 5.3 TB/s | 750 W |
| Instinct MI325X | 256 GB | HBM3e | ~6 TB/s | ~1,000 W |
| Instinct MI355X | 288 GB | HBM3e | ~8 TB/s | ~1,400 W |
| Instinct MI455X | 432 GB | HBM4 | ~23 TB/s (vendor) | sold in Helios racks, shipping from H2 2026 |
Workstation / consumer (learning, local serving)#
| GPU | VRAM | Memory type | Bandwidth | Power |
|---|---|---|---|---|
| RTX 3090 | 24 GB | GDDR6X | 0.94 TB/s | 350 W |
| RTX 4090 | 24 GB | GDDR6X | 1.01 TB/s | 450 W |
| RTX 5090 | 32 GB | GDDR7 | 1.79 TB/s | 575 W |
| Apple M-series (Max/Ultra) | up to 128-512 GB unified | LPDDR5(X) | ~0.4 - 0.8 TB/s | 60-200 W system |
Things to notice
- H100 → H200 added no compute — only 43% more bandwidth and 76% more capacity — and was a large inference upgrade. That tells you what inference is bound by.
- L4 and T4 are capacity-per-watt cards, not speed cards. An L4 has the VRAM of an A10G and half its bandwidth.
- A 4090 out-reads an L40S on bandwidth but has half the memory and no ECC.
- “H100” is two different cards. The PCIe variant has HBM2e and roughly 60% of the SXM bandwidth. Check which one your cloud instance has.
- MIG (hardware partitioning into up to 7 isolated instances) exists on A100 / H100 / H200 / Blackwell, not on T4 / L4 / A10G / L40S.
For the generation-over-generation ratios and where they are heading, see XIV.09.
4. Where the gigabytes go#
pie showData title 80 GB H100 serving an 8B model in FP16 (illustrative) "Model weights" : 16 "KV cache pool" : 52 "Activations and workspace" : 2 "CUDA context, framework, graphs" : 2 "Unreserved headroom (10%)" : 8
| Consumer | Size | Scales with | Notes |
|---|---|---|---|
| Weights | params × bytes/param | Model, precision | Fixed once loaded |
| KV cache | 2 × layers × kv_heads × head_dim × bytes × tokens | Concurrency × context | The part you actually manage (V.06) |
| Activations | batch_tokens × d_model × small factor | Tokens per step | Peaks during prefill of long prompts |
| CUDA context | ~0.3 - 0.5 GB per process | Number of processes | Paid per process, not per GPU |
| Framework overhead | 0.5 - 2 GB | Engine | cuDNN/cuBLAS workspaces, CUDA graphs, compiled kernels |
| Allocator slack | varies | Allocation pattern | reserved − allocated; fragmentation (X.05) |
Engines such as vLLM claim a fixed fraction of VRAM at startup (--gpu-memory-utilization,
default 0.9), load the weights, and turn everything left over into the KV block pool. So
nvidia-smi shows ~90% used from the first second, whether you have 1 user or 100.
Weight sizes to memorize#
| Parameters | FP32 | FP16 / BF16 | FP8 / INT8 | INT4 (≈4.5 bit effective) |
|---|---|---|---|---|
| 1 B | 4 GB | 2 GB | 1 GB | ~0.6 GB |
| 8 B | 32 GB | 16 GB | 8 GB | ~4.5 GB |
| 32 B | 128 GB | 64 GB | 32 GB | ~18 GB |
| 70 B | 280 GB | 140 GB | 70 GB | ~39 GB |
| 120 B | 480 GB | 240 GB | 120 GB | ~68 GB |
What fits where (weights only — leave room for KV)#
| Model @ precision | 16 GB | 24 GB | 48 GB | 80 GB | 141 GB | ~180 GB |
|---|---|---|---|---|---|---|
| 8B FP16 (16 GB) | no room for KV | yes | yes | yes | yes | yes |
| 8B INT4 (4.5 GB) | yes | yes | yes | yes | yes | yes |
| 32B FP16 (64 GB) | — | — | — | tight | yes | yes |
| 32B INT4 (18 GB) | — | tight | yes | yes | yes | yes |
| 70B FP16 (140 GB) | — | — | — | 2 GPUs | barely, no KV | yes |
| 70B FP8 (70 GB) | — | — | — | tight | yes | yes |
| 70B INT4 (39 GB) | — | — | tight | yes | yes | yes |
“Fits” without KV room is not serving — it is loading. A model that leaves 1 GB for KV serves almost nobody.
5. From bandwidth to tokens per second#
At batch 1, a decode step reads every weight once. So:
tokens/s ceiling (batch 1) ≈ memory bandwidth / weight bytes
realistic ≈ 60-75% of thatCeiling for a model with 16 GB of weights (e.g. 8B at FP16):
| GPU | Bandwidth | Batch-1 ceiling | Realistic |
|---|---|---|---|
| L4 | 0.30 TB/s | ~19 tok/s | ~12-14 |
| A10G | 0.60 TB/s | ~38 tok/s | ~23-28 |
| L40S | 0.86 TB/s | ~54 tok/s | ~32-40 |
| RTX 4090 | 1.01 TB/s | ~63 tok/s | ~38-47 |
| RTX 5090 | 1.79 TB/s | ~112 tok/s | ~67-84 |
| A100 80GB | 2.04 TB/s | ~127 tok/s | ~76-95 |
| H100 SXM | 3.35 TB/s | ~209 tok/s | ~125-157 |
| H200 | 4.8 TB/s | ~300 tok/s | ~180-225 |
| B200 / B300 | ~8 TB/s | ~500 tok/s | ~300-375 |
| Rubin | ~22 TB/s | ~1,375 tok/s | ~825-1,030 (predicted; measure it) |
Three ways to beat the ceiling, all of which are “move fewer bytes per token”:
flowchart LR C["Batch-1 ceiling<br/>bandwidth / weight bytes"] --> B["Batch more sequences<br/>one weight read serves many tokens"] C --> Q["Quantize weights<br/>fewer bytes to read"] C --> S["Speculative decoding<br/>several tokens per weight read"] B --> R["Until the ridge point:<br/>then you are compute-bound"] class C neutral class B,Q,S memory class R compute
The ridge point (peak FLOPs / bandwidth) for these cards and what it means for batch size is
worked through in VI.04 and
V.15.
6. The paths bytes take: interconnect speeds#
flowchart LR
DISK["NVMe SSD"] -->|"3 - 14 GB/s"| RAM["Host RAM"]
NET["Network storage"] -->|"0.1 - 5 GB/s"| RAM
RAM -->|"PCIe Gen4 x16: ~25 GB/s real<br/>Gen5 x16: ~50 GB/s real"| G0[("GPU 0 HBM<br/>2 - 22 TB/s internal")]
G0 <-->|"NVLink: 600 - 3,600 GB/s"| G1[("GPU 1 HBM")]
G0 -->|"RDMA NIC: 12 - 100 GB/s"| NODE["Other nodes"]
class G0,G1 memory
class DISK,NET,NODE io
class RAM neutral| Link | Theoretical | What it means for inference |
|---|---|---|
| HBM inside one GPU | 2,000 - 22,000 GB/s | The speed everything else is compared to |
| NVLink 3 / 4 / 5 / 6 (per GPU) | 600 / 900 / 1,800 / 3,600 GB/s | Makes tensor parallelism viable (IX.03) |
| PCIe Gen3 x16 | ~16 GB/s | T4-era hosts; slow model loads |
| PCIe Gen4 x16 | ~32 GB/s | ~25 GB/s with pinned memory |
| PCIe Gen5 x16 | ~64 GB/s | Current servers |
| 100 / 200 / 400 GbE or InfiniBand | 12.5 / 25 / 50 GB/s | KV transfer for disaggregation (XIII.06) |
| NVMe Gen4 / Gen5 | ~7 / ~14 GB/s | Cold-start weight loading (II.07) |
The gap between the first row and the rest — two orders of magnitude — is why anything that leaves the GPU on the per-token path is a design error, and why swapping KV to host memory is often slower than recomputing it.
Load-time arithmetic: 140 GB of weights ÷ 7 GB/s NVMe ≈ 20 s at best; from network storage at 1 GB/s, over two minutes. That is your cold start floor before the first kernel runs.
7. Inside a GPU server#
flowchart TB
subgraph NODE["8-GPU server (HGX / DGX class)"]
direction TB
subgraph CPUS["Host"]
C0["CPU socket 0<br/>+ DRAM (NUMA node 0)"]
C1["CPU socket 1<br/>+ DRAM (NUMA node 1)"]
end
subgraph PX["PCIe switches"]
S0["Switch A"]
S1["Switch B"]
end
subgraph GP["GPUs"]
direction LR
G0["GPU 0-3"]
G1["GPU 4-7"]
end
NVS["NVSwitch fabric<br/>every GPU to every GPU at full NVLink speed"]
NIC["RDMA NICs / DPUs"]
SSD["NVMe"]
C0 --> S0 --> G0
C1 --> S1 --> G1
G0 <--> NVS
G1 <--> NVS
S0 --> NIC
S1 --> SSD
end
class G0,G1 compute
class NVS memory
class NIC,SSD,S0,S1 io
class C0,C1 neutral| System | GPUs | Total GPU memory | GPU fabric | Power |
|---|---|---|---|---|
| Single-GPU cloud VM (T4, L4, A10G, L40S) | 1 | 16-48 GB | none | < 0.5 kW |
| 8× A100 80GB server | 8 | 640 GB | NVSwitch, 600 GB/s | ~6.5 kW |
| 8× H100 server (DGX H100 class) | 8 | 640 GB | NVSwitch, 900 GB/s | ~10 kW |
| 8× H200 server | 8 | ~1.1 TB | NVSwitch, 900 GB/s | ~10 kW |
| GB200 NVL72 rack | 72 | ~13 TB | NVLink switch spine, 1.8 TB/s, one domain | ~120 kW, liquid cooled |
| GB300 NVL72 rack (Blackwell Ultra) | 72 | ~20.7 TB | NVLink switch spine, 1.8 TB/s, one domain | ~120-140 kW, liquid cooled |
| Vera Rubin NVL72 rack | 72 | ~20.7 TB (HBM4) | NVLink 6, 3.6 TB/s, one domain | roughly 200 kW or more (reported), liquid cooled |
Host-side sizing that people get wrong:
- System RAM ≥ total VRAM as a floor (more if you offload KV or keep several models page-cached). Loading goes disk → page cache → GPU.
- NUMA locality. A GPU hangs off one socket. Pinning the process to the other socket’s
memory costs transfer bandwidth (II.03).
nvidia-smi topo -mshows the mapping. - CPU cores. Tokenization, scheduling, and detokenization are CPU work on the critical path. Budget several fast cores per GPU; a starved host shows up as an idle GPU.
- Power and cooling are the real datacenter limits. Tens of kW per air-cooled chassis, over 100 kW per liquid-cooled rack. A thermally throttled GPU silently loses clock speed.
8. Reading GPU memory: the commands#
# The overview
nvidia-smi
# The fields that matter, machine-readable, every second
nvidia-smi --query-gpu=index,name,memory.total,memory.used,memory.free,utilization.gpu,utilization.memory,temperature.gpu,power.draw,clocks_throttle_reasons.active \
--format=csv -l 1
# Who is using the memory
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv
# Rolling per-second device stats (power, util, clocks, memory, PCIe throughput)
nvidia-smi dmon -s pucvmet
# Topology: which GPU sits on which CPU socket / NUMA node, and GPU-GPU link types
nvidia-smi topo -m
# NVLink status and error counters
nvidia-smi nvlink -s
nvidia-smi nvlink -e
# Memory health
nvidia-smi -q -d ECC,ROW_REMAPPER,PAGE_RETIREMENT
# MIG layout (if enabled)
nvidia-smi -L
# Driver-level errors
dmesg -T | grep -i xidimport torch
free, total = torch.cuda.mem_get_info() # what the DRIVER says is free (bytes)
torch.cuda.memory_allocated() # bytes in live tensors
torch.cuda.memory_reserved() # bytes PyTorch holds from the driver
torch.cuda.max_memory_allocated() # peak — the number you size against
print(torch.cuda.memory_summary()) # allocator breakdownHow the three views relate:
flowchart LR T["Live tensors<br/>memory_allocated"] --> R["PyTorch cache<br/>memory_reserved<br/>(allocated + free blocks it keeps)"] R --> P["Process total in nvidia-smi<br/>(reserved + CUDA context + libraries)"] P --> G["GPU memory.used<br/>(sum over all processes)"] class T,R memory class P,G neutral
Reading traps
utilization.gpuis “percent of time at least one kernel was running”, not how hard the GPU worked (X.04).utilization.memoryis “percent of time memory was being read or written” — not how full the memory is. That ismemory.used / memory.total.memory.usednear 90% on a vLLM/SGLang node is by design (pre-allocated KV pool). The useful signal is the engine’s own KV-cache-usage metric.reservedmuch greater thanallocated→ fragmentation or a cache that never shrinks, not a leak in your tensors.
9. Sharing one GPU#
| Mechanism | Memory isolation | Fault isolation | Granularity | Use it for |
|---|---|---|---|---|
| Whole GPU per pod | Full | Full | 1 GPU | Production LLM serving (default) |
| MIG | Hardware-enforced slices | Yes | Up to 7 fixed slices (e.g. 10 GB each on H100) | Small models, strict multi-tenancy |
| MPS | Optional per-client limits, not enforced by hardware | No — one client’s fault can take down the others | Arbitrary | Several cooperating processes of one tenant |
| Time-slicing | None | No | Arbitrary replicas | Dev/test, bursty notebooks |
A MIG slice gets a fixed share of memory and of memory bandwidth and SMs, so the bandwidth arithmetic in section 5 scales down with the slice. Kubernetes exposure of each mode is in XII.03.
10. When the hardware misbehaves#
| Symptom | Likely cause | First check |
|---|---|---|
| Tokens/s dropped 20-40% with no deploy | Thermal or power throttling | clocks_throttle_reasons.active, temperature, power cap |
Xid 48, Xid 94/95 in dmesg | Uncorrectable / contained ECC memory errors | nvidia-smi -q -d ECC,ROW_REMAPPER; drain the node |
Xid 63/64 | Row remapping event / failure | Reset or RMA if remap failed |
Xid 79 | GPU fell off the bus | Power, seating, PCIe errors; node needs a reboot |
Xid 31 | GPU memory page fault | Usually a software bug (bad pointer in a kernel) |
Xid 74 | NVLink error | nvidia-smi nvlink -e; affects TP jobs first |
| One GPU of eight is slower | PCIe link trained down, or wrong NUMA node | lspci -vv link speed/width, nvidia-smi topo -m |
| OOM at steady load after hours | Fragmentation or slow leak | memory_reserved vs memory_allocated over time (X.06) |
| Slow model load | Network storage or cold page cache | Measure read GB/s; cache weights on local NVMe |
Corrected single-bit errors are normal at fleet scale. Rising counts on one device are a reason to drain it before it becomes an uncorrectable one in the middle of someone’s request.
11. Picking a GPU#
flowchart TB
S["Start: model + precision + context + concurrency"] --> W["Weight bytes + KV bytes<br/>= memory needed"]
W --> F{"Fits on one GPU<br/>with KV headroom?"}
F -->|"no"| Q{"Can you quantize<br/>within your quality budget?"}
Q -->|"yes"| W
Q -->|"no"| TP["Bigger-memory GPU (H200 / B200 / B300 / Rubin / MI3xx-4xx)<br/>or tensor parallelism over NVLink"]
F -->|"yes"| L{"Latency SLO<br/>tight on tokens/s per user?"}
L -->|"yes"| HB["HBM card<br/>bandwidth sets per-user speed"]
L -->|"no, throughput per dollar"| GD["Cheapest card where it fits<br/>(L40S / A10G / L4) + more replicas"]
TP --> CHK["Benchmark on the real workload"]
HB --> CHK
GD --> CHK
class W,HB,TP memory
class GD compute
class F,Q,L queue
class S,CHK neutralRules of thumb:
- Size memory first. If it does not fit with KV room, nothing else matters.
- Per-user speed is bandwidth; fleet throughput per dollar is often a cheaper card, replicated. Replicas beat tensor parallelism whenever the model fits on one device (IX.12).
- Price per GB of VRAM and price per TB/s are more useful columns than price per TFLOP.
- Check the exact SKU (SXM vs PCIe, 40 vs 80 GB) and whether the instance exposes NVLink.
- Then measure. Datasheet → prediction → benchmark, in that order.
12. Hands-on#
A. Inventory. On any GPU machine, run every command in section 8. Record model, VRAM,
driver, NUMA node, PCIe link speed and width into numbers.md.
B. Budget. Start vLLM (or your Project 08 engine) with a small model. Before sending
traffic, account for every GB in nvidia-smi: weights, KV pool, context, the unreserved 10%.
Your sum should land within ~1 GB.
C. Predict then measure. Using sections 4 and 5, predict batch-1 tokens/s for your model on your GPU. Measure it. Explain the gap.
D. Fill the table. Add a row to section 3 for a GPU not listed, from the vendor datasheet, and compute its ridge point and its batch-1 ceiling for a 16 GB model.
E. Break it on purpose. Run two processes on one GPU and watch per-process memory. Then
set a power cap (nvidia-smi -pl) and re-measure tokens/s.
13. Interview questions#
- Why does an H200 serve LLMs faster than an H100 when both have the same FLOPs?
- What is HBM, and why do datacenter GPUs use it instead of GDDR?
- An 80 GB GPU runs a 16 GB model. Where does the rest of the memory go, and why does
nvidia-smishow it full with no traffic? - Estimate batch-1 tokens/s for a 70B FP8 model on an H200. Show the arithmetic.
- What is the difference between
memory_allocated,memory_reserved, and whatnvidia-smireports? utilization.memoryreads 35%. Is the GPU’s memory 35% full?- MIG vs MPS vs time-slicing — which gives memory isolation, and which would you allow for two different customers?
- Why is PCIe bandwidth irrelevant to steady-state decode but critical to cold start and KV offloading?
- You see
Xid 79on one node of a tensor-parallel job. What happened and what do you do? - Two cards have the same VRAM. One costs half as much. What single spec do you check before buying it for chat serving?
14. Further reading#
- [REFERENCE] NVIDIA product datasheets for each card — the source for section 3; re-check before any purchase
- [REFERENCE] NVIDIA “Xid Errors” documentation and the GPU Deployment and Management Guide
- [REFERENCE] NVIDIA MIG User Guide
- [FUNDAMENTAL] VI.01 — GPU architecture, VI.03 — Memory hierarchy, VI.04 — Bandwidth and roofline
- [ESTABLISHED] V.06 — KV cache math, V.15 — Capacity math
- [EMERGING] XIV.09 — NVIDIA trends, XIV.10 — Alternative accelerators, XIV.11 — Memory-centric designs
- See also the hardware entries under “2025-2026 additions” in RESOURCES.md
- For reading these devices in production — DCGM fields, Xid events, power and energy — see Observability Engineering V.02