PidokuInfra

Memory-Centric and Disaggregated Futures

Expert Advanced 1h Difficulty 4/5 Topic 11 of 11

Prerequisites I.07, XIII.06, XIII.07, 09

The final file. It looks at where the field is heading, and closes the loop on the idea this curriculum opened with.


1. The thesis#

INFERENCE IS A MEMORY PROBLEM.

  decode is memory-bandwidth-bound (Section I.07)
  capacity limits concurrency (Section V.06)
  the KV cache is the binding constraint (Section V.05)
  quantization is a bandwidth optimization (Section VII.02)
  batching is a bandwidth amortization (Section I.06)
  MoE trades memory capacity for quality (Section XIII.02)
  disaggregation is about where KV lives (Section XIII.06)

→ EVERY major technique in this curriculum is, at bottom, about
  moving fewer bytes or moving them from somewhere closer.

The architectural implication: systems designed around memory rather than around compute should win. That’s the thesis of everything in this file.


2. The memory hierarchy is extending#

TODAY                              EMERGING
  registers                          registers
  SRAM (shared memory)               SRAM
  L2                                 L2
  HBM (80-192 GB)                    HBM (192+ GB)
  ─────── PCIe wall ───────          ─────── coherent link ───────
  CPU DRAM (rarely used)             CPU DRAM (a real tier, 500 GB-2 TB)
  NVMe (never used)                  CXL-attached memory (TBs)
                                     NVMe (for cold KV)
                                     remote memory (over RDMA)
WHAT MAKES A TIER USABLE
  bandwidth relative to what you'd otherwise do (Section XIII.07)
  
  CPU DRAM over PCIe Gen5:  55 GB/s → marginal
  CPU DRAM over NVLink-C2C: 900 GB/s → GENUINELY USABLE
  CXL memory:               ~64 GB/s per link, aggregatable
  NVMe:                     7 GB/s → capacity only, not speed
  
→ the coherent-link change (Grace-Hopper and successors) is what
  turns CPU DRAM from a swap tier into a real level of the hierarchy

This is the most consequential hardware trend for inference software architecture. It changes: KV tiering (Section XIII.07), expert offloading (Section XIV.06), and what “fits” means.


3. Disaggregation, generalized#

THE PATTERN, APPLIED REPEATEDLY

  DISAGGREGATE PREFILL FROM DECODE (Section XIII.06)
    because they have different bottlenecks
    
  DISAGGREGATE KV FROM COMPUTE (Mooncake, and CXL-based designs)
    because KV is data with its own lifecycle
    
  DISAGGREGATE MEMORY FROM COMPUTE (CXL memory pooling)
    because the ratio of memory to compute needed varies by workload
    
  DISAGGREGATE EXPERTS FROM ROUTING (MoE with expert offload)
    because most experts are idle for any given token

THE COMMON IDEA
  a monolithic GPU couples compute, memory capacity, and memory
  bandwidth in a fixed ratio.
  Different workloads need different ratios.
  → disaggregation lets you provision them independently.

The counter-argument, which is currently winning: every disaggregation adds a link, and links are slower than what they replace. Coupling is fast; decoupling is flexible. The trade only pays when the link is fast enough — which is why coherent interconnect matters so much.


4. CXL and memory pooling#

CXL (Compute Express Link)
  a cache-coherent interconnect over PCIe physical layers
  
  CXL.mem: attach memory to a host, or POOL memory across hosts
  → a shared memory pool that any host can allocate from

FOR INFERENCE
  ✓ a large, shared tier for KV cache and cold model weights
  ✓ provision memory independently of GPUs
  ✓ a model too large for HBM could keep cold parts in CXL memory
  
  ✗ bandwidth: ~64 GB/s per x16 link. Aggregatable, but far below HBM.
  ✗ latency: 150-300 ns, vs HBM's ~500 cycles — comparable, actually
  ✗ maturity: deployment is early

STATUS: [RESEARCH → EMERGING]. Watch for it in inference-specific
        products rather than general server designs.

5. Processing in / near memory#

THE IDEA: if moving data is the bottleneck, compute where the data is.

PIM (Processing In Memory)
  arithmetic units inside the DRAM die
  → HBM-PIM (Samsung), and others
  → for GEMV specifically — exactly what decode does
  
NEAR-MEMORY COMPUTE
  compute units on the memory controller or the HBM base die

WHY IT'S APPEALING FOR LLM DECODE
  decode is GEMV: read a weight, multiply once, discard.
  arithmetic intensity 1 (Section I.07).
  → the ideal PIM workload: the computation is trivial and the
    data movement is everything
  → a PIM device could do the multiply where the weight lives

WHY IT HASN'T HAPPENED
  ✗ programming model: how do you express a transformer for PIM?
  ✗ the arithmetic must be simple (DRAM dies have little area for logic)
  ✗ integration with the rest of the model (attention isn't GEMV)
  ✗ ecosystem: no software stack
  ✗ economics: HBM is already expensive; adding logic makes it more so

STATUS: [RESEARCH]. Genuinely well-matched to the problem, and
        persistently 5 years away.

Worth understanding because the match to LLM decode is unusually good. If anything makes PIM practical, LLM inference is the workload that would justify it.


6. What would change the field#

Ranked by potential impact:

1. COHERENT CPU-GPU MEMORY AS STANDARD
   → the hierarchy extends; offloading and tiering become architecture
     rather than compromise
   → PROBABILITY: high. Already shipping.
   → IMPACT: large. Changes what fits and what you can cache.

2. HYBRID SSM/ATTENTION AT FRONTIER QUALITY (Section XIV.02)
   → O(1) state instead of O(S) KV cache
   → PROBABILITY: moderate. Shipping in some models; not yet at
     the frontier.
   → IMPACT: very large for long context. 10-100x on the binding
     constraint.

3. NATIVE LOW-PRECISION MODELS (FP4-trained)
   → removes the quantization step and its quality question
   → PROBABILITY: high. FP8 already happened; FP4 is the same path.
   → IMPACT: 2x on memory and compute, with no quality debate.

4. INFERENCE-TIME SCALING BECOMING UNIVERSAL (Section XIV.05)
   → workload shape shifts to overwhelmingly decode-dominated
   → PROBABILITY: already happening
   → IMPACT: large. Changes which optimizations matter.

5. BANDWIDTH GROWING FASTER THAN COMPUTE
   → the ridge point falls; less batching needed
   → PROBABILITY: moderate. H200 suggests the industry is trying.
   → IMPACT: moderate. Eases the constraint without removing it.

6. PIM / NEAR-MEMORY COMPUTE
   → PROBABILITY: low near-term
   → IMPACT: would be transformative for decode

7. OPTICAL INTERCONNECT AT SCALE
   → changes multi-node economics
   → PROBABILITY: low near-term
   → IMPACT: large for very large models

Items 1-4 are the ones to plan around. They’re happening or likely, and each changes what you should invest in.


7. What stays true#

Whatever changes, these don’t:

1. THE ROOFLINE
   FLOPs, bytes, and their ratio determine performance. The numbers
   change; the model doesn't.

2. BATCHING AMORTIZES FIXED COSTS
   Whatever the memory hierarchy, reading a weight once for many
   tokens beats reading it once per token.

3. MEASURE BEFORE OPTIMIZING
   The methodology in Sections VII.01 and X.01 outlives any technique.

4. THE BOTTLENECK MOVES
   Fix one, another appears. The skill is finding the current one,
   not knowing a fixed list of optimizations.

5. QUALITY IS PART OF PERFORMANCE
   A faster system that answers worse is not faster.

6. SCHEDULING BEATS KERNELS
   The 3-10x from batching and memory management exceeds the 20-60%
   from compilation and kernel work. Order accordingly.

7. THE SIMPLE THING FIRST
   FP8 before FP4. Chunked prefill before disaggregation. Routing
   before a KV store. n-gram before EAGLE.

If the curriculum has a single thesis beyond “inference is a memory problem,” it’s number 7.


8. Where to go from here#

YOU HAVE FINISHED THE CURRICULUM. NEXT:

1. BUILD SOMETHING
   The projects (projects/) if you haven't. Or contribute to vLLM
   or SGLang — the schedulers and attention backends are where the
   education is.

2. MEASURE SOMETHING REAL
   Take a production system. Apply Section X.01's methodology.
   Find the bottleneck. Fix it. Write it up.

3. HALVE A COST
   Take one system's cost per million tokens and cut it in half.
   Document what you did. That document is your career.

4. TEACH IT
   The fastest way to find your gaps. Section 3 of every file
   (the analogy) exists partly for this.

5. STAY CALIBRATED
   Re-read Section XIV.01 when a new technique appears.
   Re-measure your `numbers.md` on new hardware.
   Re-run your capacity plan quarterly.

9. Hands-on exercise#

A. Project the hierarchy. For a hypothetical system with coherent 900 GB/s CPU-GPU memory and 1 TB of CPU DRAM, recompute: max concurrency for a 70B model at 128k context, and whether KV offloading becomes the default. What changes?

B. The hybrid projection. For a hybrid SSM/attention model with 1 attention layer per 8, compute the KV cache at 1M context. Compare to pure attention. What does it do to concurrency?

C. Revisit your priorities. Take your current optimization backlog. For each of the four likely changes in section 6, note which items become more or less valuable.

D. The unchanging test. For each of the seven things in section 7, find an example in this curriculum where it applied. Then find a case in your own work.

E. Write your own section 6. Based on your reading and your hardware roadmap, write your own ranked list of what would change the field. Revisit it in a year and see how you did.


10. Interview questions#

  1. Why is inference fundamentally a memory problem?
  2. What does coherent CPU-GPU memory change about inference architecture?
  3. What is disaggregation, generalized, and what’s the counter-argument?
  4. Why is LLM decode a good match for processing-in-memory?
  5. What would hybrid SSM/attention models change about long-context serving?
  6. What stays true regardless of hardware changes?
  7. If you had to bet on one change transforming inference in the next three years, what would it be and why?

11. Further reading#

  • [REFERENCE] CXL specification and consortium materials
  • [RESEARCH] HBM-PIM and near-memory computing literature (Samsung, SK Hynix, UPMEM)
  • [EMERGING] Grace-Hopper and coherent-interconnect architecture documentation
  • [EMERGING] Qin et al., “Mooncake” — the KV-centric framing
  • [FUNDAMENTAL] Wulf & McKee, “Hitting the Memory Wall” (1995) — where this all started
  • Back to: README.md | Projects: projects/

↑↓ navigate↵ openesc close