PidokuInfra
13Advanced LLM Inference
On this page

Advanced LLM Inference

ExpertModule 1312 topics~15h 15m

Topics, in order

01 MQA, GQA, and MLAThe KV cache, not the weights, is the capacity limit — and it's determined by an architectural choice made at training time. Advanced 1h 15m 02 Mixture of ExpertsReplace the dense FFN in each transformer layer with E expert FFNs plus a router. Each token is sent to only k of them (typically k=1 or 2). Advanced 1h 30m 03 MoE Serving and Expert ParallelismSection IX.05 covered expert parallelism as a distributed systems problem. This file covers the serving engineering: kernels, scheduling, and what actually makes MoE fast. Advanced 1h 15m 04 Long-Context InferenceSection V.07 established them; this file is about the techniques that attack each. Advanced 1h 30m 05 Chunked PrefillThe trade is explicit and favorable: one request's TTFT degrades 25% so that dozens of requests' ITL improves 12x. Advanced 1h 06 Disaggregated Prefill/DecodeRun prefill and decode on separate GPU pools, and transfer the KV cache between them. Advanced 1h 30m 07 KV Offloading and TransferMove KV cache blocks out of GPU memory to a cheaper, larger tier, and fetch them back when needed. Advanced 1h 15m 08 Request Scheduling and PrioritiesSection V.09 covered the mechanism. This file covers policy — the part that determines whether your multi-tenant platform is fair, whether your SLO tiers mean anything, and whether long … Advanced 1h 09 Speculative Decoding VariantsSection V.12 covered the algorithm; VII.13 covered deployment. This file surveys the variants and gives an honest assessment of each. Advanced 1h 15m 10 Low-Bit Inference: FP4 and BeyondThe frontier is at 4 bits with hardware support. Below that, the techniques change character: you stop quantizing a pretrained model and start training for low precision. Advanced 1h 11 Triton KernelsA Python-embedded language for writing GPU kernels, where you program at the block level rather than the thread level. The compiler handles the thread-level details. Advanced 1h 30m 12 Compilers: torch.compile and BeyondSection IV.10 covered the theory. This file covers the practical use of torch.compile for inference and the landscape around it. Advanced 1h 15m

About this module

Goal: the techniques that define the current frontier of production inference — and an honest assessment of which are ready.

Everything here is labelled:

  • [ESTABLISHED] — deployed at scale by multiple organizations
  • [EMERGING] — promising, deployed by some, still maturing
  • [RESEARCH] — interesting, not production-ready

Read the labels. Several techniques in this section are frequently presented as ready when they aren’t.

Files#

#FileLevelStatusTime
01MQA, GQA, MLAAdvancedESTABLISHED75 min
02Mixture of ExpertsAdvancedESTABLISHED90 min
03MoE serving and expert parallelismAdvancedESTABLISHED75 min
04Long-context inferenceAdvancedmixed90 min
05Chunked prefillAdvancedESTABLISHED60 min
06Disaggregated prefill/decodeAdvancedEMERGING90 min
07KV offloading and transferAdvancedEMERGING75 min
08Request scheduling and prioritiesAdvancedmixed60 min
09Speculative decoding variantsAdvancedmixed75 min
10Low-bit inference: FP4 and beyondAdvancedEMERGING60 min
11Triton kernelsAdvancedESTABLISHED90 min
12Compilers: torch.compile and beyondAdvancedESTABLISHED75 min

The thread#

flowchart TD
  N0["Shrink the KV cache architecturally<br/><b>01</b>"]
  N1["Or shrink the ACTIVE parameters per token<br/><b>02, 03</b>"]
  N2["Which matters most at long context<br/><b>04</b>"]
  N3["Where prefill must be chunked to be usable<br/><b>05</b>"]
  N4["Or separated entirely onto different hardware<br/><b>06</b>"]
  N5["With KV moving between tiers<br/><b>07</b>"]
  N6["Under a scheduler that knows about priorities<br/><b>08</b>"]
  N7["Producing multiple tokens per pass where possible<br/><b>09</b>"]
  N8["In as few bits as the hardware supports<br/><b>10</b>"]
  N9["Using kernels you can now write yourself<br/><b>11</b>"]
  N10["Or that a compiler writes for you<br/><b>12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10

  class N0,N1,N2 neutral
  class N3,N4 io
  class N5,N6 queue
  class N7,N8 compute
  class N9,N10 memory

How to read this section#

If you’re operating a system: files 01, 02, 05, and 08 affect decisions you make today. Files 04 and 10 affect model selection. The rest are for when you hit their specific bottleneck.

If you’re building: files 11 and 12 are the practical skills. 06 and 07 are the architectures to understand before you need them.

↑↓ navigate↵ openesc close