PidokuInfra

RESOURCES

Curated, organized by section. Each entry is labeled:

  • [FUNDAMENTAL] — timeless; the physics/math doesn’t change
  • [ESTABLISHED] — production-proven engineering
  • [EMERGING] — promising, not universally production-ready
  • [REFERENCE] — documentation you will return to

Prefer primary sources. When a blog post and a paper disagree, read the paper; when the paper and your profiler disagree, believe the profiler.


Cross-cutting / start here#


I — Fundamentals#

  • [FUNDAMENTAL] Williams, Waterman, Patterson, “Roofline: An Insightful Visual Performance Model for Multicore Architectures” (CACM 2009) — the origin of arithmetic intensity as a design tool.
  • [ESTABLISHED] Kaplan et al., “Scaling Laws for Neural Language Models” (2020) — read for the FLOPs-per-token accounting, not the scaling conclusions.
  • [ESTABLISHED] “LLM Inference Performance Engineering: Best Practices” — Databricks engineering blog. Good first pass on TTFT/ITL thinking.
  • [FUNDAMENTAL] Latency Numbers Every Programmer Should Know (Jeff Dean / Peter Norvig tables) — internalize the orders of magnitude.

II — Computer Systems Foundations#

  • [FUNDAMENTAL] Ulrich Drepper, “What Every Programmer Should Know About Memory” (2007) — still the best treatment of caches and NUMA. Long; read parts 2, 3, 5.
  • [FUNDAMENTAL] Brendan Gregg, Systems Performance (2nd ed.) — the USE method, perf, flame graphs.
  • [REFERENCE] Brendan Gregg’s site — https://www.brendangregg.com/ (flame graphs, eBPF)
  • [REFERENCE] man 7 numa, numactl(8), perf-stat(1)
  • [REFERENCE] PCI-SIG specifications overview; for practical numbers, NVIDIA’s bandwidthTest sample in cuda-samples.
  • [FUNDAMENTAL] Bovet & Cesati, Understanding the Linux Kernel — for scheduling and VM.

III — Machine Learning Fundamentals#

  • [FUNDAMENTAL] Vaswani et al., “Attention Is All You Need” (2017) — the transformer paper.
  • [FUNDAMENTAL] Jay Alammar, “The Illustrated Transformer” — the best visual introduction.
  • [FUNDAMENTAL] Andrej Karpathy, “Let’s build GPT: from scratch, in code, spelled out” (video) and nanoGPT — https://github.com/karpathy/nanoGPT
  • [FUNDAMENTAL] 3Blue1Brown, Essence of Linear Algebra (video series).
  • [ESTABLISHED] Micikevicius et al., “Mixed Precision Training” (2018) — where FP16 loss scaling and BF16 reasoning come from.
  • [ESTABLISHED] Micikevicius et al., “FP8 Formats for Deep Learning” (2022) — E4M3/E5M2.
  • [REFERENCE] IEEE 754 and the bfloat16 numerics note from Google Brain.

IV — Neural Network Inference#

  • [REFERENCE] NVIDIA cuBLAS docs; the “Matrix Multiplication Background User’s Guide” in NVIDIA Deep Learning Performance documentation — explains tiling and tile quantization.
  • [ESTABLISHED] “Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking” (Jia et al.) — how people actually determine what hardware does.
  • [REFERENCE] ONNX operator specification — https://onnx.ai/onnx/operators/
  • [ESTABLISHED] Chen et al., “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning” (OSDI 2018) — graph + operator level optimization.
  • [REFERENCE] PyTorch torch.compile / TorchInductor documentation.

V — LLM Inference (the core section)#

  • [ESTABLISHED] Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention” (SOSP 2023) — the vLLM paper. Read this one twice.
  • [ESTABLISHED] Yu et al., “Orca: A Distributed Serving System for Transformer-Based Generative Models” (OSDI 2022) — the origin of iteration-level (continuous) batching.
  • [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022) — the canonical analytical treatment of prefill/decode, partitioning, and the arithmetic.
  • [ESTABLISHED] Leviathan et al., “Fast Inference from Transformers via Speculative Decoding” (2022); Chen et al., “Accelerating Large Language Model Decoding with Speculative Sampling” (2023).
  • [ESTABLISHED] Shazeer, “Fast Transformer Decoding: One Write-Head is All You Need” (2019) — MQA, and the clearest statement of why KV bandwidth dominates decode.
  • [ESTABLISHED] Ainslie et al., “GQA: Training Generalized Multi-Query Transformer Models” (2023).
  • [REFERENCE] vLLM design docs on the scheduler and block manager.
  • [ESTABLISHED] Holtzman et al., “The Curious Case of Neural Text Degeneration” (2019) — where top-p (nucleus) sampling comes from.

VI — GPU Computing#

  • [REFERENCE] CUDA C++ Programming Guide (again — chapters 2, 5, and the appendix on compute capabilities).
  • [FUNDAMENTAL] NVIDIA whitepapers per architecture: Volta, Ampere (A100), Hopper (H100), Blackwell. Each explains the SM, memory system, and tensor cores of that generation.
  • [FUNDAMENTAL] “CUDA Refresher” blog series on NVIDIA Developer Blog.
  • [ESTABLISHED] Mark Harris, “How to Optimize Data Transfers in CUDA C/C++” and the “CUDA Pro Tip” series.
  • [REFERENCE] Nsight Systems and Nsight Compute user guides.
  • [FUNDAMENTAL] Programming Massively Parallel Processors (Kirk & Hwu), 4th ed. — the standard textbook.
  • [REFERENCE] CUTLASS — https://github.com/NVIDIA/cutlass — how production GEMMs are built.

VII — Inference Optimization#

  • [ESTABLISHED] Dao et al., “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness” (2022).
  • [ESTABLISHED] Dao, “FlashAttention-2” (2023).
  • [EMERGING→ESTABLISHED] Shah et al., “FlashAttention-3” (2024) — Hopper-specific (async, FP8).
  • [ESTABLISHED] Frantar et al., “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers” (2022).
  • [ESTABLISHED] Lin et al., “AWQ: Activation-aware Weight Quantization” (2023).
  • [ESTABLISHED] Xiao et al., “SmoothQuant” (2022).
  • [ESTABLISHED] Dettmers et al., “LLM.int8()” (2022) — outlier features, and why naive INT8 fails.
  • [REFERENCE] TensorRT-LLM docs and its docs/source/performance pages.
  • [REFERENCE] NVIDIA Transformer Engine (FP8) documentation.
  • [EMERGING] “KIVI”, “KVQuant” and related KV-cache quantization papers.

VIII — Inference Serving Systems#

  • [ESTABLISHED] vLLM source: vllm/core/scheduler.py, vllm/core/block_manager*, vllm/attention/. Reading real schedulers beats reading about schedulers.
  • [ESTABLISHED] Zheng et al., “SGLang: Efficient Execution of Structured Language Model Programs” (2023) — RadixAttention / prefix caching.
  • [REFERENCE] Hugging Face TGI architecture docs.
  • [REFERENCE] NVIDIA Triton Inference Server docs — especially dynamic batching and the backend API.
  • [REFERENCE] ONNX Runtime performance tuning docs.
  • [FUNDAMENTAL] “The Tail at Scale” (Dean & Barroso, CACM 2013) — required reading for anyone owning a latency SLO.
  • [FUNDAMENTAL] Little’s Law — any queueing theory primer. L = λW.

IX — Distributed Inference#

  • [ESTABLISHED] Shoeybi et al., “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism” (2019) — where TP layouts come from.
  • [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022) — again; it is the best partitioning analysis available.
  • [REFERENCE] NCCL documentation and the NCCL tests repo (nccl-tests) for measuring real collective bandwidth.
  • [REFERENCE] NVIDIA NVLink / NVSwitch technical briefs.
  • [FUNDAMENTAL] Any treatment of the ring-allreduce algorithm (Baidu’s original writeup).
  • [ESTABLISHED] Huang et al., “GPipe” (2018) — pipeline bubbles.

X — Memory & Performance Engineering#

  • [FUNDAMENTAL] Roofline paper (see I).
  • [REFERENCE] Nsight Compute metrics reference — learn dram__throughput, sm__throughput, l2_tex__t_bytes.
  • [FUNDAMENTAL] Brendan Gregg, flame graphs; py-spy for Python-side stalls.
  • [REFERENCE] PyTorch Profiler and torch.cuda.memory_summary() / memory snapshot tooling.
  • [REFERENCE] DCGM exporter + Prometheus + Grafana for GPU fleet metrics.
  • [ESTABLISHED] “USE Method” (Utilization, Saturation, Errors) — Gregg.

XI — Production Inference Engineering#

  • [FUNDAMENTAL] Google SRE Book, chapters on SLOs, overload, and handling cascading failures — https://sre.google/books/
  • [FUNDAMENTAL] “The Tail at Scale” (again).
  • [ESTABLISHED] Public engineering writeups on LLM cost-per-token modeling (treat specific numbers as dated; take the method).
  • [REFERENCE] OpenTelemetry semantic conventions for GenAI spans.
  • [REFERENCE] OWASP Top 10 for LLM Applications — for the security/abuse section.

XII — Inference Platform Engineering#

  • [REFERENCE] Kubernetes docs: device plugins, scheduling framework, topology manager, ResourceQuota, HPA/KEDA.
  • [REFERENCE] NVIDIA GPU Operator and DCGM.
  • [ESTABLISHED] KServe and Ray Serve architecture docs — two different answers to the same problem.
  • [EMERGING] Kubernetes Gateway API Inference Extension — model-aware routing.
  • [FUNDAMENTAL] “Control plane vs data plane” as articulated in networking literature; the distinction transfers exactly.

XIII — Advanced LLM Inference#

  • [ESTABLISHED] Fedus et al., “Switch Transformers” (2021); Lepikhin et al., “GShard” (2020) — MoE routing.
  • [ESTABLISHED] DeepSeek-V2 / V3 technical reports — MLA, MoE at scale, and unusually candid inference engineering detail.
  • [ESTABLISHED] Agrawal et al., “Sarathi-Serve: Taming Throughput-Latency Tradeoff in LLM Inference” (2024) — chunked prefill.
  • [EMERGING] Zhong et al., “DistServe: Disaggregating Prefill and Decoding” (2024); Splitwise (Microsoft, 2024).
  • [EMERGING] Cai et al., “Medusa” (2024); Li et al., “EAGLE” / “EAGLE-2” (2024).
  • [REFERENCE] OpenAI Triton language docs — https://triton-lang.org/
  • [REFERENCE] torch.compile internals; TorchInductor design notes.

XIV — Research & Frontier#

  • [REFERENCE] arXiv cs.LG + cs.DC — filter for “inference”, “serving”, “KV cache”.
  • [REFERENCE] MLSys, OSDI, SOSP, ASPLOS, ISCA proceedings — the venues where inference systems work actually lands.
  • [REFERENCE] MLPerf Inference results — https://mlcommons.org/benchmarks/inference/ — the closest thing to apples-to-apples hardware comparison.
  • [REFERENCE] AWS Neuron SDK docs (Inferentia/Trainium); Google Cloud TPU docs; AMD ROCm / composable_kernel docs.
  • [EMERGING] Processing-in-memory and near-memory literature (UPMEM, HBM-PIM papers).

2025-2026 additions#

Last reviewed: 2026-10-03. The sections above are the durable core. This block is the part that ages: newer papers, engine write-ups, and the feeds worth following. Everything here is [EMERGING] unless marked otherwise — apply the paper-reading rules at the bottom of this file, especially to vendor blogs, which report their best configuration.

The book-length treatment#

V / VIII — Engine internals (read alongside the vLLM and SGLang sections)#

VI / VII — Kernels, compilers, quantization#

V.12 / VII.13 / XIII.09 — Speculative decoding#

IX / XIII — Disaggregation, wide expert parallelism, KV as a managed resource#

The largest shift since this curriculum’s core papers: prefill/decode disaggregation and tiered KV memory moved from research into the default large-scale configuration.

XI / XII — Platform, routing, Kubernetes#

X / XIV — Benchmarks and hardware#

Observability#

Status changes worth knowing#

  • TGI is archived. Hugging Face moved Text Generation Inference to maintenance mode in December 2025 and archived the repository in March 2026; its docs recommend vLLM or SGLang.
  • vLLM was at v0.30 (22 September 2026) at the time of review.

Feeds to follow#

Check these roughly weekly; they are where changes to this curriculum’s [EMERGING] tier show up first.

FeedWhyURL
vLLM blogEngine internals, new techniques landing in productionhttps://vllm.ai/blog
SGLang blog + releasesThe other reference engine; often first on cache and P/D workhttps://www.sglang.io/blog · https://github.com/sgl-project/sglang/releases
llm-d blogKubernetes-side routing, scheduling, autoscalinghttps://llm-d.ai/blog
NVIDIA Developer blogHardware, Dynamo, TensorRT-LLM, CUDAhttps://developer.nvidia.com/blog
Tri Dao’s blogAttention kernels from the sourcehttps://tridao.me/blog/
InferenceXWhether last month’s claims held up on real hardwarehttps://inferencex.semianalysis.com/
SemiAnalysis newsletterHardware economics and roadmapshttps://newsletter.semianalysis.com/
Baseten blogPractitioner write-ups from a serving providerhttps://www.baseten.co/blog/
Brendan GreggPerformance methodology, now applied to AI systemshttps://www.brendangregg.com/
arXiv cs.DC + cs.LGFilter: “LLM serving”, “KV cache”, “speculative decoding”https://arxiv.org/list/cs.DC/recent
MLSys / OSDI / SOSP / ASPLOS programsWhere serving-systems work is peer reviewed—

Community study repos#


How to read a paper in this field#

  1. Read the abstract, then jump to the evaluation setup (hardware, model, batch sizes, sequence lengths). If the setup does not resemble your workload, the speedup number does not transfer.
  2. Find the baseline. Half of reported speedups are against a weak baseline.
  3. Find the bottleneck they claim to remove and check it against the roofline. If a paper claims a 3x decode speedup without reducing bytes moved or improving arithmetic intensity, be suspicious.
  4. Ask what breaks at long context, large batch, or under multi-tenancy. Papers optimize for one regime.

See XIV-research-and-frontier/01-reading-papers.md for the long version.

↑↓ navigate↵ openesc close