PidokuInfra

GPU Engineering

The GPU itself — what is inside the chip, how it is programmed, why it is fast, and how thousands are run together.

Start reading
  • 5 levels
  • 7 modules
  • 32 topics
  • ~27h 15m

Contents

  1. Foundations

    Build the mental model.

    4 topics · ~2h 30m

    1. 01 Why GPUs ExistA GPU is not a faster CPU. It is a different bargain: give up speed on any single task, and in exchange do an enormous number of simple tasks at the same time. 4 topics · ~2h 30m
      1. 00Module overview
      2. 01What Is a GPU?Beginner30 min
      3. 02Latency vs ThroughputBeginner40 min
      4. 03From Pixels to TensorsBeginner35 min
      5. 04Reading a Spec SheetBeginner45 min
  2. Basic

    Understand the core mechanisms.

    5 topics · ~4h 10m

    1. 02 Inside the ChipModule I said a GPU has "thousands of lanes" and "fast memory". This module opens the lid and names every part, using NVIDIA's H100 as the worked example. 5 topics · ~4h 10m
      1. 00Module overview
      2. 01SMs, Warps and SIMTBeginner1h
      3. 02The Memory HierarchyBeginner50 min
      4. 03HBM and Memory BandwidthIntermediate45 min
      5. 04Tensor Cores and Number FormatsIntermediate55 min
      6. 05Power, Heat and the Host LinkIntermediate40 min
  3. Intermediate

    Learn the optimization techniques.

    10 topics · ~9h 45m

    1. 03 The Programming ModelYou know what the hardware is. Now: how do you tell it what to do? 5 topics · ~4h 50m
      1. 00Module overview
      2. 01Kernels, Threads, Blocks, GridsIntermediate1h
      3. 02Your First Kernel, Called from GoIntermediate1h 30m
      4. 03Host and Device MemoryIntermediate50 min
      5. 04Asynchronous Execution and StreamsIntermediate50 min
      6. 05The Software StackIntermediate40 min
    2. 04 PerformanceA GPU can be a thousand times faster than a CPU or barely faster at all. The difference is almost never "better arithmetic". It is whether the program respects four constraints: memory … 5 topics · ~4h 55m
      1. 00Module overview
      2. 01The Roofline ModelIntermediate1h 10m
      3. 02Memory Access PatternsIntermediate1h
      4. 03Occupancy and DivergenceIntermediate55 min
      5. 04Launch Overhead and FusionIntermediate50 min
      6. 05Measuring a GPUIntermediate1h
  4. Advanced

    Study systems at production scale.

    8 topics · ~7h 50m

    1. 05 GPUs for AIModules II–IV were about GPUs in general. This module applies them to the workload that now buys most GPUs: neural networks, and large language models in particular. 4 topics · ~4h 5m
      1. 00Module overview
      2. 01Matrix Multiplication on a GPUIntermediate1h 15m
      3. 02Precision and Quantization on HardwareIntermediate1h
      4. 03Memory Planning for LLMsIntermediate1h 10m
      5. 04Training vs Inference on a GPUIntermediate40 min
    2. 06 Multi-GPU and SharingOne GPU is a component. Real systems have two opposite problems: a job too big for one GPU, and a GPU too big for one job. This module covers both, and how a cluster scheduler hands GPUs … 4 topics · ~3h 45m
      1. 00Module overview
      2. 01Interconnects and TopologyAdvanced55 min
      3. 02Splitting Work Across GPUsAdvanced1h
      4. 03Sharing One GPUAdvanced50 min
      5. 04GPUs in KubernetesAdvanced1h
  5. Expert

    Design platforms and read the frontier.

    4 topics · ~3h

    1. 07 Data Center and FrontierThe last module zooms out: from one chip to product generations, to competing designs, to the buildings full of them, and finally to where the technology is heading. 4 topics · ~3h
      1. 00Module overview
      2. 01GPU GenerationsAdvanced45 min
      3. 02Other AcceleratorsAdvanced50 min
      4. 03Racks, Power and CostAdvanced50 min
      5. 04Where This Is GoingAdvanced35 min

Build

Reference

About

Understand the GPU completely: what it is, how it is built, how you program it, why it is fast, why it is sometimes slow, and how thousands of them are run together.

This course starts at “what is a GPU?” and ends at “plan the GPUs, network and power for a cluster”. It assumes you can program in Go. It does not assume any hardware, graphics, C or machine-learning background.


How this course works#

Every lesson follows the same shape, so you always know where you are:

  1. The idea in one minute — the whole lesson in a few sentences.
  2. An analogy — something from everyday life with the same shape.
  3. A picture — a hand-drawn diagram of the mechanism.
  4. How it really works — the precise version, with real numbers.
  5. Code — a small Go program you can run, usually without owning a GPU.
  6. Remember this — the three or four facts worth keeping.
  7. Try it and Check yourself — exercises and questions.

Why Go, when GPUs are programmed in C?#

A GPU runs kernels — small functions written in CUDA C (or generated by a compiler). That will not change, and this course shows you those kernels in CUDA C where they matter.

Everything around the kernel is ordinary systems programming, and Go is very good at it:

  • Models and simulators. Most GPU behaviour — warps, coalescing, the roofline, memory planning — can be reproduced in fifty lines of Go and run on a laptop. You learn the mechanism by building it.
  • Talking to the GPU. NVIDIA’s own Kubernetes device plugin, GPU operator and container toolkit are written in Go, on top of the go-nvml bindings. Reading GPU state from Go is a first-class path.
  • Calling kernels. cgo lets a Go program call CUDA C directly. You will do this in module III.
  • Serving. The layers above the GPU — schedulers, gateways, autoscalers — are where most Go engineers meet GPUs, and they are the subject of the sister course, Inference Engineering.
flowchart LR
  subgraph GO["Written in Go"]
    APP["Your service<br/>scheduler, gateway"]
    SIM["Simulators and<br/>capacity models"]
    MON["Monitoring<br/>go-nvml"]
  end
  subgraph C["Written in CUDA C"]
    K["Kernels<br/>the code that runs on the GPU"]
  end
  DRV["NVIDIA driver"]
  GPU[("GPU")]
  APP -->|"cgo"| K
  MON --> DRV
  K --> DRV --> GPU
  class APP,SIM,MON compute
  class K io
  class DRV neutral
  class GPU memory

The path#

flowchart LR
  I["I Why GPUs exist"] --> II["II Inside the chip"]
  II --> III["III Programming model"]
  III --> IV["IV Performance"]
  IV --> V["V GPUs for AI"]
  V --> VI["VI Multi-GPU and sharing"]
  VI --> VII["VII Data center and frontier"]
  class I,II neutral
  class III,IV compute
  class V memory
  class VI queue
  class VII io
ModuleYou will be able toLevel
I — Why GPUs ExistExplain what a GPU is, why it beats a CPU at some work and loses at other work, and read a spec sheetBeginner
II — Inside the ChipDescribe SMs, warps, the memory hierarchy, HBM and tensor cores, with numbersBeginner → Intermediate
III — The Programming ModelWrite a CUDA kernel, call it from Go, and manage host/device memory and streamsIntermediate
IV — PerformancePredict speed with the roofline, fix slow memory access, and profile a GPUIntermediate
V — GPUs for AIExplain how matmul, precision and memory planning decide AI performanceIntermediate → Advanced
VI — Multi-GPU and SharingReason about NVLink, collectives, MPS/MIG and GPUs in KubernetesAdvanced
VII — Data Center and FrontierCompare GPU generations and other accelerators; plan power, cooling and costAdvanced
ProjectsBuild five Go programs that make the ideas stickAll

What you need#

  • Go 1.22 or newer.
  • A terminal.
  • No GPU required for about 90% of the course. Lessons that need real hardware say so and give a no-GPU alternative. A free Colab T4 or any rented NVIDIA card covers the rest.

A promise about numbers#

GPU specifications change every year. This course teaches you the mechanisms, which do not change, and uses real parts (mostly NVIDIA’s H100, because it is the best documented) as worked examples. When a number is a vendor announcement rather than something measured, the lesson says so. Always check the current data sheet before spending money.

Product facts were last checked on 3 October 2026: Rubin (HBM4) had been shipping since August, Blackwell and Blackwell Ultra made up most new capacity, and Hopper remained the largest installed base. VII.01 holds the current table.

Where to go next#

↑↓ navigate↵ openesc close