Understand the GPU completely: what it is, how it is built, how you program it, why it is fast, why it is sometimes slow, and how thousands of them are run together.
This course starts at “what is a GPU?” and ends at “plan the GPUs, network and power for a cluster”. It assumes you can program in Go. It does not assume any hardware, graphics, C or machine-learning background.
How this course works#
Every lesson follows the same shape, so you always know where you are:
- The idea in one minute — the whole lesson in a few sentences.
- An analogy — something from everyday life with the same shape.
- A picture — a hand-drawn diagram of the mechanism.
- How it really works — the precise version, with real numbers.
- Code — a small Go program you can run, usually without owning a GPU.
- Remember this — the three or four facts worth keeping.
- Try it and Check yourself — exercises and questions.
Why Go, when GPUs are programmed in C?#
A GPU runs kernels — small functions written in CUDA C (or generated by a compiler). That will not change, and this course shows you those kernels in CUDA C where they matter.
Everything around the kernel is ordinary systems programming, and Go is very good at it:
- Models and simulators. Most GPU behaviour — warps, coalescing, the roofline, memory planning — can be reproduced in fifty lines of Go and run on a laptop. You learn the mechanism by building it.
- Talking to the GPU. NVIDIA’s own Kubernetes device plugin, GPU operator and container
toolkit are written in Go, on top of the
go-nvmlbindings. Reading GPU state from Go is a first-class path. - Calling kernels.
cgolets a Go program call CUDA C directly. You will do this in module III. - Serving. The layers above the GPU — schedulers, gateways, autoscalers — are where most Go engineers meet GPUs, and they are the subject of the sister course, Inference Engineering.
flowchart LR
subgraph GO["Written in Go"]
APP["Your service<br/>scheduler, gateway"]
SIM["Simulators and<br/>capacity models"]
MON["Monitoring<br/>go-nvml"]
end
subgraph C["Written in CUDA C"]
K["Kernels<br/>the code that runs on the GPU"]
end
DRV["NVIDIA driver"]
GPU[("GPU")]
APP -->|"cgo"| K
MON --> DRV
K --> DRV --> GPU
class APP,SIM,MON compute
class K io
class DRV neutral
class GPU memoryThe path#
flowchart LR I["I Why GPUs exist"] --> II["II Inside the chip"] II --> III["III Programming model"] III --> IV["IV Performance"] IV --> V["V GPUs for AI"] V --> VI["VI Multi-GPU and sharing"] VI --> VII["VII Data center and frontier"] class I,II neutral class III,IV compute class V memory class VI queue class VII io
| Module | You will be able to | Level |
|---|---|---|
| I — Why GPUs Exist | Explain what a GPU is, why it beats a CPU at some work and loses at other work, and read a spec sheet | Beginner |
| II — Inside the Chip | Describe SMs, warps, the memory hierarchy, HBM and tensor cores, with numbers | Beginner → Intermediate |
| III — The Programming Model | Write a CUDA kernel, call it from Go, and manage host/device memory and streams | Intermediate |
| IV — Performance | Predict speed with the roofline, fix slow memory access, and profile a GPU | Intermediate |
| V — GPUs for AI | Explain how matmul, precision and memory planning decide AI performance | Intermediate → Advanced |
| VI — Multi-GPU and Sharing | Reason about NVLink, collectives, MPS/MIG and GPUs in Kubernetes | Advanced |
| VII — Data Center and Frontier | Compare GPU generations and other accelerators; plan power, cooling and cost | Advanced |
| Projects | Build five Go programs that make the ideas stick | All |
What you need#
- Go 1.22 or newer.
- A terminal.
- No GPU required for about 90% of the course. Lessons that need real hardware say so and give a no-GPU alternative. A free Colab T4 or any rented NVIDIA card covers the rest.
A promise about numbers#
GPU specifications change every year. This course teaches you the mechanisms, which do not change, and uses real parts (mostly NVIDIA’s H100, because it is the best documented) as worked examples. When a number is a vendor announcement rather than something measured, the lesson says so. Always check the current data sheet before spending money.
Product facts were last checked on 3 October 2026: Rubin (HBM4) had been shipping since August, Blackwell and Blackwell Ultra made up most new capacity, and Hopper remained the largest installed base. VII.01 holds the current table.
Where to go next#
- Golang Engineering — the language used throughout, in depth.
- Inference Engineering — the serving systems built on these GPUs.
- Observability Engineering — how to measure them in production: GPU telemetry, power, and cost per token.