A GPU can be a thousand times faster than a CPU or barely faster at all. The difference is almost never “better arithmetic”. It is whether the program respects four constraints: memory bandwidth, memory access patterns, how well the SMs are kept busy, and per-launch overhead.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | The Roofline Model | What is the fastest this operation could possibly run? |
| 02 | Memory Access Patterns | Why does the order of memory access change speed 10x? |
| 03 | Occupancy and Divergence | Are the SMs actually busy? |
| 04 | Launch Overhead and Fusion | Why are many small kernels slow? |
| 05 | Measuring a GPU | What do the tools tell you — and what do they hide? |
When you finish you can look at an operation, predict its speed from a data sheet, measure it, and explain the gap.