Before any tool: what are you trying to learn about a running system, and how do you state “working” precisely enough that a machine can check it?
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Monitoring vs Observability | What is the difference, and why does it matter? |
| 02 | The Signals | What kinds of telemetry exist, and what is each one good for? |
| 03 | SLIs, SLOs and Error Budgets | How do you write down “good enough” as a number? |
| 04 | Percentiles and Tails | Why is the average the wrong number, and what replaces it? |
When you finish you can look at any service and say which three numbers describe whether its users are happy, and why none of them is an average.