Everything so far applies to any service. This module is about what changes when the service is a model on a GPU: new resources to watch, new units of work, new failure modes — and a product whose output can be wrong while every infrastructure metric is green.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | What Is Different | Which assumptions from web-service observability break? |
| 02 | GPU Telemetry | What does a GPU report, and which numbers lie? |
| 03 | Inference Engine Metrics | What should an LLM server expose, and how do I read it? |
| 04 | Tracing LLM Requests | How do I trace a model call, a tool call, an agent? |
| 05 | Platform and Fleet | What do I watch across a cluster, and what do I scale on? |
| 06 | Cost, Power and Efficiency | What does a token cost in dollars and joules? |
| 07 | Quality and Evals | How do I observe whether the answers are any good? |
When you finish you can list the metrics that matter at each layer of an inference platform, explain why GPU utilization is not one of them, and compute cost per million tokens from counters.
This module refers often to Inference Engineering (prefill, decode, KV cache, continuous batching) and GPU Engineering (memory bandwidth, power). Each lesson links the specific background it needs.