PidokuInfra

Observability Engineering

Seeing inside running systems — metrics, logs, traces and profiles, then the GPU, token and cost telemetry that AI infrastructure adds.

Start reading
  • 5 levels
  • 6 modules
  • 30 topics
  • ~24h 25m

Contents

  1. Foundations

    Build the mental model.

    4 topics · ~2h 40m

    1. 01 Why ObservabilityBefore any tool: what are you trying to learn about a running system, and how do you state "working" precisely enough that a machine can check it? 4 topics · ~2h 40m
      1. 00Module overview
      2. 01Monitoring vs ObservabilityBeginner30 min
      3. 02The SignalsBeginner40 min
      4. 03SLIs, SLOs and Error BudgetsBeginner45 min
      5. 04Percentiles and TailsBeginner45 min
  2. Basic

    Understand the core mechanisms.

    9 topics · ~7h 30m

    1. 02 MetricsMetrics are the cheapest signal and the one every alert is built on. This module builds a metrics library by hand, then uses the real one, then learns to query it without being fooled. 5 topics · ~4h 5m
      1. 00Module overview
      2. 01Metric Types and the Data ModelBeginner45 min
      3. 02Instrumenting a Go ServiceBeginner50 min
      4. 03PromQL EssentialsBeginner55 min
      5. 04HistogramsIntermediate55 min
      6. 05CardinalityIntermediate40 min
    2. 03 Logs, Traces and ProfilesMetrics tell you something is wrong. These three signals keep the detail needed to say what. 4 topics · ~3h 25m
      1. 00Module overview
      2. 01Structured Logs and Wide EventsBeginner40 min
      3. 02Distributed TracingIntermediate55 min
      4. 03OpenTelemetryIntermediate1h
      5. 04Profiling and eBPFIntermediate50 min
  3. Intermediate

    Learn the optimization techniques.

    5 topics · ~4h 10m

    1. 04 Running ItEmitting telemetry is the easy half. This module is about the system that receives it, the screens and alerts built on it, what it costs, and how to use it when something is on fire. 5 topics · ~4h 10m
      1. 00Module overview
      2. 01The Telemetry PipelineIntermediate55 min
      3. 02DashboardsIntermediate40 min
      4. 03Alerting on SLOsIntermediate55 min
      5. 04Sampling and CostIntermediate50 min
      6. 05Debugging an IncidentIntermediate50 min
  4. Advanced

    Study systems at production scale.

    7 topics · ~6h 30m

    1. 05 AI InfrastructureEverything so far applies to any service. This module is about what changes when the service is a model on a GPU: new resources to watch, new units of work, new failure modes — and a product … 7 topics · ~6h 30m
      1. 00Module overview
      2. 01What Is DifferentAdvanced45 min
      3. 02GPU TelemetryAdvanced1h
      4. 03Inference Engine MetricsAdvanced1h 10m
      5. 04Tracing LLM RequestsAdvanced55 min
      6. 05Platform and FleetAdvanced1h
      7. 06Cost, Power and EfficiencyAdvanced50 min
      8. 07Quality and EvalsAdvanced50 min
  5. Expert

    Design platforms and read the frontier.

    4 topics · ~3h 35m

    1. 06 Tools and FrontierThe mechanisms in modules I–V are stable. The products that implement them are not: this module maps what exists, explains the direction the field is moving, and gives you a method for … 4 topics · ~3h 35m
      1. 00Module overview
      2. 01The Tool LandscapeAdvanced50 min
      3. 02How It Is ChangingExpert1h
      4. 03What to WatchExpert45 min
      5. 04Designing Observability for an Inference PlatformExpert1h

Build

Reference

About

Learn to see inside running systems: what to measure, how the measurements are collected and stored, how to turn them into decisions — and what changes when the system is a rack of GPUs generating tokens.

This course starts at “what is a metric?” and ends at “design the telemetry for a multi-tenant inference platform”. It assumes you can program in Go. It does not assume you have run Prometheus, read a trace, or been on call.

It leans on two other paths for the systems being observed — Inference Engineering and GPU Engineering — and on Golang Engineering if you want the language itself in depth. You can take modules I–IV without any of them.


How this course works#

Every lesson follows the same shape:

  1. The idea in one minute — the whole lesson in a few sentences.
  2. An analogy — something from everyday life with the same shape.
  3. A picture — a hand-drawn diagram of the mechanism.
  4. How it really works — the precise version, with real names and numbers.
  5. Code — a small Go program you can run. Almost all of them use only the standard library.
  6. Remember this — the three or four facts worth keeping.
  7. Try it and Check yourself — exercises and questions.

Why build the tools before using them?#

You will not use a hand-written metrics library in production; you will use Prometheus client libraries and OpenTelemetry. But a counter is twenty lines of Go, a histogram is forty, and a trace context is one HTTP header. Building each one once removes the mystery, and the mystery is what makes people instrument badly: labels with unbounded values, averages of percentiles, alerts on causes instead of symptoms. Each lesson builds the mechanism first and then shows the production tool that does the same job.

flowchart LR
  APP["Your service<br/>counters, spans, logs"] --> COL["Collection<br/>scrape or push"]
  COL --> STORE[("Storage<br/>time series, columns, objects")]
  STORE --> Q["Query<br/>PromQL, SQL, trace search"]
  Q --> DASH["Dashboards"]
  Q --> ALERT["Alerts"]
  ALERT --> HUMAN["A person<br/>or an agent"]
  HUMAN -->|"asks a new question"| Q
  class APP compute
  class COL io
  class STORE memory
  class Q,ALERT queue
  class DASH,HUMAN neutral

The path#

flowchart LR
  I["I Why observability"] --> II["II Metrics"]
  II --> III["III Logs, traces, profiles"]
  III --> IV["IV Running it"]
  IV --> V["V AI infrastructure"]
  V --> VI["VI Tools and frontier"]
  class I neutral
  class II,III compute
  class IV queue
  class V memory
  class VI io
ModuleYou will be able toLevel
I — Why ObservabilitySay what to measure and why; define an SLO; read a percentile correctlyFoundations
II — MetricsInstrument a Go service, write PromQL, choose histogram buckets, keep cardinality boundedBasic
III — Logs, Traces and ProfilesEmit structured events, propagate a trace, use OpenTelemetry, read a profileBasic
IV — Running ItDesign the pipeline, build dashboards that answer questions, alert on SLO burn, control cost, work an incidentIntermediate
V — AI InfrastructureObserve GPUs, inference engines, LLM requests, a serving fleet, cost and power, and output qualityAdvanced
VI — Tools and FrontierMap the tool landscape, explain how the field is changing, and know what to watchExpert
ProjectsBuild five Go programs that make the ideas stickAll

What you need#

  • Go 1.22 or newer and a terminal.
  • Nothing else for modules I–III: every program runs offline.
  • Docker (or Podman) for the optional exercises that start Prometheus, Grafana or an OpenTelemetry Collector.
  • No GPU required. Module V explains GPU and engine telemetry with simulators and real metric names; the exercises that read a real GPU say so and give a no-GPU alternative.

Getting started in fifteen minutes#

  1. Read I.01 and run its program.
  2. Read II.01 and II.02; curl your own /metrics endpoint.
  3. If you came for AI infrastructure and already know Prometheus, jump to V.01 and come back to module II’s histograms and cardinality lessons when a percentile or a bill surprises you.

A promise about names and versions#

Metric names, tool versions and project statuses in this course were checked on 3 October 2026. Mechanisms — what a counter is, why a percentile cannot be averaged, why a GPU-utilization gauge misleads — do not change. Names do: a lesson says so wherever a name comes from a specification that is still marked unstable, and VI.03 lists where to check what has moved.

↑↓ navigate↵ openesc close