PidokuInfra

Monitoring vs Observability

Foundations Beginner 30 min Difficulty 1/5 Topic 01 of 04

The idea in one minute#

Monitoring answers questions you thought of in advance: is the error rate above 1%? Observability is the property of a system that lets you answer questions you did not think of in advance: why are requests from one customer, on one model, slow only when the prompt is long? — without shipping new code to find out.

Monitoring tells you that something is wrong. Observability lets you find out why. You need both, and the second is built from richer data, not from more dashboards.

An analogy#

A car dashboard is monitoring: a few fixed gauges and a warning light. A mechanic’s diagnostic port is observability: it lets them ask the engine arbitrary questions about what it was doing when the light came on.

The light is useful. It is not enough to fix the car.

A picture#

flowchart TB
  SYS["Running system"] -->|"emits telemetry"| TEL[("Telemetry<br/>metrics, logs, traces, profiles")]
  TEL --> KNOWN["Known questions<br/>dashboards and alerts"]
  TEL --> UNKNOWN["New questions<br/>slice by any field"]
  KNOWN -->|"something is wrong"| P["A person"]
  P -->|"why?"| UNKNOWN
  UNKNOWN -->|"which requests, which hosts, which change"| FIX["Cause found"]
  class SYS compute
  class TEL memory
  class KNOWN,UNKNOWN queue
  class P,FIX neutral

How it really works#

Known unknowns and unknown unknowns#

MonitoringObservability
QuestionDecided before the incidentInvented during the incident
DataA few aggregated numbersDetailed events with many fields
Typical toolDashboard, alert ruleAd-hoc query, trace search, profile diff
Fails whenThe failure is newThe data needed was never recorded

A system is observable to the degree that its outputs let you reconstruct its internal state. The term comes from control theory; in software it reduces to one practical test:

Can I explain why this particular request was slow, using only what the system already emitted?

What makes a system observable#

  1. It emits events with context. Not request failed, but which tenant, which model, which replica, which version, how many tokens, how long each phase took.
  2. The context is consistent across components. The same request ID or trace ID appears in the gateway, the server and the log line, so you can follow one request through all of them.
  3. You can slice by any of those fields afterwards. The fields you need are rarely the ones you predicted.

What observability is not#

  • Not three pillars you buy. Metrics, logs and traces are data types (lesson 02). Having all three in separate tools that cannot be joined is still poor observability.
  • Not more dashboards. Fifty dashboards answer fifty predicted questions.
  • Not free. Every event costs CPU, network and storage. Module IV is largely about deciding what to keep.

The loop you are building#

detect   → an alert on a user-facing symptom fires
triage   → how bad, who is affected, since when
localize → which component, version, tenant, host
explain  → what that component was doing (trace, profile, logs)
fix      → roll back, scale, patch
learn    → add the missing signal so next time is faster

Each module of this course strengthens one step. Detection is modules I, II and IV; localizing and explaining is module III; module V repeats the whole loop for GPUs and LLM serving.

Code#

A service that only counts errors can tell you that 3% of requests fail. The same service recording one event per request can tell you which 3%.

Go
// why.go — the same failures seen through a counter and through events.
package main

import (
	"fmt"
	"math/rand"
	"sort"
)

type Event struct {
	Tenant, Model, Replica string
	PromptTokens           int
	Failed                 bool
}

func main() {
	rng := rand.New(rand.NewSource(1))
	tenants := []string{"acme", "globex", "initech"}
	models := []string{"small", "large"}
	replicas := []string{"r0", "r1", "r2", "r3"}

	var events []Event
	for i := 0; i < 20000; i++ {
		e := Event{
			Tenant:       tenants[rng.Intn(len(tenants))],
			Model:        models[rng.Intn(len(models))],
			Replica:      replicas[rng.Intn(len(replicas))],
			PromptTokens: 100 + rng.Intn(8000),
		}
		// The hidden bug: replica r2 runs out of memory on long prompts for the large model.
		if e.Replica == "r2" && e.Model == "large" && e.PromptTokens > 6000 {
			e.Failed = rng.Float64() < 0.9
		}
		events = append(events, e)
	}

	// Monitoring view: one number.
	failed := 0
	for _, e := range events {
		if e.Failed {
			failed++
		}
	}
	fmt.Printf("monitoring: error rate = %.2f%%  (something is wrong; no idea what)\n\n",
		100*float64(failed)/float64(len(events)))

	// Observability view: group failures by each field and see where they concentrate.
	for _, dim := range []struct {
		name string
		key  func(Event) string
	}{
		{"tenant", func(e Event) string { return e.Tenant }},
		{"model", func(e Event) string { return e.Model }},
		{"replica", func(e Event) string { return e.Replica }},
		{"long prompt", func(e Event) string { return fmt.Sprint(e.PromptTokens > 6000) }},
	} {
		total, bad := map[string]int{}, map[string]int{}
		for _, e := range events {
			k := dim.key(e)
			total[k]++
			if e.Failed {
				bad[k]++
			}
		}
		keys := make([]string, 0, len(total))
		for k := range total {
			keys = append(keys, k)
		}
		sort.Strings(keys)
		fmt.Printf("by %-12s", dim.name)
		for _, k := range keys {
			fmt.Printf("  %s=%.1f%%", k, 100*float64(bad[k])/float64(total[k]))
		}
		fmt.Println()
	}
}

The counter says roughly 3%. The grouped view shows the failures are evenly spread across tenants, and concentrated on one replica, one model and long prompts. That is the difference this course is about.

Remember this#

  • Monitoring answers predicted questions; observability lets you answer new ones.
  • Observability comes from events with rich, consistent context — not from more dashboards.
  • The practical test: can you explain one specific slow request from what was already emitted?
  • Telemetry has a cost; deciding what to keep is part of the engineering.

Try it#

  1. Run why.go. Change the hidden bug to depend on Tenant instead of Replica. Does the grouped view still find it?
  2. Add a field the bug depends on but that you do not record (say, a kernel version). What does the grouped view show now? What does that teach you about choosing fields?
  3. For a service you know, write down one question you could not answer during its last incident. Which field was missing?

Check yourself#

  1. State the difference between monitoring and observability in one sentence.
  2. Why can a system with metrics, logs and traces still be hard to debug?
  3. What are the six steps of the incident loop?

↑↓ navigate↵ openesc close