PidokuInfra

Projects

Reading builds recognition. Building builds understanding. These five projects are all in Go, and none requires a GPU.

#ProjectAfter moduleNeedsTime
1A metrics library and scraperIINothing4–6 h
2A tracer with propagation and tail samplingIIINothing5–8 h
3An SLO burn-rate alerterIVNothing3–5 h
4An LLM load tester that measures goodputVAny OpenAI-compatible endpoint, or the included fake5–8 h
5An inference observability stackVIDocker; a GPU is optional8–12 h

For each project, “done” means every box under Done when is ticked and you can explain the result to someone else.


1 — A metrics library and scraper#

Goal. Make counters, gauges, histograms and the scrape model concrete by building both sides.

Build. A Go package with Counter, Gauge and Histogram types supporting labels, and an HTTP handler that serves them in Prometheus text format. Then a second program that scrapes several such endpoints on an interval, stores samples in memory with a retention limit, records up for every scrape, and answers three queries: rate, sum by, and a histogram quantile.

Done when

  • A real Prometheus can scrape your endpoint without errors.
  • rate handles a counter reset correctly when you restart the target.
  • Your quantile matches Prometheus’s histogram_quantile on the same data.
  • You can print the number of series and show how one extra label changes it.

Stretch. Add an exponential (native-style) histogram and compare its error with fixed buckets on the same observations.


2 — A tracer with propagation and tail sampling#

Goal. Understand spans, context propagation and sampling by implementing them.

Build. A tracer with Start(ctx, name) returning a span, W3C traceparent inject and extract for net/http, and an exporter that sends finished spans as OTLP/JSON to a collector you also write. The collector buffers spans by trace ID, waits for the trace to finish, and keeps every trace with an error or over a latency threshold plus a configurable percentage of the rest.

Done when

  • A request through three of your services produces one trace with correct parentage.
  • Context survives a goroutine hand-off and a worker pool.
  • A real OpenTelemetry Collector accepts your exported payload.
  • Tail sampling keeps 100% of error traces and the configured share of the rest, and you can report how much memory the buffer used.

Stretch. Derive RED metrics from spans before sampling, and attach exemplars.


3 — An SLO burn-rate alerter#

Goal. Turn an SLO into pages that fire for real problems and stay quiet for blips.

Build. A program that reads an SLO definition (target, window, good and total queries), evaluates it against a stream of counters — your project 1 store or a Prometheus API — and implements the multi-window, multi-burn-rate rules. It prints, or posts to a webhook, a notification with the remaining budget and the burn rate.

Done when

  • A brief blip (5% errors for five minutes) does not page; a 30-minute partial outage does, within minutes.
  • A slow leak produces a ticket-severity notification, not a page.
  • The alert resolves within minutes of the problem ending.
  • It can also emit the equivalent Prometheus recording and alerting rules as YAML.

Stretch. Support a latency SLI from histogram buckets, and a low-traffic mode.


4 — An LLM load tester that measures goodput#

Goal. Measure an LLM server the way its users experience it.

Build. An open-loop load generator for an OpenAI-compatible streaming endpoint: requests are sent on a schedule (Poisson arrivals at a target rate) regardless of responses. For every request record the timestamp of each streamed chunk and compute TTFT, TPOT, the ITL distribution, output tokens and finish reason. Sweep the arrival rate and report throughput and goodput against TTFT and TPOT thresholds. Include a small fake server (fixed prefill time, per-token delay growing with concurrency) so the tool can be developed without a model.

Done when

  • Latency is measured from the intended send time (no coordinated omission).
  • The report shows, per rate: TTFT p50/p95/p99, TPOT p95, ITL p99, tokens/s, goodput.
  • You can state the server’s capacity as “the highest rate with goodput ≥ 99%”.
  • Against a real engine, your client-side TTFT is consistent with its server-side histogram, and you can explain the difference.

Stretch. Replay a prompt set with shared prefixes and report how prefix-cache hit rate changes capacity.


5 — An inference observability stack#

Goal. Assemble the whole course into one running system.

Build. A docker compose (or Kubernetes) setup with: an inference server — vLLM or llama.cpp’s server if you have hardware, otherwise your fake from project 4 exposing the same metric families; a small Go gateway that authenticates fake tenants, counts tokens, emits one wide event per request and propagates trace context; an OpenTelemetry Collector; Prometheus; Grafana; a trace store. If you have an NVIDIA GPU, add dcgm-exporter; if not, write a fake exporter that serves plausible GPU series.

Done when

  • One dashboard shows, top to bottom: SLO status, TTFT and TPOT, tokens/s, waiting requests and KV usage, GPU power and temperature.
  • A burn-rate alert fires when you overload the server with project 4.
  • From a latency spike you can click through to a trace, and from the trace find the tenant and the replica.
  • A cost panel shows dollars per million tokens and tokens per joule, and both move the right way when you change the load.
  • Prompt content is absent from every store unless you turn a documented flag on.

Stretch. Add a canary replica with different settings and a panel comparing finish reasons and output lengths between the two.

↑↓ navigate↵ openesc close