PidokuInfra

The Tool Landscape

Expert Advanced 50 min Difficulty 2/5 Topic 01 of 04

Prerequisites IV.01, V

The idea in one minute#

There are hundreds of observability products and a small number of jobs: instrument, collect, store, query, alert — and, for AI, read the accelerator, read the engine, trace the model call, score the output. Every tool does one or more of these jobs. Place a tool on that map and you know what it replaces and what it still needs beside it.

Choose in this order: open standards for instrumentation first (so nothing else is permanent), then storage by the shape of your questions, then the interface people will actually use at 3 a.m.

An analogy#

A kitchen. There are countless brands, and only so many jobs: cut, heat, cool, store, serve. You do not need to know every brand. You need to know which job each appliance does, and not to buy three ovens and no fridge.

A picture#

flowchart TB
  subgraph INS["Instrument"]
    I1["OpenTelemetry SDKs"]
    I2["Prometheus client libraries"]
    I3["eBPF: OBI, Beyla, Pixie"]
    I4["Exporters: node, kube-state, DCGM"]
  end
  subgraph COLL["Collect and process"]
    C1["OTel Collector, Grafana Alloy"]
    C2["Prometheus scrape, vmagent"]
    C3["Fluent Bit, Vector"]
  end
  subgraph STO["Store and query"]
    S1["Metrics: Prometheus, Mimir, Thanos, VictoriaMetrics"]
    S2["Logs: Loki, OpenSearch, ClickHouse"]
    S3["Traces: Tempo, Jaeger"]
    S4["Profiles: Pyroscope, Parca"]
  end
  subgraph USE["Use"]
    U1["Grafana, vendor UIs"]
    U2["Alertmanager, on-call tools"]
    U3["SLO tools: Sloth, Pyrra"]
  end
  subgraph AI["AI-specific"]
    A1["GPU: DCGM, vendor exporters"]
    A2["Engine metrics: vLLM, SGLang, Dynamo"]
    A3["LLM tracing and evals: Langfuse, Phoenix, LangSmith, ..."]
    A4["Load and benchmark: GuideLLM, AIPerf, inference-perf"]
  end
  INS --> COLL --> STO --> USE
  AI --> COLL
  class I1,I2,I3,I4 compute
  class C1,C2,C3 io
  class S1,S2,S3,S4 memory
  class U1,U2,U3 queue
  class A1,A2,A3,A4 neutral

How it really works#

General-purpose stack, by job#

JobOpen sourceNotes
Instrument codeOpenTelemetry SDKs; Prometheus client librariesOTel for traces and logs; either for metrics
Instrument without codeOBI (OpenTelemetry eBPF Instrumentation), Grafana Beyla, Pixie, CorootBreadth, protocol-level detail
Export system statenode_exporter, cAdvisor, kube-state-metrics, blackbox exporterThe standard Kubernetes set
Collect and processOpenTelemetry Collector; Grafana Alloy (a Collector distribution); Fluent Bit; VectorThe Collector is the default hub
Metrics storagePrometheus; Grafana Mimir, Thanos, Cortex; VictoriaMetricsSingle node → horizontally scaled with object storage
Log storageGrafana Loki; OpenSearch / Elasticsearch; ClickHouseLabel-indexed vs full-text vs columnar
Trace storageGrafana Tempo; Jaeger (v2 is built on the OTel Collector); ClickHouse
Profile storageGrafana Pyroscope; Parca
All signals in oneSigNoz, OpenObserve, Uptrace, ClickStack (ClickHouse + HyperDX)OTel-native, columnar back ends
DashboardsGrafana; Perses
Alert routingAlertmanager; Grafana AlertingPlus an on-call scheduler
SLOsSloth, Pyrra; OpenSLO as the specGenerate burn-rate rules
Kubernetes packagingkube-prometheus-stack, the Prometheus Operator, the OpenTelemetry OperatorStart here rather than assembling by hand
EnergyKeplerPer-pod energy estimates

Commercial platforms covering most of the table in one product include Datadog, Dynatrace, New Relic, Splunk, Honeycomb, Chronosphere and Grafana Cloud, and each major cloud has its own (CloudWatch, Azure Monitor, Google Cloud Observability, including managed Prometheus services).

AI infrastructure, by layer#

Layer (V.01)ToolsWhat you get
AcceleratorNVIDIA DCGM + dcgm-exporter (installed by the GPU Operator); go-nvml; nvidia-smi, nvtop, nvitop for ad-hoc use; AMD’s device metrics exporter; cloud-provider accelerator metricsPower, energy, temperature, memory, errors, links
Node healthNode Problem Detector; NVIDIA NVSentinel; DRA device taintsDetect, cordon, drain, remediate
Device profilingNVIDIA Nsight Systems and Nsight Compute; PyTorch profiler; eBPF-based GPU profilersKernel-level time on the device
EngineBuilt-in Prometheus endpoints of vLLM, SGLang, NVIDIA Dynamo and Triton, TensorRT-LLM, llama.cpp server, KServe, Ray ServeQueue, TTFT, tokens, KV cache
Router / gatewayGateway API Inference Extension endpoint picker and llm-d; Envoy AI Gateway; LiteLLM; Kong and other API gateways with AI pluginsPer-tenant tokens, routing decisions
PlatformKubernetes metrics; KEDA; Kueue and other batch schedulers’ metricsScaling, pending, quotas
LLM application tracingOpenTelemetry GenAI instrumentations; OpenLLMetry (Traceloop); OpenInference (Arize)Standard spans for model, tool and agent calls
LLM observability and evalsLangfuse; Arize Phoenix / Arize AX; LangSmith; Braintrust; Opik (Comet); W&B Weave; MLflow Tracing; Helicone; LLM modules of Datadog, Grafana, New Relic, Dynatrace, HoneycombTrace views built for prompts, token cost, datasets, evaluators
Load testing and benchmarksGuideLLM; NVIDIA AIPerf (successor to GenAI-Perf); inference-perf (Kubernetes SIG); llm-d-benchmark; public references InferenceX and MLPerf InferenceCapacity and goodput numbers to compare your metrics against

LLM-observability tools: how they differ#

They all show a trace of model calls with tokens and cost. The differences that matter:

QuestionWhy it matters
Does it ingest plain OTLP with the GenAI conventions?Otherwise you are locked to its SDK
Open source and self-hostable? Under which licence?Prompts are sensitive; many teams must keep them in-house. Langfuse’s core is MIT; Phoenix uses Elastic License 2.0; others are hosted-only
Is evaluation built in (datasets, judges, human review queues)?Tracing without scoring is half the job (V.07)
Does it join to infrastructure telemetry?Otherwise an LLM trace and a GPU metric live in different worlds
Is it tied to one framework?LangSmith is deepest with LangChain/LangGraph; that is a strength only if you use them
What happens at volume?Many were built for development-time tracing, not millions of requests per hour

The category is consolidating into the general platforms: ClickHouse acquired Langfuse in January 2026 (the open-source licence and self-hosting were kept), and every large observability vendor now ships an LLM module. Expect “LLM observability” to become a feature of your observability platform rather than a separate product — which is one more reason to emit standard OTel.

How to choose#

  1. Instrument with OpenTelemetry and Prometheus exposition. This decision outlasts all the others.
  2. Start from what your platform installs. On Kubernetes with GPUs: kube-prometheus-stack, the GPU Operator’s dcgm-exporter, your engine’s /metrics. That is a working system in an afternoon.
  3. Pick storage by question shape. Known questions and alerts → a Prometheus-compatible TSDB. Unknown questions sliced by many fields → a column store with wide events.
  4. Decide where prompts may live before choosing an LLM tracing tool.
  5. Count operating cost honestly. Self-hosting is free until it pages you.
  6. Run a real incident drill on the candidate before committing. The test is how fast a tired engineer gets from the page to the cause.

A sensible default, small to large#

ScaleStack
One team, one clusterPrometheus + Grafana + Alertmanager; Loki; dcgm-exporter; engine metrics; an OTel Collector; one LLM tracing tool if you build applications
Several clustersAdd remote write to Mimir / Thanos / VictoriaMetrics or a managed Prometheus; Tempo with tail sampling; continuous profiling
Large fleetCollector gateways with per-tenant limits; a column store for wide events; SLO tooling; automated node remediation; showback

Code#

A tool is a set of jobs. This program checks a proposed stack for gaps and overlaps — the question to ask of any architecture diagram.

Go
// stackcheck.go — does this stack cover every job, and where does it overlap?
package main

import (
	"fmt"
	"sort"
	"strings"
)

func main() {
	jobs := []string{
		"instrument", "collect", "metrics store", "log store", "trace store", "profiles",
		"dashboards", "alert routing", "gpu telemetry", "engine metrics", "llm tracing", "evals", "load test",
	}
	stack := map[string][]string{
		"OpenTelemetry SDK + Collector": {"instrument", "collect"},
		"Prometheus":                    {"collect", "metrics store"},
		"Grafana":                       {"dashboards"},
		"Alertmanager":                  {"alert routing"},
		"Loki":                          {"log store"},
		"Tempo":                         {"trace store"},
		"dcgm-exporter":                 {"gpu telemetry"},
		"vLLM /metrics":                 {"engine metrics"},
		"Langfuse":                      {"llm tracing", "evals", "trace store"},
	}

	covered := map[string][]string{}
	for tool, js := range stack {
		for _, j := range js {
			covered[j] = append(covered[j], tool)
		}
	}
	fmt.Println("job              covered by")
	var gaps []string
	for _, j := range jobs {
		tools := covered[j]
		sort.Strings(tools)
		note := ""
		switch {
		case len(tools) == 0:
			note = "← GAP"
			gaps = append(gaps, j)
		case len(tools) > 1:
			note = "← overlap: decide which is the source of truth"
		}
		fmt.Printf("%-15s  %-48s %s\n", j, strings.Join(tools, ", "), note)
	}
	fmt.Printf("\n%d gaps: %s\n", len(gaps), strings.Join(gaps, ", "))
}

Remember this#

  • Tools are implementations of a few jobs. Map a tool to its jobs before comparing it.
  • Open instrumentation first; it keeps every later choice reversible.
  • For AI: accelerator exporter + engine metrics + gateway accounting + standard GenAI traces + an evaluator. LLM observability is becoming a feature of general platforms.
  • Decide where prompts are allowed to be stored before you pick a tracing tool.

Try it#

  1. Run stackcheck.go with the stack you actually use. What are the gaps?
  2. For one LLM observability tool, answer the six questions in the table from its documentation.
  3. Design the smallest stack that covers every job for a single-node inference server.

Check yourself#

  1. Which decision in a stack is the hardest to reverse, and how do you make it reversible?
  2. What distinguishes a column store from a time-series database as a home for telemetry?
  3. Name one tool per AI layer.

↑↓ navigate↵ openesc close