PidokuInfra
05AI Infrastructure

AI Infrastructure

AdvancedModule 057 topics~6h 30m

Topics, in order

01 What Is DifferentA web request costs microseconds of CPU and is either right or an error. An LLM request costs seconds on a device that rents for dollars an hour, its cost depends on how many tokens go in … Advanced 45 min 02 GPU TelemetryA GPU reports four kinds of thing: how busy it is, how much memory is in use, its power and temperature, and its errors. The first two are the ones people look at, and both mislead for LLM … Advanced 1h 03 Inference Engine MetricsAn LLM server is a queue in front of a batch loop. A request waits, is prefilled (the prompt is read, producing the first token), then decodes one token per step alongside everyone else in … Advanced 1h 10m 04 Tracing LLM RequestsA trace of an AI request has the usual HTTP and database spans plus three new kinds: a model call (which model, how many tokens in and out, why it stopped), a tool call (the model asked for … Advanced 55 min 05 Platform and FleetOne engine's metrics tell you about one replica. A platform has many replicas, several models, a router deciding who gets which request, an autoscaler deciding how many replicas exist, and … Advanced 1h 06 Cost, Power and EfficiencyFor AI infrastructure, cost is not an accounting afterthought; it is an engineering metric to graph next to latency. Two figures summarize a serving system: dollars per million tokens and … Advanced 50 min 07 Quality and EvalsAn inference platform can be fast, cheap and available while producing worse answers than last week. Nothing in the first five layers detects that. Quality has to be measured as its own … Advanced 50 min

About this module

Everything so far applies to any service. This module is about what changes when the service is a model on a GPU: new resources to watch, new units of work, new failure modes — and a product whose output can be wrong while every infrastructure metric is green.

#LessonThe question it answers
01What Is DifferentWhich assumptions from web-service observability break?
02GPU TelemetryWhat does a GPU report, and which numbers lie?
03Inference Engine MetricsWhat should an LLM server expose, and how do I read it?
04Tracing LLM RequestsHow do I trace a model call, a tool call, an agent?
05Platform and FleetWhat do I watch across a cluster, and what do I scale on?
06Cost, Power and EfficiencyWhat does a token cost in dollars and joules?
07Quality and EvalsHow do I observe whether the answers are any good?

When you finish you can list the metrics that matter at each layer of an inference platform, explain why GPU utilization is not one of them, and compute cost per million tokens from counters.

This module refers often to Inference Engineering (prefill, decode, KV cache, continuous batching) and GPU Engineering (memory bandwidth, power). Each lesson links the specific background it needs.

↑↓ navigate↵ openesc close