PidokuInfra

Project 15 — Mini Inference Platform

Expert 20h Difficulty 5/5 Topic 15 of 15

Prerequisites Projects 06, 08, 12; Section XI; Section XII

The capstone. Put a control plane around everything you have built: declare a model, and the platform deploys it, routes to it, scales it, rolls it, and bills for it.


1. What you build#

A small but complete multi-tenant inference platform on Kubernetes:

                         CONTROL PLANE
   ┌───────────────┐   ┌────────────────┐   ┌──────────────┐
   │ Model registry│ → │  Controller /  │ → │  Autoscaler  │
   │ (CRD or YAML) │   │  reconciler    │   │ (queue-based)│
   └───────────────┘   └────────────────┘   └──────────────┘
                              │ creates / updates
                              ▼
                          DATA PLANE
   client → Gateway (P12) → Engine pods (P08 or vLLM) → GPUs / CPUs
                │                    │
                └──── metrics ───────┴──→ Prometheus → Grafana, alerts, cost report

One kubectl apply -f model.yaml should take a model from a registry entry to a served, routed, autoscaled, observable endpoint.

Diagram — The platform#

flowchart TB
  DEV["kubectl apply ModelDeployment"] --> REG
  subgraph CP["Control plane"]
    REG["Model registry"] --> CTL["Controller / reconciler"]
    CTL --> AS["Autoscaler<br/>queue-based"]
    CTL --> RO["Rollout manager<br/>canary + gates"]
  end
  subgraph DP["Data plane"]
    GW["Gateway - Project 12"] --> ENG["Engine pods - Project 08 or vLLM"]
    ENG --> HW["GPUs / CPUs"]
  end
  U["Tenants"] --> GW
  CTL -.->|"creates, updates"| ENG
  RO -.->|"traffic weights"| GW
  AS -.->|"replicas"| ENG
  GW -.-> PROM["Prometheus"]
  ENG -.-> PROM
  PROM --> AS
  PROM --> GRAF["Grafana, alerts, cost report"]

  class REG,CTL,AS,RO queue
  class GW,ENG,HW compute
  class PROM,GRAF io
  class DEV,U neutral

2. Why it matters#

This is the job. Platform and infrastructure roles in this field are not asked to invent attention kernels; they are asked to make fleets of engines dependable and cheap. Everything in Sections XI and XII becomes one working system here — and one portfolio piece that speaks directly to that role.


3. Read first#


4. Spec#

YAML
apiVersion: lab.inference/v1
kind: ModelDeployment
metadata: { name: chat-small }
spec:
  model:    { uri: "hf://Qwen/Qwen2.5-0.5B-Instruct", revision: "<commit sha>", sha256: "..." }
  engine:   { image: "ghcr.io/you/p08-engine:1.4", maxNumSeqs: 32, maxLen: 4096 }
  resources:{ gpu: 0, cpu: "4", memory: "8Gi" }
  scaling:  { min: 1, max: 6, metric: queue_wait_p95_ms, target: 200,
              scaleDownStabilizationSeconds: 300 }
  rollout:  { strategy: canary, steps: [5, 25, 100], gate: { ttft_p95_ms: 800, error_rate: 0.01 } }
  tenants:  [ { name: team-a, tpm: 200000, priority: high },
              { name: team-b, tpm:  50000, priority: batch } ]

Components:

Registry      immutable, content-addressed model versions (URI + revision + checksum)
Controller    reconciles ModelDeployment → Deployment, Service, gateway route, ServiceMonitor
              (kopf / controller-runtime / a plain loop over the API — your choice)
Model cache   weights on a PVC or node-local cache; init container verifies the checksum
Probes        startup (model loading can take minutes), readiness (loaded AND not saturated),
              liveness (process only)
Gateway       Project 12, configured by the controller
Autoscaler    scales on queue wait / running sequences per replica — NOT on GPU utilization
Rollout       weighted canary through the gateway, automatic promote or rollback on the gate
Observability dashboards: TTFT, ITL, queue wait, running batch size, tokens/s, KV usage,
              errors, per-tenant usage ; SLO burn-rate alert
Cost          per-tenant $ = tokens-share-weighted GPU-seconds × price (XI.03)

5. Milestones#

  1. Cluster. kind or k3d. If you have a GPU node: NVIDIA GPU Operator, and request devices via the device plugin or DRA.
  2. Hand-written manifests first. Deploy one engine + gateway manually. Get probes right.
  3. Registry + controller. kubectl apply a ModelDeployment; the controller creates everything. Deleting it cleans everything up. Changing maxNumSeqs triggers a rolling update.
  4. Observability. Prometheus scraping engines and gateway; one Grafana dashboard per model; an SLO alert.
  5. Autoscaling. Drive load with your generator. Watch replicas follow queue wait. Measure cold-start time and decide your min from it.
  6. Canary rollout. Ship a new engine image at 5% → 25% → 100%. Then ship a deliberately slow one and watch the gate roll it back.
  7. Multi-tenancy. Quotas, priority, and a noisy-neighbour test.
  8. Cost report. Per-tenant cost for a one-hour replay. Idle capacity must be accounted for, not hidden.
  9. Game day. Kill a pod mid-stream, drain a node, corrupt a weight file, exhaust a quota. Write a short runbook entry for each.

6. Starter skeleton#

Python
import kopf, kubernetes as k8s

@kopf.on.create("lab.inference", "v1", "modeldeployments")
@kopf.on.update("lab.inference", "v1", "modeldeployments")
def reconcile(spec, name, namespace, patch, **_):
    desired = render_deployment(name, spec)          # pure function: spec → manifests
    kopf.adopt(desired)                              # owner refs → garbage collection on delete
    apply(desired)                                   # server-side apply; idempotent
    gateway.upsert_route(model=name, service=f"{name}.{namespace}.svc", tenants=spec["tenants"])
    patch.status["observedRevision"] = spec["model"]["revision"]

@kopf.timer("lab.inference", "v1", "modeldeployments", interval=15)
def autoscale(spec, name, namespace, status, **_):
    s = spec["scaling"]
    cur = current_replicas(name, namespace)
    val = prom(f'histogram_quantile(0.95, sum(rate(queue_wait_seconds_bucket{{model="{name}"}}[1m])) by (le))') * 1000
    want = min(s["max"], max(s["min"], math.ceil(cur * val / s["target"])))
    if want > cur:
        scale(name, namespace, want)                              # scale up immediately
    elif want < cur and stable_for(name, s["scaleDownStabilizationSeconds"]):
        scale(name, namespace, cur - 1)                           # scale down slowly, one at a time

7. What to measure#

MeasurementExpectation to write down first
Time from kubectl apply to first successful tokenDominated by image pull + weight load
Cold start breakdown: schedule / pull / load / warmupKnow each term
Scale-up reaction time vs load rampIs the SLO violated before capacity arrives?
TTFT p95 during a rolling updateNo visible blip if draining works
Canary: time to detect and roll back a bad versionMinutes, automatically
Noisy tenant at 10× quota: effect on others’ p95None
Utilization: served tokens / capacity tokensProbably far lower than you hoped
Cost per million tokens, per tenantWith idle capacity included

8. Done when#

  • One manifest deploys a model end to end; deleting it removes everything.
  • Replicas follow load, with a recorded cold-start number justifying min.
  • A bad canary is rolled back without human action.
  • Quotas and priorities hold under a noisy-neighbour test.
  • A dashboard answers “is it healthy, is it fast, who is using it, what does it cost.”
  • You wrote a two-page design doc: architecture, SLOs, capacity model, failure modes, what you would change for 100 GPUs. That document is the real deliverable.
  • You can answer Checkpoint F from ROADMAP.md with this system as your worked example.

9. Common pitfalls#

Autoscaling on GPU utilization. It reads ~100% from one request to saturation (X.04). Scale on queue wait or running sequences.

Liveness probe that depends on the model. Slow load or a busy engine → restart loop.

Readiness that only checks “loaded”. A saturated pod keeps receiving traffic.

Scaling down as fast as up. With multi-minute cold starts, flapping is very expensive.

No graceful drain. SIGTERM kills in-flight streams on every rollout. Use a preStop hook and a long enough termination grace period.

Mutable model references (main, latest). Two replicas end up serving different weights. Pin the revision and verify the checksum.

Downloading weights on every pod start. Cache them on the node or a shared volume.

Control plane on the request path. If the controller is down, existing traffic must keep flowing.


10. Stretch goals#

  • Swap your engine for vLLM and your gateway’s routing for the Gateway API Inference Extension or llm-d; compare what they give you against what you built.
  • Scale to zero with request buffering at the gateway; measure the first-request penalty.
  • Multi-LoRA: one base model, per-tenant adapters loaded on demand.
  • Bin-packing: several small models on one GPU (MPS / time-slicing / MIG), with interference measured (XII.04).
  • A second “region” (second cluster) with failover and a DR drill (XI.06).
  • GitOps: ModelDeployment manifests in a repo, reconciled by Argo CD or Flux.

11. Interview questions this project answers#

  1. Design a multi-tenant LLM serving platform. What is in the control plane, what is in the data plane, and why does the split matter?
  2. Which metric do you autoscale on, and why not GPU utilization?
  3. How do you roll out a new model version safely?
  4. How do you attribute cost per tenant on shared GPUs?
  5. How do you keep cold starts from violating your SLO?
  6. What happens to in-flight streams during a deploy?

12. After this#

You have built the stack from a NumPy matmul to a platform. See “After XIV” in ROADMAP.md: contribute upstream, reproduce a paper, or halve a real system’s cost per million tokens and write it up.

↑↓ navigate↵ openesc close