PidokuInfra

Model Registry

Expert Intermediate 1h Difficulty 2/5 Topic 02 of 11

Prerequisites VIII.10, XI.07


1. What it is#

The authoritative record of what models exist, what versions of each, where their artifacts live, and how to serve them.

Not a model zoo or an experiment tracker. A deployment registry: the thing the platform consults to answer “what should I run and how?”


2. What a registry entry must contain#

YAML
model_id: llama-3-70b-instruct
versions:
  - version: v1.5.0
    status: production          # draft | staged | canary | production | deprecated

    # --- IDENTITY ---
    artifact:
      uri: s3://models/llama-3-70b-instruct/v1.5.0/
      format: safetensors
      sha256_manifest: sha256:a3f2...
      size_bytes: 70600000000
    base_model: meta-llama/Meta-Llama-3-70B-Instruct
    tokenizer_uri: s3://models/.../tokenizer/
    chat_template_sha256: sha256:9b1c...

    # --- SERVING CONFIGURATION (part of the identity!) ---
    quantization:
      method: fp8
      scheme: per-tensor
      kv_cache_dtype: fp8
      calibration_dataset: s3://calib/prod-sample-2026-02/
      calibration_sha256: sha256:7e4d...
      excluded_layers: [lm_head, embed_tokens]
    engine:
      type: vllm
      version: "0.x.y"
      image: registry/vllm:0.x.y-cuda12.4
    parallelism:
      tensor_parallel: 8
      pipeline_parallel: 1
    engine_args:
      max_model_len: 8192
      max_num_seqs: 160
      max_num_batched_tokens: 2048
      gpu_memory_utilization: 0.90
      enable_prefix_caching: true
      enable_chunked_prefill: true

    # --- REQUIREMENTS ---
    resources:
      gpus_per_replica: 8
      gpu_type: [H100-80GB, H200-141GB]
      requires_nvlink_domain: true
      min_cpu_cores: 64
      min_host_memory_gb: 512
      min_shm_gb: 16

    # --- OPERATIONAL ---
    capabilities: [chat, streaming, structured_output, multi_lora]
    max_context: 8192
    default_sampling:
      temperature: 0.7
      top_p: 0.9

    # --- PROVENANCE AND VALIDATION ---
    created_at: "2026-03-01T10:00:00Z"
    created_by: "platform-team"
    evaluation:
      suite: prod-eval-v3
      results_uri: s3://evals/llama-3-70b-v1.5.0.json
      summary: {mmlu: 0.792, gsm8k: 0.881, our_task: 0.847}
      reference_version: v1.4.0
      delta_vs_reference: {mmlu: -0.003, our_task: -0.002}
      approved_by: "ml-quality-team"
      approved_at: "2026-03-02T14:30:00Z"
    rollout:
      strategy: progressive
      canary_percent: 1
      guards_ref: standard-llm-guards-v2

The section labelled “serving configuration is part of the identity” is the one that distinguishes a real registry from a file listing. Two deployments of the same checkpoint with different quantization are different models from the user’s perspective (Section VIII.10).


3. Why the serving config belongs in the registry#

QUESTION: "Which version of Llama-3-70B is in production?"

WRONG ANSWER: "v1.5.0"
RIGHT ANSWER: "v1.5.0, FP8 per-tensor with FP8 KV, vLLM 0.x.y, TP=8,
               chat template hash 9b1c..., evaluated at MMLU 0.792"

Because if any of those change, the outputs change.

Practical consequence: when someone reports “the model got worse,” you look up the version record and immediately see everything that could have caused it. Without the record, you’re guessing.


4. The API#

# Discovery
GET  /v1/models                            list all
GET  /v1/models/{id}                       model details
GET  /v1/models/{id}/versions              all versions
GET  /v1/models/{id}/versions/{v}          full record
GET  /v1/models/{id}/production            what's serving now

# Lifecycle
POST /v1/models/{id}/versions              register a new version (draft)
POST /v1/models/{id}/versions/{v}/evaluate attach evaluation results
POST /v1/models/{id}/versions/{v}/promote  draft → staged → canary → production
POST /v1/models/{id}/versions/{v}/rollback revert to the previous production
POST /v1/models/{id}/versions/{v}/deprecate

# Platform integration
GET  /v1/deployments                       what should be running where
GET  /v1/models/{id}/resources             what a replica needs (for the scheduler)

The promotion path with gates is what makes it a registry rather than a database:

draft ──[evaluation passes]──▶ staged ──[shadow ok]──▶ canary
      ──[guards hold 2h]──▶ production ──[new version]──▶ deprecated

Each transition has a required condition, checked by the registry, not by whoever runs the deployment.


5. Artifact management#

STORAGE LAYOUT
  s3://models/{model_id}/{version}/
    ├── manifest.json           checksums for every file
    ├── config.json
    ├── tokenizer/
    ├── model-00001-of-00030.safetensors
    ├── ...
    └── engine_artifacts/       ← per-GPU-architecture builds
        ├── h100-tp8/
        │   ├── engine.plan     (TensorRT) or
        │   └── compile_cache/  (torch.compile)
        └── h200-tp8/

REGIONAL REPLICATION
  push to each region's bucket
  → never pull cross-region per node (egress cost, Section XI.06)

NODE-LOCAL CACHE
  /var/cache/models/{model_id}/{version}/
  populated by a DaemonSet or init container
  LRU eviction with a size limit
  → cold start goes from minutes to seconds (Section II.07)

The engine_artifacts directory is often forgotten and is what saves you 60-600 seconds of compile time per cold start.


6. Integration with the platform#

REGISTRY → SCHEDULER
  "model X needs 8 GPUs of type [H100, H200] in one NVLink domain"
  → the scheduler uses this for placement (Section XII.04)

REGISTRY → DEPLOYMENT CONTROLLER
  "the production version of model X is v1.5.0 with these engine args"
  → reconciles the actual state to match

REGISTRY → ROUTER
  "model X is served by these replicas, with these capabilities"
  → the routing table (Section XII.01)

REGISTRY → GATEWAY
  "model alias 'chat-large' maps to model X"
  → user-facing names decoupled from internal versions

REGISTRY → BILLING
  "model X costs W per input token, Y per output token"

The alias layer is worth calling out: users request chat-large, and the registry maps that to a specific model and version. This lets you change the underlying model without changing every client.


7. What to build vs adopt#

For a small platform (< 10 models):
  a Git repo of YAML + an object store + a CI pipeline
  → genuinely sufficient. Don't over-build.

For a medium platform (10-100 models, several teams):
  a real service with the API from section 4
  backed by Postgres + object storage
  → this is where the investment pays

For a large platform:
  the above, plus: multi-region replication, approval workflows,
  automated evaluation triggering, cost attribution integration

DON'T:
  use an experiment tracker (MLflow, W&B) as your deployment registry.
  They're built for a different purpose and lack the serving-config
  and lifecycle semantics you need. (You can use them alongside.)

8. Production implications#

  • The registry is the source of truth. If something is serving that isn’t in the registry, that’s an incident.
  • Serving configuration is part of the version identity. Record all of it.
  • Gate promotions on evaluation results. Make the registry enforce it, not a human’s memory.
  • Store engine artifacts per GPU architecture. It’s the difference between 30-second and 10-minute cold starts.
  • Version and checksum everything. Including the calibration dataset for quantization.
  • Keep deprecated versions deployable for 30+ days for rollback.
  • Use aliases so client-facing names are stable.
  • Emit the version in every metric and response header (Section XI.07).

9. Common mistakes#

Recording only the checkpoint. The serving config determines the output.

Using an experiment tracker as a deployment registry.

No checksums. A corrupted shard produces plausible garbage.

No engine artifact storage. Recompiling on every cold start.

No aliases. Every client change requires coordination.

Deleting old versions immediately. No rollback path.

Promotion without gates. The evaluation gets skipped under pressure.

No record of the calibration dataset for a quantized model. You can’t reproduce it.


10. Hands-on exercise#

A. Design the schema. Write the full registry entry (section 2 format) for a model you serve. Which fields did you not have recorded anywhere?

B. Build a minimal registry. Implement the API from section 4 backed by SQLite and a filesystem. Include the promotion gates.

C. Verify reproducibility. From your registry entry alone, can someone reproduce the exact deployment? Try it: give the entry to a colleague and have them deploy from it.

D. Artifact pipeline. Build the pipeline: checkpoint → quantize → build engine artifacts per GPU arch → checksum → push to regional storage → register. Time it end to end.

E. Node cache. Implement a DaemonSet (or a script) that populates a node-local model cache from the registry, with LRU eviction. Measure the cold-start improvement.


11. Interview questions#

  1. What must a model registry entry contain, beyond the checkpoint URI?
  2. Why is the serving configuration part of the model’s identity?
  3. What’s the promotion path, and what gates each transition?
  4. Why store engine artifacts, and how are they keyed?
  5. Why use aliases rather than exposing version IDs to clients?
  6. Why is an experiment tracker not a deployment registry?
  7. What do you keep for rollback and for how long?

12. Further reading#

  • [REFERENCE] MLflow Model Registry, KServe InferenceService — for comparison
  • [REFERENCE] OCI artifact specifications (some platforms store models as OCI artifacts)
  • [ESTABLISHED] Sculley et al., “Hidden Technical Debt in Machine Learning Systems” (2015)
  • Next: 03 — Kubernetes for GPUs

↑↓ navigate↵ openesc close