1. What it is#
The authoritative record of what models exist, what versions of each, where their artifacts live, and how to serve them.
Not a model zoo or an experiment tracker. A deployment registry: the thing the platform consults to answer “what should I run and how?”
2. What a registry entry must contain#
model_id: llama-3-70b-instruct
versions:
- version: v1.5.0
status: production # draft | staged | canary | production | deprecated
# --- IDENTITY ---
artifact:
uri: s3://models/llama-3-70b-instruct/v1.5.0/
format: safetensors
sha256_manifest: sha256:a3f2...
size_bytes: 70600000000
base_model: meta-llama/Meta-Llama-3-70B-Instruct
tokenizer_uri: s3://models/.../tokenizer/
chat_template_sha256: sha256:9b1c...
# --- SERVING CONFIGURATION (part of the identity!) ---
quantization:
method: fp8
scheme: per-tensor
kv_cache_dtype: fp8
calibration_dataset: s3://calib/prod-sample-2026-02/
calibration_sha256: sha256:7e4d...
excluded_layers: [lm_head, embed_tokens]
engine:
type: vllm
version: "0.x.y"
image: registry/vllm:0.x.y-cuda12.4
parallelism:
tensor_parallel: 8
pipeline_parallel: 1
engine_args:
max_model_len: 8192
max_num_seqs: 160
max_num_batched_tokens: 2048
gpu_memory_utilization: 0.90
enable_prefix_caching: true
enable_chunked_prefill: true
# --- REQUIREMENTS ---
resources:
gpus_per_replica: 8
gpu_type: [H100-80GB, H200-141GB]
requires_nvlink_domain: true
min_cpu_cores: 64
min_host_memory_gb: 512
min_shm_gb: 16
# --- OPERATIONAL ---
capabilities: [chat, streaming, structured_output, multi_lora]
max_context: 8192
default_sampling:
temperature: 0.7
top_p: 0.9
# --- PROVENANCE AND VALIDATION ---
created_at: "2026-03-01T10:00:00Z"
created_by: "platform-team"
evaluation:
suite: prod-eval-v3
results_uri: s3://evals/llama-3-70b-v1.5.0.json
summary: {mmlu: 0.792, gsm8k: 0.881, our_task: 0.847}
reference_version: v1.4.0
delta_vs_reference: {mmlu: -0.003, our_task: -0.002}
approved_by: "ml-quality-team"
approved_at: "2026-03-02T14:30:00Z"
rollout:
strategy: progressive
canary_percent: 1
guards_ref: standard-llm-guards-v2The section labelled “serving configuration is part of the identity” is the one that distinguishes a real registry from a file listing. Two deployments of the same checkpoint with different quantization are different models from the user’s perspective (Section VIII.10).
3. Why the serving config belongs in the registry#
QUESTION: "Which version of Llama-3-70B is in production?"
WRONG ANSWER: "v1.5.0"
RIGHT ANSWER: "v1.5.0, FP8 per-tensor with FP8 KV, vLLM 0.x.y, TP=8,
chat template hash 9b1c..., evaluated at MMLU 0.792"
Because if any of those change, the outputs change.Practical consequence: when someone reports “the model got worse,” you look up the version record and immediately see everything that could have caused it. Without the record, you’re guessing.
4. The API#
# Discovery
GET /v1/models list all
GET /v1/models/{id} model details
GET /v1/models/{id}/versions all versions
GET /v1/models/{id}/versions/{v} full record
GET /v1/models/{id}/production what's serving now
# Lifecycle
POST /v1/models/{id}/versions register a new version (draft)
POST /v1/models/{id}/versions/{v}/evaluate attach evaluation results
POST /v1/models/{id}/versions/{v}/promote draft → staged → canary → production
POST /v1/models/{id}/versions/{v}/rollback revert to the previous production
POST /v1/models/{id}/versions/{v}/deprecate
# Platform integration
GET /v1/deployments what should be running where
GET /v1/models/{id}/resources what a replica needs (for the scheduler)The promotion path with gates is what makes it a registry rather than a database:
draft ──[evaluation passes]──▶ staged ──[shadow ok]──▶ canary
──[guards hold 2h]──▶ production ──[new version]──▶ deprecatedEach transition has a required condition, checked by the registry, not by whoever runs the deployment.
5. Artifact management#
STORAGE LAYOUT
s3://models/{model_id}/{version}/
├── manifest.json checksums for every file
├── config.json
├── tokenizer/
├── model-00001-of-00030.safetensors
├── ...
└── engine_artifacts/ ← per-GPU-architecture builds
├── h100-tp8/
│ ├── engine.plan (TensorRT) or
│ └── compile_cache/ (torch.compile)
└── h200-tp8/
REGIONAL REPLICATION
push to each region's bucket
→ never pull cross-region per node (egress cost, Section XI.06)
NODE-LOCAL CACHE
/var/cache/models/{model_id}/{version}/
populated by a DaemonSet or init container
LRU eviction with a size limit
→ cold start goes from minutes to seconds (Section II.07)The engine_artifacts directory is often forgotten and is what saves you 60-600 seconds of
compile time per cold start.
6. Integration with the platform#
REGISTRY → SCHEDULER
"model X needs 8 GPUs of type [H100, H200] in one NVLink domain"
→ the scheduler uses this for placement (Section XII.04)
REGISTRY → DEPLOYMENT CONTROLLER
"the production version of model X is v1.5.0 with these engine args"
→ reconciles the actual state to match
REGISTRY → ROUTER
"model X is served by these replicas, with these capabilities"
→ the routing table (Section XII.01)
REGISTRY → GATEWAY
"model alias 'chat-large' maps to model X"
→ user-facing names decoupled from internal versions
REGISTRY → BILLING
"model X costs W per input token, Y per output token"The alias layer is worth calling out: users request chat-large, and the registry maps that
to a specific model and version. This lets you change the underlying model without changing
every client.
7. What to build vs adopt#
For a small platform (< 10 models):
a Git repo of YAML + an object store + a CI pipeline
→ genuinely sufficient. Don't over-build.
For a medium platform (10-100 models, several teams):
a real service with the API from section 4
backed by Postgres + object storage
→ this is where the investment pays
For a large platform:
the above, plus: multi-region replication, approval workflows,
automated evaluation triggering, cost attribution integration
DON'T:
use an experiment tracker (MLflow, W&B) as your deployment registry.
They're built for a different purpose and lack the serving-config
and lifecycle semantics you need. (You can use them alongside.)8. Production implications#
- The registry is the source of truth. If something is serving that isn’t in the registry, that’s an incident.
- Serving configuration is part of the version identity. Record all of it.
- Gate promotions on evaluation results. Make the registry enforce it, not a human’s memory.
- Store engine artifacts per GPU architecture. It’s the difference between 30-second and 10-minute cold starts.
- Version and checksum everything. Including the calibration dataset for quantization.
- Keep deprecated versions deployable for 30+ days for rollback.
- Use aliases so client-facing names are stable.
- Emit the version in every metric and response header (Section XI.07).
9. Common mistakes#
Recording only the checkpoint. The serving config determines the output.
Using an experiment tracker as a deployment registry.
No checksums. A corrupted shard produces plausible garbage.
No engine artifact storage. Recompiling on every cold start.
No aliases. Every client change requires coordination.
Deleting old versions immediately. No rollback path.
Promotion without gates. The evaluation gets skipped under pressure.
No record of the calibration dataset for a quantized model. You can’t reproduce it.
10. Hands-on exercise#
A. Design the schema. Write the full registry entry (section 2 format) for a model you serve. Which fields did you not have recorded anywhere?
B. Build a minimal registry. Implement the API from section 4 backed by SQLite and a filesystem. Include the promotion gates.
C. Verify reproducibility. From your registry entry alone, can someone reproduce the exact deployment? Try it: give the entry to a colleague and have them deploy from it.
D. Artifact pipeline. Build the pipeline: checkpoint → quantize → build engine artifacts per GPU arch → checksum → push to regional storage → register. Time it end to end.
E. Node cache. Implement a DaemonSet (or a script) that populates a node-local model cache from the registry, with LRU eviction. Measure the cold-start improvement.
11. Interview questions#
- What must a model registry entry contain, beyond the checkpoint URI?
- Why is the serving configuration part of the model’s identity?
- What’s the promotion path, and what gates each transition?
- Why store engine artifacts, and how are they keyed?
- Why use aliases rather than exposing version IDs to clients?
- Why is an experiment tracker not a deployment registry?
- What do you keep for rollback and for how long?
12. Further reading#
- [REFERENCE] MLflow Model Registry, KServe InferenceService — for comparison
- [REFERENCE] OCI artifact specifications (some platforms store models as OCI artifacts)
- [ESTABLISHED] Sculley et al., “Hidden Technical Debt in Machine Learning Systems” (2015)
- Next: 03 — Kubernetes for GPUs