1. What is it?#
The contract between clients and your inference service: protocol, schema, streaming semantics, errors, and limits.
For LLMs, the industry has converged on the OpenAI-compatible API as the de facto standard. Deviating from it costs you every existing client library, framework integration, and evaluation harness.
2. Why the convergence matters#
If your API is OpenAI-compatible:
✓ works with openai-python, LangChain, LlamaIndex, Instructor, every eval harness
✓ clients can switch providers by changing a base_url
✓ your users already know the schema
If it isn't:
✗ every integration is bespoke
✗ you maintain client libraries
✗ evaluation tooling doesn't workImplement the OpenAI schema even if you add extensions. The cost of compatibility is a day; the cost of incompatibility is permanent.
3. Simple analogy#
Electrical sockets. A better plug design doesn’t help if nothing plugs into it. The standard is worth more than the improvement.
4. The essential endpoints#
POST /v1/chat/completions the main one
POST /v1/completions legacy text completion
POST /v1/embeddings if you serve embedding models
GET /v1/models model discovery
GET /health, /health/ready liveness and readiness (different!)
GET /metrics Prometheus
POST /tokenize, /detokenize useful extensionsRequest:
{
"model": "llama-3-70b-instruct",
"messages": [
{"role": "system", "content": "You are helpful."},
{"role": "user", "content": "Explain gravity"}
],
"max_tokens": 500,
"temperature": 0.7,
"top_p": 0.9,
"stream": true,
"stop": ["\n\n"],
"seed": 42,
"response_format": {"type": "json_schema", "json_schema": {...}}
}Streaming response chunks:
data: {"id":"...","object":"chat.completion.chunk","created":1234,"model":"...",
"choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"...","choices":[{"index":0,"delta":{"content":"Gravity"},"finish_reason":null}]}
data: {"...","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],
"usage":{"prompt_tokens":12,"completion_tokens":143,"total_tokens":155}}
data: [DONE]Note: usage in the final chunk is an extension many providers now support
(stream_options: {"include_usage": true}). Clients need token counts for cost tracking.
5. Technical explanation#
Protocol choice#
HTTP/1.1 + SSE HTTP/2 gRPC
Public API ✓ THE STANDARD ok needs grpc-web
Internal service ok ✓ ✓ best
Multiplexing 1 req/connection many many
Per-message overhead ~80 bytes text ~20 binary ~10 protobuf
Browser support native EventSource native no
Middlebox friendly ✓ (it's just HTTP) mostly often not
Bidirectional no no (in practice) ✓Practical answer: SSE over HTTP/1.1 for the public API, gRPC between gateway and engines. The 80 bytes of SSE framing per token is negligible (5,000 tok/s × 80 B = 400 KB/s), and the compatibility is worth everything.
Validation — what to enforce at the edge#
max input tokens protects activation memory and prefill time
max output tokens protects KV cache and slot occupancy
max total tokens prompt + max_tokens ≤ model's context limit
temperature ∈ [0, 2]
top_p ∈ (0, 1]
top_k ≥ 1 or absent
n (completions) ≤ small n=100 multiplies your cost by 100
stop sequences: ≤ 4, each ≤ 64 chars
message count and total sizeEvery one of these is a capacity control, not just hygiene. A client sending
max_tokens: 100000, n: 20 can occupy a large fraction of your fleet.
Reject with a clear 400, naming the limit and the observed value.
Errors#
400 invalid request (bad params, too long)
401 unauthenticated
403 unauthorized for this model
404 unknown model
408 request timeout
413 payload too large
422 semantic validation failure
429 rate limited ← include Retry-After
499 client closed request (nginx convention; useful in logs)
500 internal error
503 overloaded — no capacity ← include Retry-After
504 upstream timeout429 vs 503 matters. 429 means “you’re over your quota”; 503 means “we’re over capacity.” Clients should back off differently. Section VIII.05.
Mid-stream errors must go in-band (Section V.14):
data: {"error":{"message":"...","type":"server_error","code":"engine_crash"}}
data: [DONE]Extensions worth having#
Beyond the OpenAI schema, useful additions:
"priority": 0-9 scheduling priority (multi-tenant)
"ignore_eos": bool for benchmarking
"min_tokens": int force a minimum length
"repetition_penalty": float not in the OpenAI schema, widely supported
"min_p": float better than top_p (Section V.13)
"guided_json" / "guided_regex" structured output
"logit_bias": {token_id: bias} in the OpenAI schema
"echo": bool return the prompt too
"lora_request": name which adapter to use (multi-LoRA)Put extensions in a namespaced field or document them clearly, so clients know what’s portable.
Health checks — two different things#
GET /health LIVENESS: is the process alive? Return 200 always if it
responds. Used to decide whether to RESTART the pod.
GET /health/ready READINESS: is the model loaded, warmed up, and able to serve?
Returns 503 during model load and warmup.
Used to decide whether to SEND TRAFFIC.Conflating them is a common and damaging mistake: if liveness fails during a 5-minute model
load, Kubernetes kills the pod and it never starts. Set initialDelaySeconds generously on
liveness, or use a startup probe.
Health checks must be cheap. Never run a real generation on every probe — at a 5-second probe interval across 100 replicas, that’s real GPU capacity spent on health checks.
6-9. Under the hood, performance, production, mistakes#
Under the hood — per-request CPU cost in the API layer:
TLS + HTTP parse 50-200 µs
JSON decode 20-500 µs (scales with message size)
Validation 5-20 µs
Chat template 10-100 µs
Tokenization 50 µs - 5 ms (scales with prompt length)
Per-token: JSON encode + SSE frame + write 20-100 µsAt 5,000 output tokens/sec across all streams, the per-token cost alone is 0.1-0.5 of a core.
Use a fast JSON library (orjson, msgspec) and consider batching token emissions.
Production:
- Be OpenAI-compatible. Test with the actual
openaiclient library in CI. - Enforce all limits at the edge, with clear errors.
- Separate liveness and readiness probes, with a startup probe for model loading.
- Return
usagein the final streaming chunk. - Include
Retry-Afteron 429 and 503. - Log the request shape (input tokens, max_tokens, model, tenant) for every request — it’s the basis of all capacity analysis.
- Version your API (
/v1/) and don’t break it.
Mistakes:
- Inventing your own schema. Loses the ecosystem.
- Conflating liveness and readiness. Restart loops during model load.
- Expensive health checks.
- No limits on
normax_tokens. One client can consume the fleet. - Returning 500 for overload instead of 503. Clients retry immediately and make it worse.
- Not handling mid-stream errors.
- Slow JSON parsing on large messages.
10. Hands-on exercise#
A. Build an OpenAI-compatible endpoint. Implement /v1/chat/completions with streaming and
non-streaming modes. Test it with the real openai Python client — that’s the compatibility
test that matters.
B. Validation suite. Write tests for every limit in section 5. Verify each returns a clear 400 with a useful message.
C. Probe design. Implement liveness, readiness, and startup probes. Simulate a 5-minute model load and verify Kubernetes-style probing doesn’t kill the pod.
D. Measure the API cost. Profile the API layer at 500 req/s with 2,000-token prompts. How many cores does it need? What’s the biggest cost — parsing, tokenizing, or streaming?
E. Error semantics. Implement 429 and 503 with Retry-After. Write a client that backs off
correctly for each and verify the behavior differs.
11. Interview questions#
- Why implement the OpenAI-compatible API rather than your own?
- When would you use gRPC instead of HTTP+SSE?
- What’s the difference between liveness and readiness, and what breaks if you conflate them?
- Which request parameters would you validate at the edge, and why is each a capacity control?
- How do you return an error after streaming has started?
- What’s the difference between 429 and 503, and how should a client treat each?
- What per-request information would you log, and what would you use it for?
12. Further reading#
- [REFERENCE] OpenAI API reference — the de facto schema
- [REFERENCE] vLLM’s OpenAI-compatible server implementation
- [REFERENCE] Kubernetes probe documentation
- [FUNDAMENTAL] Google SRE Book, chapter on load shedding (for error semantics)
- Next: 03 — Queues, scheduling, admission control