PidokuInfra

Inference Gateways

Expert Advanced 1h 15m Difficulty 4/5 Topic 08 of 11

Prerequisites VIII.02, XI.10, 06


1. What a gateway does#

The single front door for all inference traffic. Everything common to every request lives here, so it doesn’t have to live in every engine.

CLIENT ──► GATEWAY ──► ROUTER ──► ENGINE

GATEWAY RESPONSIBILITIES
  TLS termination
  authentication and authorization
  rate limiting and quota enforcement
  request validation (limits, schema)
  model alias resolution
  chat template application
  tokenization
  request/response logging and metrics
  streaming (SSE) framing
  detokenization
  error normalization
  usage accounting for billing

The gateway is the place where the platform’s policy lives. Engines should know nothing about tenants, quotas, or billing.


2. Why a separate component#

IF YOU PUT THIS IN THE ENGINE
  ✗ every engine process runs auth, rate limiting, and billing logic
  ✗ the GIL contends with the engine loop (Section VIII.01)
  ✗ policy changes require redeploying models (minutes each)
  ✗ N copies of the quota state to keep consistent
  ✗ engines become platform-specific rather than swappable

WITH A GATEWAY
  ✓ policy changes deploy in seconds
  ✓ engines stay generic and swappable
  ✓ one place for auth, billing, and observability
  ✓ scales independently (CPU-bound, cheap to replicate)

3. The request pipeline#

1.  TLS terminate
2.  Parse HTTP, decode JSON               (orjson/msgspec — 3-10x faster)
3.  Authenticate                          (JWT verify, or API key lookup)
4.  Authorize                             (may this key call this model?)
5.  Validate                              (limits, parameter ranges)
6.  Resolve alias                         ("chat-large" → model+version)
7.  Apply chat template                   (from the registry)
8.  Tokenize                              (thread pool; Rust tokenizer)
9.  Rate limit / quota check              (tokens, concurrency, requests)
10. Route                                 (Section XII.06)
11. Forward to the engine                 (gRPC or HTTP)
12. Stream response back
      - detokenize incrementally          (Section V.02)
      - frame as SSE
      - detect client disconnect → abort  (Section V.04)
13. Record usage, emit metrics and logs
14. Release the concurrency slot, refund unused quota

Steps 8 and 12 are the CPU-heavy ones. At 5,000 output tokens/sec across all streams, the per-token detokenize + SSE frame + write costs 0.1-0.5 of a core.


4. Sizing and scaling the gateway#

PER-REQUEST CPU COST
  parse + validate               50-200 µs
  auth (cached)                  10-30 µs
  auth (uncached, JWT verify)    100-500 µs
  chat template                  10-100 µs
  tokenize (per 1k tokens)       ~50 µs
  routing decision               10-50 µs

PER-TOKEN CPU COST (streaming)
  detokenize + SSE frame + write 20-100 µs

SIZING
  at 100 req/s with 2k prompts and 400 output tokens:
    per-request: 100 × 500 µs = 0.05 core
    tokenize:    100 × 100 µs = 0.01 core
    per-token:   100 × 400 × 50 µs = 2.0 cores    ← dominates
  → ~3 cores plus headroom. Small.
  
  at 1,000 req/s: ~25 cores. Run 4-6 gateway pods.

The gateway is cheap relative to the GPUs, which is why putting work here rather than in the engine is almost always right.

Scale it independently. Gateway pods are stateless (with cached auth and routing tables) and scale in seconds — unlike engines.


5. Auth and authorization#

API KEYS
  hash them; store the hash. Never log the key.
  cache the lookup (key_hash → tenant, scopes) with a short TTL.
  → the cache is what keeps auth off the critical path when the auth
    store is slow or down (Section XI.05: fail-closed with a cache
    that rides out brief outages)

JWT
  verify the signature locally with a cached JWKS.
  ✓ no per-request network call
  ✗ revocation is hard (short expiry + a revocation list)

SCOPES
  what can this key do?
    models: [chat-large, chat-small]
    capabilities: [streaming, structured_output]
    max_context: 32768
    tier: standard
  → enforce at the gateway, so engines don't need to know

Capability scoping is under-used and valuable: restricting which models and features a key can access limits both cost exposure and blast radius.


6. Streaming through the gateway#

This is where most gateway implementations get it wrong.

Go
func (g *Gateway) chat(w http.ResponseWriter, r *http.Request) {
	rc, err := g.preamble(r) // auth, validate, tokenize
	if err != nil {
		writeError(w, err)
		return
	}
	ctx := r.Context() // cancelled when the client disconnects

	if !rc.Body.Stream {
		result, err := g.engine.Generate(ctx, rc)
		if err != nil {
			writeError(w, err)
			return
		}
		g.recordUsage(rc, result.OutputTokens)
		json.NewEncoder(w).Encode(toOpenAI(result))
		return
	}

	w.Header().Set("Content-Type", "text/event-stream")
	w.Header().Set("Cache-Control", "no-cache")
	w.Header().Set("X-Accel-Buffering", "no")
	w.Header().Set("X-Model-Version", rc.ModelVersion)
	flusher := w.(http.Flusher)
	send := func(v any) {
		b, _ := json.Marshal(v)
		fmt.Fprintf(w, "data: %s\n\n", b)
		flusher.Flush()
	}

	detok := NewIncrementalDetokenizer(rc.Tokenizer)
	stop := NewStopSequenceMatcher(rc.Body.Stop)
	nOut := 0
	defer func() { // runs on every exit path, including disconnects
		g.recordUsage(rc, nOut)
		g.releaseQuota(rc, nOut) // refund unused
	}()

	tokens, errs := g.engine.GenerateStream(ctx, rc) // the engine aborts when ctx is cancelled
	for tokens != nil {
		select {
		case <-ctx.Done(): // 1. disconnect — CRITICAL (Section V.14)
			return
		case err := <-errs:
			// in-band error — we already sent 200 (Section V.14)
			send(map[string]any{"error": map[string]string{"message": err.Error(), "type": "server_error"}})
			tokens = nil
		case id, ok := <-tokens:
			if !ok {
				tokens = nil
				break
			}
			delta := detok.Add(id) // 2. incremental detokenization (Section V.02)
			if delta == "" {
				continue // incomplete UTF-8
			}
			emit, done := stop.Feed(delta) // 3. stop sequence hold-back
			if emit != "" {
				nOut++
				send(chunk(emit))
			}
			if done {
				tokens = nil
			}
		}
	}
	send(finalChunk(usage(rc, nOut)))
	fmt.Fprint(w, "data: [DONE]\n\n")
	flusher.Flush()
}

Five things this gets right that implementations commonly get wrong:

  1. Disconnect detection and abort propagation.
  2. Incremental detokenization (no mojibake).
  3. Stop-sequence hold-back.
  4. In-band error reporting after streaming has started.
  5. Quota refund in finally.

7. Deployment shape#

                    ┌──────────────────────────┐
    Internet ──────►│  L7 LB / Ingress         │  TLS, WAF, DDoS
                    │  (proxy_buffering OFF!)  │
                    └────────────┬─────────────┘
                                 │
                    ┌────────────▼─────────────┐
                    │  GATEWAY (4-8 pods)      │  stateless, CPU-only
                    │  auth, quota, tokenize   │  scales in seconds
                    └────────────┬─────────────┘
                                 │
                    ┌────────────▼─────────────┐
                    │  ROUTER (2-4 pods)       │  routing table + load state
                    └────────────┬─────────────┘
                                 │
                    ┌────────────▼─────────────┐
                    │  ENGINES                 │  GPU
                    └──────────────────────────┘

Sometimes the gateway and router are one component. That's fine at small
scale; separate them when the routing logic grows.

Every layer above the engines must have buffering disabled (Section II.08). Test the whole path from a real client.


8. Build vs adopt#

GENERIC API GATEWAYS (Kong, Envoy, APISIX, cloud gateways)
  ✓ auth, TLS, basic rate limiting, observability
  ✗ don't understand tokens (can't rate limit by tokens)
  ✗ don't tokenize or apply chat templates
  ✗ don't understand streaming semantics for LLMs
  ✗ can't do prefix-aware routing
  → USE FOR: TLS, WAF, coarse rate limiting at the edge
  → NOT SUFFICIENT as the inference gateway

LLM-SPECIFIC GATEWAYS (LiteLLM, Portkey, Kong AI Gateway, and others)
  ✓ multi-provider abstraction, token-aware limiting, caching
  ✗ maturity and performance vary; evaluate carefully
  → worth evaluating; may cover 80% of your needs

BUILD
  ✓ full control over routing, quota semantics, and observability
  ✗ real ongoing work
  → the usual answer for a serious platform, often layered behind a
    generic gateway that handles TLS and edge concerns

A common and sensible architecture: a generic gateway at the edge (TLS, WAF, IP rate limits) plus a custom inference gateway behind it (tokens, quotas, routing, streaming).


9. Production implications#

  • The gateway owns policy; engines own inference. Keep the separation.
  • Disable buffering at every layer above the engine.
  • Cache auth lookups so the auth store isn’t in the critical path.
  • Rate limit by tokens and concurrency at the gateway (Section XI.10).
  • Implement disconnect propagation. Highest-value item in the streaming path.
  • Use fast JSON (orjson/msgspec) and Rust tokenizers.
  • Emit X-Model-Version so client-side issues can be correlated with deployments.
  • Scale the gateway independently. It’s cheap and fast to scale.
  • Record usage in a finally block. Otherwise aborted requests aren’t billed or accounted.

10. Common mistakes#

Putting policy in the engine. Slow to change, contends with the engine loop.

Proxy buffering enabled somewhere. Destroys streaming (Section X, case 8).

Per-request auth store lookups. Latency and a hard dependency.

Request-based rate limiting. (Section XI.10.)

No disconnect detection. 10-30% capacity waste.

Naive detokenization. Mojibake for non-Latin text.

No stop-sequence hold-back. Stop strings flash in the output.

Slow JSON parsing. Measurable at high QPS with large prompts.

Not recording usage on the error path.


11. Hands-on exercise#

A. Build the gateway. Implement the pipeline from section 3, with the streaming handler from section 6. Test with the real openai client library. This is Project 12.

B. Test the streaming correctness. Verify: multi-byte UTF-8 across token boundaries, stop sequences spanning tokens, in-band errors, disconnect abort, and usage recording on all paths.

C. Measure the CPU cost. Load-test at 500 req/s with 2k prompts. Profile the gateway (Section X.09). Where does CPU go? How many cores per 1,000 req/s?

D. Auth caching. Implement the auth cache with TTL. Kill the auth store and verify requests continue for the TTL duration. Measure the latency difference cached vs uncached.

E. End-to-end streaming test. Deploy behind a real ingress. Measure client-side TTFT. Deliberately enable buffering somewhere and confirm your test catches it.

F. Quota refund. Verify that a request with max_tokens: 4096 that generates 100 tokens refunds the difference, and that an aborted request refunds correctly.


12. Interview questions#

  1. What belongs in the gateway rather than the engine, and why?
  2. Walk through the request pipeline in order.
  3. What’s the CPU cost of a gateway, and what dominates it?
  4. How do you handle an error after streaming has started?
  5. Why cache auth lookups, and what’s the failure mode without it?
  6. Why is a generic API gateway insufficient for LLM traffic?
  7. What must the gateway do in a finally block, and why?

13. Further reading#

  • [REFERENCE] OpenAI API specification — the schema to implement
  • [REFERENCE] Envoy, Kong documentation for the general patterns
  • [REFERENCE] LiteLLM and similar LLM gateway projects — for comparison
  • [REFERENCE] MDN Server-Sent Events
  • Next: 09 — Cascades and fallbacks

↑↓ navigate↵ openesc close