PidokuInfra

Multi-Tenancy and Quotas

Expert Advanced 1h 15m Difficulty 4/5 Topic 05 of 11

Prerequisites XI.10, XI.09


1. What multi-tenancy has to provide#

1. ISOLATION       one tenant cannot see another's data
2. FAIRNESS        one tenant cannot starve another
3. ACCOUNTABILITY  you know who used what
4. FLEXIBILITY     tenants can have different SLAs and limits

Section XI.09 covered isolation (security) and XI.10 covered limits (enforcement). This file is about the platform-level design: how tenants share a cluster.


2. The economic case for sharing#

5 teams, each with a dedicated deployment:
  each sized for their peak, at 60% target utilization
  each gets ~15 req/s peak, needs 2 replicas → 10 replicas total
  average batch per replica: 11
  
Shared platform:
  combined peak 75 req/s (peaks don't coincide — this is the key)
  needs 6 replicas at the same target utilization
  average batch per replica: 28
  
  → 40% fewer GPUs, AND 2.5x better per-GPU throughput (Section IX.12)
  → combined effect: ~2.5-4x better economics

Two effects compound: peak smoothing (peaks don’t coincide) and batch concentration. This is why a shared platform beats per-team deployments, and it’s the core value proposition of Section XII.

The organizational cost: teams give up control, and you must provide fairness guarantees convincing enough that they accept it. That is largely a trust problem, solved by transparency (section 6).


3. The isolation spectrum#

LEVEL                     ISOLATION   EFFICIENCY   USE
Shared model, shared      weak        highest      internal, mutually
  engine process                                   trusting teams
Shared model, separate    weak-med    high         different teams,
  prefix cache namespaces                          same trust domain
Separate engine process   medium      medium       untrusted tenants,
  per tenant                                       same model
MIG partition             strong      medium-low   regulatory isolation
Separate node             strongest   low          air-gapped
Separate cluster          strongest   lowest       sovereign / contractual

Most platforms should default to the second row: shared model, per-tenant prefix cache namespaces. It preserves nearly all the efficiency while eliminating the timing side channel (Section XI.09).

Go
// Prefix cache namespacing: include the tenant in the block hash,
// so one tenant can never hit (or time) another tenant's cached prefix.
func blockHash(tenantID string, prevHash uint64, tokenIDs []int32) uint64 {
	h := fnv.New64a()
	h.Write([]byte(tenantID))
	binary.Write(h, binary.LittleEndian, prevHash)
	binary.Write(h, binary.LittleEndian, tokenIDs)
	return h.Sum64()
}

Cost: a shared system prompt used by 5 tenants is cached 5 times. For a 2,000-token system prompt at 128 KiB/token, that’s 5 × 256 MB = 1.28 GB instead of 256 MB. Usually acceptable.


4. Fairness under contention#

Quotas bound the maximum; fairness governs what happens when everyone wants their maximum at once.

WITHOUT FAIRNESS
  tenant A (quota 100k tok/min) and tenant B (quota 10k tok/min)
  both submit continuously.
  FCFS scheduling → A gets ~90% of capacity because it submits more.
  B is effectively starved despite being within quota.

WITH WEIGHTED FAIR QUEUEING
  weights proportional to committed capacity (or tier)
  → each gets its share regardless of submission rate
Go
// drr.go — deficit round robin over tenants, weighted by their share.
package main

import "fmt"

type Request struct {
	Tenant string
	Cost   float64 // weighted tokens
}

type FairScheduler struct {
	weights map[string]float64
	deficit map[string]float64
	queues  map[string][]Request
	order   []string
	idx     int
}

func NewFairScheduler(order []string, weights map[string]float64) *FairScheduler {
	return &FairScheduler{weights, map[string]float64{}, map[string][]Request{}, order, 0}
}

func (s *FairScheduler) Enqueue(r Request) { s.queues[r.Tenant] = append(s.queues[r.Tenant], r) }

func (s *FairScheduler) Dequeue() (Request, bool) {
	pending := false
	for _, q := range s.queues {
		pending = pending || len(q) > 0
	}
	for pending {
		t := s.order[s.idx]
		q := s.queues[t]
		if len(q) > 0 && s.deficit[t] >= q[0].Cost { // keep serving this tenant while its credit lasts
			s.deficit[t] -= q[0].Cost
			s.queues[t] = q[1:]
			return q[0], true
		}
		if len(q) == 0 {
			s.deficit[t] = 0 // don't accumulate credit while idle
		}
		// move on: the next tenant earns its weight in credit
		s.idx = (s.idx + 1) % len(s.order)
		s.deficit[s.order[s.idx]] += s.weights[s.order[s.idx]]
	}
	return Request{}, false
}

func main() {
	s := NewFairScheduler([]string{"big", "small"}, map[string]float64{"big": 300, "small": 100})
	for i := 0; i < 40; i++ { // both tenants flood the queue with identical requests
		s.Enqueue(Request{"big", 100})
		s.Enqueue(Request{"small", 100})
	}
	served := map[string]int{}
	for i := 0; i < 40; i++ {
		r, _ := s.Dequeue()
		served[r.Tenant]++
	}
	fmt.Println(served) // map[big:30 small:10] — a 3:1 split, matching the weights
}

The deficit[t] = 0.0 reset when idle is important: without it, a tenant that was idle accumulates credit and then bursts, starving others.


5. Quota models#

HARD PARTITION
  tenant A gets nodes 0-3, tenant B gets 4-7.
  ✓ perfect isolation, trivially fair
  ✗ no statistical multiplexing — you lose the whole economic benefit
  → only when contractually required

RESERVED + BURST
  tenant A: 20 req/s reserved (guaranteed), may burst to 50 if capacity is free
  ✓ guarantees for those who need them, efficiency from sharing
  ✗ needs admission control that distinguishes reserved from burst traffic
  → THE USUAL ANSWER

PURE SHARED WITH WEIGHTS
  no guarantees; each tenant gets weight/Σweights of capacity under contention
  ✓ maximum efficiency
  ✗ no guarantees; hard to sell to teams with SLAs
  → good for internal platforms with cooperative teams

CREDIT / SPEND
  tenants have a budget; usage draws it down; no rate limit until exhausted
  ✓ aligns with billing
  ✗ doesn't prevent instantaneous contention
  → combine with one of the above

Reserved + burst is what most platforms converge on, because it lets you make credible guarantees while capturing most of the sharing benefit.

IMPLEMENTATION
  Σ reserved ≤ 70% of capacity        (leave room for burst and headroom)
  reserved traffic: admitted always, highest priority
  burst traffic: admitted if capacity available, shed first under load

6. Accountability and transparency#

This is what makes teams accept sharing. If a team can’t see what they’re getting, they’ll demand dedicated capacity.

PER-TENANT DASHBOARD
  requests, tokens (input/output), and cost — daily and monthly
  latency percentiles for THEIR traffic
  quota consumption and headroom
  rejection rate and reasons
  their share of cluster capacity
  comparison to their reserved allocation

PER-TENANT COST (Section XI.03)
  input_tokens × w_in + output_tokens × w_out + kv_block_seconds × w_kv
  
  Publish it. Internally, showback (visibility) changes behavior even
  without chargeback (actual billing).

Showback alone typically reduces waste by 10-30%, because teams discover they’re spending more than they thought on something they didn’t need.


7. Noisy neighbours#

WAYS ONE TENANT DEGRADES ANOTHER

  1. Consuming all KV cache with long-context requests
     → per-tenant KV block limits (Section XI.10)

  2. Long prefills blocking decode
     → chunked prefill (mandatory), and route long prompts to a
       separate pool (Section VIII.06)

  3. Submitting continuously and winning FCFS
     → weighted fair queueing (section 4)

  4. Triggering preemption by pushing memory to the limit
     → per-tenant memory limits, and headroom

  5. Pathological requests (huge n, degenerate grammars)
     → validation limits (Section VIII.02)

  6. Polluting the prefix cache with unique prefixes
     → per-tenant cache namespaces with per-tenant size limits

Item 6 is subtle: a tenant sending millions of unique prompts fills the prefix cache with never-reused entries, evicting other tenants’ useful cached prefixes. Per-tenant cache quotas fix it.


8. Onboarding a tenant#

CHECKLIST
  □ tenant ID, auth credentials (scoped API keys)
  □ tier assignment (free / standard / premium / reserved)
  □ quotas: token rate, concurrency, KV blocks, requests/min
  □ model access list (which models can they call?)
  □ capability scoping (can they use tools? structured output? logprobs?)
  □ data handling agreement (logging, retention, training use)
  □ cost center for showback/chargeback
  □ SLO tier and any contractual commitments
  □ contact for incidents affecting them
  □ dashboard access

CAPACITY IMPACT
  □ does their projected demand fit in existing headroom?
  □ if reserved: does Σ reserved still ≤ 70%?
  □ do they need a model that isn't currently deployed?

The “does their demand fit” check is the one that gets skipped, and it’s how a platform ends up over-subscribed.


9. Production implications#

  • Default to shared model with per-tenant prefix cache namespaces.
  • Implement weighted fair queueing. Quotas alone don’t ensure fairness.
  • Reserved + burst is the quota model that works for platforms with mixed SLA needs.
  • Cap Σ reserved at ~70% of capacity.
  • Publish per-tenant cost and usage. Showback changes behavior.
  • Per-tenant KV block limits, not just token rate limits.
  • Per-tenant prefix cache quotas to prevent cache pollution.
  • Have an onboarding checklist including the capacity check.
  • Give each tenant visibility into their own latency, not just aggregate.

10. Common mistakes#

Quotas without fair queueing. A tenant within quota starves others.

Rate limiting tokens but not concurrency or KV. The real contention is memory.

Shared prefix cache across untrusted tenants. Timing side channel (XI.09).

No per-tenant cache quota. One tenant pollutes the cache.

No showback. Teams have no incentive to be efficient.

Over-committing reserved capacity. No room for burst or headroom.

Onboarding without a capacity check.

Only aggregate dashboards. Teams can’t see their own experience and lose trust.


11. Hands-on exercise#

A. Quantify the sharing benefit. For 5 synthetic tenants with realistic (non-coincident) traffic patterns, compute GPUs needed for dedicated deployments versus a shared platform. Include the batch-size effect.

B. Implement fair queueing. Build the FairScheduler from section 4. Test with tenants submitting at very different rates. Verify each gets its weighted share and that idle tenants don’t accumulate credit.

C. Noisy neighbour. Simulate a tenant sending long-context requests in a shared pool. Measure the impact on other tenants’ latency. Then add per-tenant KV limits and re-measure.

D. Cache pollution. Simulate one tenant sending unique prompts and another sending a shared prefix. Measure the second tenant’s cache hit rate with and without per-tenant cache quotas.

E. Build the tenant dashboard. Implement per-tenant metrics: requests, tokens, cost, latency percentiles, quota consumption. Would a team trust this enough to give up dedicated capacity?


12. Interview questions#

  1. What’s the economic case for a shared inference platform? Quantify it.
  2. Why do quotas alone not ensure fairness?
  3. Explain weighted fair queueing and why idle tenants must not accumulate credit.
  4. What’s the reserved + burst quota model and why is it common?
  5. Name six ways one tenant can degrade another, and the mitigation for each.
  6. Why namespace the prefix cache per tenant, and what does it cost?
  7. Why does showback reduce waste even without chargeback?

13. Further reading#

  • [ESTABLISHED] Weighted fair queueing and deficit round robin literature
  • [FUNDAMENTAL] Google’s Borg paper — multi-tenancy at scale
  • [REFERENCE] Kubernetes ResourceQuota and LimitRange
  • Next: 06 — Routing strategies

↑↓ navigate↵ openesc close