1. Why request-based rate limiting is wrong#
"1,000 requests per minute" allows:
Tenant A: 1,000 requests × 50 input + 100 output tokens
= 150,000 tokens/min
Tenant B: 1,000 requests × 100,000 input + 4,000 output tokens
= 104,000,000 tokens/min
Same limit. 693x the resource consumption.Every LLM rate limiting scheme must be denominated in the resource that’s actually scarce, which is a combination of tokens and memory-time.
2. The three dimensions to limit#
1. TOKEN RATE (throughput)
input tokens/min and output tokens/min, separately
(output costs ~5x input — Section XI.03)
2. CONCURRENCY (memory occupancy)
simultaneous in-flight requests, or better:
simultaneous KV blocks held
→ this is what actually denies capacity to others
3. REQUEST RATE (overhead)
requests/min — bounds gateway and scheduler overhead
→ a secondary limit; the first two matter moreLimiting only #1 is the common mistake. A tenant with 50 concurrent 128k-context requests has modest token throughput and enormous memory occupancy.
3. The implementation#
// limiter.go — per-tenant limiting in three dimensions: concurrency, memory, token rate.
package main
import (
"fmt"
"sync"
"time"
)
// TokenBucket is a standard token bucket, denominated in WEIGHTED tokens.
type TokenBucket struct {
capacity, refillPerSec, tokens float64
last time.Time
}
func NewTokenBucket(capacity, refillPerSec float64) *TokenBucket {
return &TokenBucket{capacity, refillPerSec, capacity, time.Now()}
}
// TryConsume returns ok, or how long to wait (→ the Retry-After header).
func (b *TokenBucket) TryConsume(amount float64) (ok bool, retryAfter time.Duration) {
now := time.Now()
b.tokens = min(b.capacity, b.tokens+now.Sub(b.last).Seconds()*b.refillPerSec)
b.last = now
if b.tokens >= amount {
b.tokens -= amount
return true, 0
}
return false, time.Duration((amount - b.tokens) / b.refillPerSec * float64(time.Second))
}
const wIn, wOut = 1.0, 4.0 // an output token costs ~4x an input token
type Request struct{ PromptTokens, MaxTokens, EstKVBlocks int }
type TenantLimiter struct {
mu sync.Mutex
tokens, requests *TokenBucket
maxConcurrent, maxKVBlocks int
inFlight, kvBlocksHeld int
}
// Admit returns "" to admit, or the reason for rejecting plus a Retry-After hint.
func (l *TenantLimiter) Admit(r Request) (reason string, retryAfter time.Duration) {
l.mu.Lock()
defer l.mu.Unlock()
// 1. concurrency and memory (checked first — cheapest and most important)
if l.inFlight >= l.maxConcurrent {
return "concurrency_limit", time.Second
}
if l.kvBlocksHeld+r.EstKVBlocks > l.maxKVBlocks {
return "memory_limit", 2 * time.Second
}
// 2. token rate — charge the ESTIMATE up front
cost := float64(r.PromptTokens)*wIn + float64(r.MaxTokens)*wOut
if ok, wait := l.tokens.TryConsume(cost); !ok {
return "token_rate_limit", wait
}
// 3. request rate
if ok, wait := l.requests.TryConsume(1); !ok {
return "request_rate_limit", wait
}
l.inFlight++
l.kvBlocksHeld += r.EstKVBlocks
return "", 0
}
// Complete REFUNDS the difference: we charged for MaxTokens, they used fewer.
func (l *TenantLimiter) Complete(r Request, actualOutputTokens int) {
l.mu.Lock()
defer l.mu.Unlock()
refund := float64(r.MaxTokens-actualOutputTokens) * wOut
l.tokens.tokens = min(l.tokens.capacity, l.tokens.tokens+refund)
l.inFlight--
l.kvBlocksHeld -= r.EstKVBlocks
}
func main() {
l := &TenantLimiter{tokens: NewTokenBucket(30000, 2000), requests: NewTokenBucket(10, 5),
maxConcurrent: 4, maxKVBlocks: 1000}
r := Request{PromptTokens: 1000, MaxTokens: 1000, EstKVBlocks: 125}
for i := 0; i < 6; i++ {
reason, wait := l.Admit(r)
fmt.Printf("request %d: %-18q retry after %v\n", i, reason, wait.Round(time.Millisecond))
}
l.Complete(r, 50) // finished early: 950 unused output tokens are refunded
reason, _ := l.Admit(r)
fmt.Printf("after a completion + refund: %q\n", reason)
}The charge-then-refund pattern is important. You must charge for max_tokens at admission
(you don’t know the actual output length yet), or a tenant can request max_tokens: 100000
repeatedly and only be charged after the fact. Refunding the unused portion keeps it fair.
4. Where to enforce#
LAYER ENFORCES WHY THERE
CDN / WAF gross abuse, IP bans cheapest rejection
API gateway per-key limits before any expensive work
(tokens, concurrency)
Router per-model limits capacity-aware
Engine scheduler admission control knows actual system state
(Section VIII.03)Enforce as early as possible. Rejecting at the CDN costs nothing; rejecting at the engine means you already tokenized, routed, and queued.
But the engine’s admission control is the only layer that knows the actual system state, so it must also be able to reject — a tenant within their quota can still be rejected when the system is over capacity.
5. Quota design#
QUOTA TYPES
hard limit reject when exceeded
soft limit allow, but deprioritize
burst short-term allowance above the sustained rate
credit pre-purchased budget, consumed over time
TIME WINDOWS
per second smooths bursts, high enforcement cost
per minute the usual granularity
per day/month billing-aligned; needs persistent state
A GOOD SCHEME COMBINES:
per-second token bucket → smooth, prevents instantaneous bursts
per-minute limit → the published limit users understand
per-month credit → the billing relationship
concurrency limit → memory protectionPublish the per-minute limit; enforce with the per-second bucket. Users understand “60,000 tokens/minute”; the bucket enforces it smoothly rather than allowing 60,000 in the first second and nothing for 59 more.
6. Fairness across tenants#
Under contention, quotas alone don’t ensure fairness. You need scheduling policy too.
WEIGHTED FAIR QUEUEING
Each tenant gets a share of throughput proportional to its weight.
Implementation: track per-tenant service received; admit from the
tenant with the lowest (service / weight) ratio.
→ prevents one tenant with a huge quota from starving others
during contention// selectNext picks the waiting tenant furthest below its fair share.
func selectNext(waiting map[string][]*Request, serviceReceived, weights map[string]float64) *Request {
best, bestShare := "", math.Inf(1)
for tenant, queue := range waiting {
if share := serviceReceived[tenant] / weights[tenant]; len(queue) > 0 && share < bestShare {
best, bestShare = tenant, share
}
}
if best == "" {
return nil
}
r := waiting[best][0]
waiting[best] = waiting[best][1:]
return r
}With service_received measured in weighted tokens (input × w_in + output × w_out +
kv_block_seconds × w_kv), this gives genuinely fair sharing of the scarce resource.
Add aging so a low-weight tenant doesn’t starve indefinitely.
7. Detecting abuse#
SIGNALS
request pattern: systematic enumeration (sequential IDs, exhaustive prompts)
content: repeated near-identical prompts (distillation)
known jailbreak patterns
economics: cost per tenant far exceeding their tier
timing: perfectly regular intervals (bot behavior)
errors: high rate of limit violations (probing)
output: requesting full logprobs systematically
RESPONSES, ESCALATING
1. log and monitor
2. tighten limits for that key
3. add friction (CAPTCHA at the application layer, slower responses)
4. soft ban (accept but heavily deprioritize)
5. hard banDistillation detection is genuinely hard and the signal is weak: legitimate high-volume use looks similar. Practical approach: monitor cost per tenant against tier, investigate outliers manually, and rely on terms of service for enforcement.
8. The user experience of limits#
GOOD 429 RESPONSE
HTTP/1.1 429 Too Many Requests
Retry-After: 4
X-RateLimit-Limit-Tokens: 60000
X-RateLimit-Remaining-Tokens: 0
X-RateLimit-Reset-Tokens: 1710412800
X-RateLimit-Limit-Requests: 500
X-RateLimit-Remaining-Requests: 143
{"error": {
"message": "Rate limit exceeded: 60000 tokens/min. Retry in 4s.",
"type": "rate_limit_error",
"code": "token_rate_limit"}}
WHY THIS MATTERS
- Retry-After lets clients back off correctly instead of hammering
- the headers let clients self-throttle proactively
- a specific message tells them WHICH limit and what to do
- a machine-readable code lets client libraries handle itHeaders that let clients self-throttle reduce your load, because well-behaved clients slow down before hitting the limit.
9. Production implications#
- Limit tokens, concurrency, AND requests. All three.
- Weight output tokens ~5x input tokens.
- Charge
max_tokensat admission, refund the difference. - Limit concurrent KV blocks, not just concurrent requests.
- Return proper 429s with
Retry-Afterand limit headers. - Implement weighted fair queueing for multi-tenant fairness under contention.
- Monitor cost per tenant and alert on outliers.
- Enforce at the earliest possible layer, but keep engine-level admission control too.
- Publish your limits clearly.
10. Common mistakes#
Request-based rate limiting. 693x variance in actual consumption.
No concurrency limit. One tenant occupies all the KV cache.
Charging only actual output tokens. A tenant can request huge max_tokens freely.
No Retry-After. Clients retry immediately and make it worse.
Quotas without fair queueing. A tenant within quota can still starve others under contention.
Fail-closed rate limiter. Redis blips and the whole service rejects (Section XI.05).
Not monitoring cost per tenant. You discover abuse from the bill.
Limits that aren’t published. Users can’t design around them.
11. Hands-on exercise#
A. Implement the three-dimensional limiter. Build the TenantLimiter from section 3.
Include the charge-and-refund logic. Test with a workload that requests large max_tokens but
generates few tokens.
B. Demonstrate the request-limit failure. Simulate two tenants under a request-based limit, one sending tiny requests and one sending huge ones. Measure actual resource consumption. Then switch to token+concurrency limiting and re-measure.
C. Fair queueing. Implement weighted fair queueing across three tenants with different weights. Verify each receives its share under contention, and that aging prevents starvation.
D. Client headers. Implement the full 429 response with headers. Write a client that self-throttles based on the headers. Measure the reduction in rejected requests.
E. Cost monitoring. Compute cost per tenant (Section XI.03) and identify the outliers in a synthetic multi-tenant workload. What threshold would you alert on?
F. KV occupancy limit. Implement a per-tenant limit on concurrent KV blocks. Verify it prevents a single tenant from monopolizing memory with long-context requests.
12. Interview questions#
- Why is request-based rate limiting wrong for LLM services? Quantify it.
- What three dimensions would you limit, and why each?
- Why charge
max_tokensat admission and refund afterwards? - Why limit concurrent KV blocks rather than concurrent requests?
- What is weighted fair queueing and why do quotas alone not ensure fairness?
- What should a 429 response contain?
- How would you detect a tenant attempting to distill your model?
13. Further reading#
- [FUNDAMENTAL] Token bucket and leaky bucket algorithms
- [ESTABLISHED] Weighted fair queueing literature (Demers, Keshav, Shenker)
- [REFERENCE] Published rate limit documentation from LLM API providers — note that they all limit tokens and concurrency, not just requests
- Next: 11 — Scenario: 70B, 10,000 users, fixed budget