PidokuInfra

Long-Context Inference

Expert Advanced 1h 30m Difficulty 4/5 Topic 04 of 12

Prerequisites V.07, 01


1. The four costs, restated#

Section V.07 established them; this file is about the techniques that attack each.

COST                        SCALING   ATTACKED BY
KV memory                   O(S)      GQA/MLA, KV quant, sliding window,
                                      eviction, offload
KV read bandwidth per step  O(S)      same, plus sparse attention
Prefill compute             O(S²)     chunked prefill (doesn't reduce, but
                                      makes it schedulable), sparse attention,
                                      prefix caching
Concurrency                 O(1/S)    everything above

2. The techniques, honestly labelled#

[ESTABLISHED] — deploy these
  GQA / MQA                      4-8x KV reduction, no quality cost
  MLA                            10-60x KV reduction (DeepSeek architecture)
  FP8 KV cache                   2x, ~free
  Sliding window attention       caps KV at W (architectural)
  Chunked prefill                makes long prefill schedulable
  Prefix caching                 eliminates redundant prefill
  FlashAttention / FlashDecoding required for feasibility at all

[EMERGING] — evaluate carefully
  INT4 KV cache                  4x, measurable quality cost
  Attention sinks (StreamingLLU) bounded memory for infinite streams
  KV offload to CPU/NVMe         capacity, not bandwidth
  Disaggregated prefill/decode   (file 06)
  Cross-layer KV sharing         2x, limited production deployment

[RESEARCH] — interesting, not production
  KV eviction (H2O, SnapKV)      3-5x, SILENT failure mode
  Learned KV compression
  Sparse attention patterns      (beyond sliding window)
  Linear attention / SSM hybrids O(1) state, different quality profile

The eviction row is the one to be careful about. It appears in many papers with impressive benchmark results and has a failure mode — silently losing information the model needed — that your monitoring cannot detect.


3. Context extension: how models get long context#

Models are trained at some length and extended. The methods matter to you because they change rope_theta, and using the wrong value silently degrades output at long positions.

POSITION INTERPOLATION (PI)
  Compress positions into the trained range: pos' = pos × (L_train/L_target)
  ✓ simple
  ✗ degrades short-context performance

NTK-AWARE SCALING
  Scale rope_theta rather than positions: theta' = theta × s^(d/(d-2))
  ✓ preserves short-context performance better
  → what most models use

YARN
  NTK scaling + attention temperature correction + selective
  interpolation by frequency band
  ✓ better than NTK alone
  ✗ more parameters to get right

LONGROPE
  Search for per-dimension rescaling factors
  ✓ best reported results
  ✗ requires the search
JSON
// In config.json — READ THIS before deploying
{
  "rope_theta": 500000.0,          // not the base model's 10000!
  "rope_scaling": {
    "type": "yarn",
    "factor": 8.0,
    "original_max_position_embeddings": 8192
  },
  "max_position_embeddings": 131072
}

Verify your engine reads and applies rope_scaling. If it applies plain RoPE with the original theta, output at position 50,000 will be degraded with no error.


4. Effective context vs advertised context#

A model advertising 128k may not USE 128k well.

MEASUREMENT: needle-in-a-haystack and its successors
  place a fact at depth d in a context of length L
  ask a question requiring it
  measure retrieval accuracy as a function of (L, d)

TYPICAL FINDINGS
  ✓ accurate at the beginning and end of the context
  ✗ degraded in the middle ("lost in the middle", Liu et al. 2023)
  ✗ accuracy falls with L, often well before the advertised limit
  ✗ multi-fact retrieval degrades much faster than single-fact
  ✗ reasoning over long context degrades faster than retrieval

Before building a product on 128k context, measure whether the model uses it for your task. The cost is 20x (Section V.07); paying it for capability you don’t get is expensive.

A PRACTICAL TEST SUITE
  1. single needle, varying depth and length      (easiest)
  2. multiple needles requiring aggregation
  3. a question requiring reasoning over two distant facts
  4. your actual task at varying context lengths
  
  → plot accuracy vs context length. Where does YOUR task break?

5. RAG vs long context — the cost argument#

TASK: answer questions about a 200-page document (~150k tokens).

OPTION A — STUFF THE CONTEXT
  prefill 150k tokens per question
  prefill FLOPs ≈ 2×P×S + 2×L×S²×d
    for an 8B model: 2.4e15 + 5.9e15 = 8.3 PFLOP
    at 600 TFLOP/s: 14 seconds of TTFT
  KV: 150k × 128 KiB = 19.2 GB per sequence
  → ~3 concurrent users on an H100
  
  WITH PREFIX CACHING (same document, many questions):
    first question: 14 s
    subsequent:     ~0.1 s (the document's KV is cached)
    → transforms the economics IF questions share the document

OPTION B — RAG
  embed and index the document once (offline)
  per question: retrieve 5 chunks × 800 tokens = 4k tokens
  prefill 4k tokens: 0.06 s TTFT
  KV: 0.5 GB per sequence → ~120 concurrent users
  → 40x cheaper per question

QUALITY
  RAG wins when the answer is in a retrievable chunk.
  Long context wins when the answer requires synthesis across the
  whole document, or when retrieval fails.

The right answer is usually: RAG for retrieval-shaped questions, long context with prefix caching for synthesis-shaped ones, and measure which your task is.

Note that prefix caching changes the calculation dramatically when many questions share one document — which is the common case. Long context with a warm prefix cache is competitive with RAG.


6. Sparse and windowed attention#

SLIDING WINDOW (Mistral)  [ESTABLISHED]
  each token attends to the last W tokens
  → KV capped at W regardless of S
  → information beyond W propagates through layers (L×W theoretical reach)

INTERLEAVED (Gemma 2, others)  [ESTABLISHED]
  most layers sliding-window, a few full attention
  → caps most of the KV while retaining true long-range access
  → e.g. 5 local layers per 1 global layer → ~83% KV reduction

ATTENTION SINKS (StreamingLLM)  [EMERGING]
  keep the first few tokens (which act as attention "sinks") plus a
  recent window
  → enables unbounded streaming with bounded memory
  → the model doesn't GAIN long-range ability; it just doesn't break
  → some models now train with explicit sink tokens

STRUCTURED SPARSITY  [RESEARCH]
  strided, dilated, or learned sparse patterns
  → promising in papers; limited production deployment
  → the kernel efficiency of irregular patterns is the obstacle

Interleaved local/global attention is the design most likely to become standard, because it captures most of the memory saving without giving up long-range capability.


7. KV eviction — why to be careful#

THE IDEA: not all cached tokens matter. Keep the important ones.

H2O:      keep tokens with high cumulative attention scores
SnapKV:   at the end of prefill, use the last queries' attention to
          select which prompt tokens to keep
PyramidKV: allocate more KV budget to lower layers

REPORTED: 3-5x KV reduction with "minimal" quality loss.

THE PROBLEM
  Quality loss is TASK-DEPENDENT and the failure mode is SILENT.
  
  If the model needed an evicted token, it doesn't error — it
  CONFABULATES. Your monitoring sees a normal response.
  
  Benchmarks that show minimal loss typically test tasks where the
  relevant information is recent or highly attended. Tasks requiring
  a specific fact from the middle of a long document fail badly.

IF YOU DEPLOY IT
  □ evaluate on YOUR task, specifically on cases requiring information
    from evicted positions
  □ needle-in-a-haystack at varying depths, WITH eviction enabled
  □ have a fallback: detect low confidence and re-run without eviction
  □ don't use it for tasks where a wrong answer is costly

This is the technique in this curriculum most often presented as ready when it isn’t. It works; the question is whether it works for your task, and you must measure that specifically.


8. Production implications#

  • Cap context by tier. Long context is genuinely expensive; price and limit it.
  • Segregate long-context requests into their own pool (Section VIII.06). One 128k request poisons a pool tuned for 4k.
  • Chunked prefill is mandatory if you accept long prompts and have an ITL SLO.
  • Prefix caching is the highest-value long-context optimization when documents recur.
  • Verify rope_scaling is applied.
  • Measure effective context for your task before promising it.
  • Evaluate RAG as an alternative. Often 40x cheaper.
  • FP8 KV cache first, before considering eviction.
  • Be conservative with eviction. Silent failures.

9. Common mistakes#

Deploying a long-context fine-tune without checking rope_scaling.

Promising 128k without measuring effective context.

Mixing long and short requests in one pool.

Not chunking prefill. An 11-second prefill freezes everyone.

Deploying eviction based on paper benchmarks.

Not evaluating RAG as an alternative.

Assuming cost scales linearly with context. It’s superlinear (Section V.07).


10. Hands-on exercise#

A. Measure effective context. Build a needle-in-a-haystack test: place a fact at depths 0%, 25%, 50%, 75%, 100% in contexts of 4k, 16k, 64k, 128k. Plot accuracy. Where does your model actually break?

B. Multi-fact. Extend A to require two facts at different depths. How much faster does accuracy degrade?

C. RAG comparison. For a document-QA task, implement both: full-context and RAG with retrieval. Measure cost per question and answer quality. Where’s the crossover?

D. Prefix caching effect. Measure the cost of 20 questions about one 100k-token document, with and without prefix caching. Quantify the difference.

E. Eviction risk. Implement a simple eviction scheme (keep first 4 + last N). Run your needle test with it enabled. At what eviction ratio does retrieval fail? Is the failure silent?

F. rope_scaling. Take a long-context model and deliberately serve it with the base rope_theta. Measure output quality at position 1k vs 50k. How bad is it, and would you have noticed?


11. Interview questions#

  1. What are the four costs of long context, and how does each scale?
  2. What is rope_scaling and what happens if you ignore it?
  3. What is “effective context” and how would you measure it?
  4. When is RAG better than long context? Give the cost argument.
  5. What is interleaved local/global attention and what does it save?
  6. Why are KV eviction methods risky in production?
  7. What would you deploy today for a 128k-context service?

12. Further reading#

  • [ESTABLISHED] Peng et al., “YaRN” (2023); Su et al., “RoFormer” (RoPE, 2021)
  • [ESTABLISHED] Liu et al., “Lost in the Middle” (2023)
  • [ESTABLISHED] Jiang et al., “Mistral 7B” (2023) — sliding window
  • [EMERGING] Xiao et al., “StreamingLLM” (2023)
  • [RESEARCH] Zhang et al., “H2O” (2023); Li et al., “SnapKV” (2024)
  • Next: 05 — Chunked prefill

↑↓ navigate↵ openesc close