1. What it is#
Replace the dense FFN in each transformer layer with E expert FFNs plus a router. Each token is sent to only k of them (typically k=1 or 2).
DENSE FFN MoE FFN
x → W_up → act → W_down → y x → router → picks experts {3, 17}
every parameter used x → expert_3 →┐
for every token x → expert_17 →┴→ weighted sum → y
only 2 of E experts usedCONSEQUENCE
total parameters: large (E × FFN size, × L layers)
active parameters: small (k × FFN size, × L layers)
DeepSeek-V3: 671B total, 37B active per token
Mixtral 8x7B: 47B total, 13B active per tokenDiagram — Only a few experts run per token#
flowchart LR
X["Token hidden state"] --> RT["Router / gate"]
RT -->|"top-k"| E3["Expert 3"]
RT -->|"top-k"| E7["Expert 7"]
RT -.-> EO["Other experts<br/>not executed for this token"]
E3 --> SUM(("weighted sum"))
E7 --> SUM
SUM --> Y["Layer output"]
class RT queue
class E3,E7 compute
class EO neutral
class SUM memory
class X,Y neutral2. Why it exists#
Model quality scales with parameters; compute scales with active parameters. MoE decouples them.
A dense 37B model: 37B params, 37B active. Quality of a 37B model.
DeepSeek-V3: 671B params, 37B active. Quality closer to a much
larger dense model, at 37B-model compute per token.For training, this is a straightforward win. For inference, it’s a trade, and the trade is the interesting part:
Dense 37B MoE 671B/37B active
FLOPs per token 74 GFLOP 74 GFLOP same
Memory (weights) 74 GB 1,342 GB 18x more
Decode bytes read 74 GB ~74 GB* similar
(*only the active experts' weights,
IF you can read selectively)
GPUs to hold it 1-2 16-20
Quality 37B-level much betterMoE trades memory capacity for quality at constant compute. Whether that’s a good trade depends entirely on whether you have the memory.
3. Simple analogy#
A hospital with specialists versus general practitioners.
A dense model is a GP who has studied everything and applies all of it to every patient.
An MoE is a hospital with 256 specialists. Each patient sees 2 of them. The hospital knows more in total, and each consultation costs the same as seeing a GP — but you must employ all 256 specialists whether or not they see patients today.
The employment cost is the memory cost, and it’s the whole difficulty.
4. The router#
def moe_layer(x, router_weight, experts, k=2):
# x: (num_tokens, d)
logits = x @ router_weight.T # (num_tokens, E)
probs = softmax(logits, dim=-1)
topk_probs, topk_idx = probs.topk(k, dim=-1) # (num_tokens, k)
topk_probs = topk_probs / topk_probs.sum(-1, keepdim=True) # renormalize
out = torch.zeros_like(x)
for j in range(k):
for e in range(len(experts)):
mask = (topk_idx[:, j] == e)
if mask.any():
out[mask] += topk_probs[mask, j:j+1] * experts[e](x[mask])
return outThat naive loop is exactly what you must not do in production (Section VI.10 — maximal warp divergence, uncoalesced access). The production version sorts tokens by expert and uses a grouped GEMM (Section XIII.03).
Router variants#
TOP-K (standard) each token picks its k best experts
EXPERT CHOICE each EXPERT picks its top tokens
→ perfectly balanced by construction
→ but some tokens may get 0 experts
SHARED EXPERT (DeepSeek) one expert is ALWAYS active, plus k routed ones
→ captures common patterns; improves stability
FINE-GRAINED (DeepSeek) many small experts (256) instead of few large (8)
→ better specialization, harder to balance
AUXILIARY-LOSS-FREE (DeepSeek-V3) bias terms adjusted during training to
balance load without an auxiliary lossThe shared-expert design is worth knowing: DeepSeek keeps one expert always on, which handles the “generic” part of the computation and lets the routed experts specialize. It also means the model degrades gracefully if routing is poor.
5. The inference consequences#
1. MEMORY IS THE CONSTRAINT, NOT COMPUTE
671B at FP8 = 671 GB. Needs 9+ H100s just to hold, before KV cache.
→ MoE models require multi-GPU (and often multi-node) deployment
for capacity reasons, not latency reasons.
2. DECODE READS ONLY THE ACTIVE EXPERTS' WEIGHTS... IN THEORY
At batch 1, one token uses 2 of 256 experts per layer.
Bytes read ≈ (attention weights) + (2/256 of FFN weights)
→ decode SHOULD be as cheap as a small dense model.
IN PRACTICE at higher batch: different tokens in the batch pick
DIFFERENT experts. At batch 64 with k=2 and E=256, you'll touch
most experts.
→ bytes read approaches the FULL model at moderate batch.
→ this is the key asymmetry: MoE's compute advantage holds,
but its memory-traffic advantage erodes with batch size.
3. LOAD IMBALANCE (Section IX.05)
Routing is data-dependent and uneven. The All-to-All is a barrier.
→ the slowest expert determines the step time.
4. ALL-TO-ALL COMMUNICATION
Two per MoE layer, when using expert parallelism.
→ expensive across nodes.Point 2 deserves emphasis because it’s counterintuitive.
Expected experts touched, batch B, top-k, E experts
≈ E × (1 - (1 - k/E)^B)
Mixtral 8x7B (E=8, k=2):
B=1 → 2.0 experts (25% of FFN weights read)
B=4 → 5.4 (67%)
B=16 → 7.9 (99%)
B=64 → 8.0 (100%)
DeepSeek-V3 (E=256, k=8):
B=1 → 8.0 experts (3% of FFN weights)
B=16 → 105 (41%)
B=64 → 217 (85%)
B=256 → 256 (100%)By batch 16-64, you’re reading essentially all the weights. So MoE’s decode advantage is largest at small batch and disappears at large batch — the opposite of most optimizations.
6. The serving profile#
MoE MODELS ARE:
✓ cheap in FLOPs per token (like their active-parameter count)
✗ expensive in memory capacity (like their total-parameter count)
✗ expensive in memory bandwidth at moderate-to-high batch
✓ excellent quality per FLOP
✗ complicated to serve (All-to-All, imbalance, grouped GEMM)
THEY SUIT:
✓ deployments with abundant GPU memory
✓ workloads where quality per FLOP matters
✓ high-batch throughput serving (the FLOP advantage holds)
THEY DON'T SUIT:
✗ memory-constrained deployments
✗ low-batch latency-critical serving on limited hardware
✗ simple single-GPU deploymentsThe honest summary: an MoE with 37B active parameters does NOT serve like a 37B dense model. It serves like something between a 37B and a 671B model, depending on your batch size, and it requires the memory of the 671B.
7. Quantization and MoE#
MoE is unusually amenable to quantization because memory is the constraint:
671B at BF16: 1,342 GB → 17 H100s
671B at FP8: 671 GB → 9 H100s
671B at INT4: 336 GB → 5 H100s
→ quantization changes the deployment shape dramatically.
→ DeepSeek trains and serves in FP8 partly for this reason.
CAVEAT: the ROUTER weights are tiny and disproportionately important.
A wrong routing decision sends a token to the wrong expert entirely.
→ keep the router in higher precision.8. Production implications#
- Compute the memory requirement from TOTAL parameters, not active. This is the number that determines your GPU count.
- Compute the decode bandwidth from expected experts touched at your batch size, not from active parameters.
- MoE strongly favors quantization because memory is the constraint.
- Keep router weights at higher precision.
- Monitor expert load distribution as a first-class metric (Section IX.05).
- Verify your engine uses grouped GEMM for the expert computation, not a per-expert loop.
- Understand that “37B active” is a training-cost statement, not a serving-cost statement.
9. Common mistakes#
Sizing memory from active parameters. Off by 18x.
Assuming MoE decode is as cheap as a dense model of the active size. True at batch 1, false by batch 32.
Naive per-expert loops. Massive warp divergence and uncoalesced access.
Quantizing router weights aggressively. A wrong routing decision is catastrophic.
Ignoring load imbalance. Halves throughput invisibly.
Deploying MoE on memory-constrained hardware because the active parameter count looked manageable.
10. Hands-on exercise#
A. Compute the expert-touching curve. For Mixtral (E=8, k=2) and a DeepSeek-scale model (E=256, k=8), compute expected experts touched vs batch size. Plot it. At what batch does each reach 90%?
B. Measure it. Run an MoE model and instrument the router. Record actual experts touched per batch at batch 1, 8, 32, 128. Compare to the formula.
C. Memory vs bandwidth. For an MoE model, compute: memory required (total params) and bytes read per decode step at batch 1 and batch 64. How do they differ from a dense model of the active size?
D. Load imbalance. Plot the distribution of tokens per expert at various batch sizes. How does the imbalance change? What’s the effective utilization at each?
E. Router precision. Quantize an MoE model with the router at FP16 and at INT4. Measure routing agreement and downstream quality. How sensitive is it?
11. Interview questions#
- What is MoE and what does it decouple?
- Does a “671B total, 37B active” model serve like a 37B model? Explain carefully.
- How many experts are touched at batch B, and why does that matter?
- Why does MoE’s decode advantage shrink with batch size?
- Why is MoE unusually amenable to quantization?
- Why keep router weights at higher precision?
- What’s the wrong way to implement the expert computation?
12. Further reading#
- [ESTABLISHED] Shazeer et al., “Outrageously Large Neural Networks” (2017) — the original
- [ESTABLISHED] Lepikhin et al., “GShard” (2020); Fedus et al., “Switch Transformers” (2021)
- [ESTABLISHED] Jiang et al., “Mixtral of Experts” (2024)
- [ESTABLISHED] DeepSeek-V2 and V3 technical reports — fine-grained experts, shared expert, auxiliary-loss-free balancing
- Next: 03 — MoE serving and expert parallelism