TOPIC #209Advanced 14 min read

Mixture of Experts (MoE)

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

How sparse Mixture-of-Experts architectures decouple total parameters from active parameters: a learned router sends each token to only the top-k expert FFNs, so capacity grows while per-token compute stays small — the design behind Mixtral, DeepSeek-V3, and Kimi K2.

MoE Layer: Sparse Expert Routing 🔐

Each token activates only the top-k expert FFNs (k typically 1 or 2) chosen by a learned router; the remaining experts are not computed, so FLOPs per token scale with active, not total, parameters.

MoE Layer: Sparse Expert Routing 🔐
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: Knowledge Is Cheap, Compute Per Token Is Not

A normal ("dense") transformer charges you the same price for every word: every parameter fires on every token.

Want a model that knows more? Add parameters. But then every single token pays for the whole enlarged model — inference gets slower and more expensive per word, even when most of that knowledge was irrelevant to the sentence at hand.

So the question becomes

Insight

Can a model STORE a lot of knowledge but only CONSULT a little of it per token?

A hospital already solves this exact problem. It houses dozens of specialists — cardiologists, neurologists, dermatologists — a huge total amount of expertise. But no patient sees all of them. A triage nurse routes each patient to the two or three relevant specialists, and their opinions get combined into one treatment plan.

Mixture of Experts (MoE) is the same architecture for neural networks.

02.The Idea in Plain Words: Many Specialist FFNs, One Switchboard

MoE in one line

Insight

Replace each feed-forward block with N parallel expert FFNs plus a tiny learned router — and compute only the top-k experts per token.

Unpack the pieces:

  • In a transformer, the FFN (feed-forward network, the MLP block after attention) is where most stored knowledge lives. That is what gets multiplied.
  • Each expert is just its own FFN of the same shape.
  • The router (gating network) is a small linear layer that reads the token's hidden state, scores all N experts with a softmax, and picks the best k — k=1 in Switch Transformer, k=2 in Mixtral and GShard-style models.
  • The outputs are merged as a weighted sum using the router's own scores (e.g. 0.62 × Expert3 + 0.38 × Expert7).
  • Everything else — attention, embeddings, the first layer — stays dense. Experts only replace MLP blocks, typically from layer 2 onward.

The payoff is the separation of two numbers everyone used to conflate:

  • Total parameters = capacity, knowledge stored in weights. Grows massively.
  • Active parameters per token = what fires when you generate one word. Stays modest — this is what sets FLOPs and latency.

The scoreboard, straight from production models:

  • Mixtral 8x7B (2024): 47B total parameters, but only ~13B active per token — it outperformed Llama-2 70B at the inference cost of a dense 13B.
  • DeepSeek-V3 (Dec 2024): 671B total, 37B active per token; 256 routed experts per MoE layer, top-8, plus 1 shared expert every token uses regardless of routing.
  • Qwen3-235B-A22B (2025): 128 experts, top-8; ~22B active (the "A22B" literally means "active 22B").
  • Kimi K2 (2025): 1T total parameters, 32B active.

03.A Simple Worked Example: One Token, Eight Experts

Suppose a layer has 8 experts and top-2 routing. Each expert is 5B parameters; attention and everything else totals 7B.

The token "photosynthesis" arrives. The router scores all 8 experts:

code
 expert:   1     2     3     4     5     6     7     8
 score:  0.02  0.05  0.62  0.01  0.02  0.04  0.23  0.01
                           ^top-1        ^top-2

Only Expert 3 and Expert 7 are computed:

  • output = 0.62 × Expert3(x) + 0.23 × Expert7(x) (scores renormalized over the chosen k, in practice).
  • The other 6 experts — 30B of weights — sit idle for this token.

Count the bill:

code
 active per token: 7B (dense parts) + 2 experts × 5B = 17B
 total resident:   7B + 8 experts × 5B = 47B
 → 47B of knowledge, 17B of compute per token

Different tokens pick different pairs — biology tokens visit Expert 3, code tokens visit Expert 0, and so on. Specialization is not programmed; the router learns it from data.

04.Visual Intuition: The Hub-and-Spoke Layer

code
                    ┌─ Expert 1   (skipped: score ~0)
                    ├─ Expert 2   (skipped)
   token ──► Router │─ Expert 3 ★ 0.62 ─┐
  hidden state     ├─ Expert 4   (skip) ├─► weighted sum ─► + residual ─► next layer
   (softmax        ├─ ...               └─ Expert 7 ★ 0.23 ─┘
    over N)        └─ Expert 8   (skipped)

And across a whole model, experts specialize, which you can picture as a directory the router learns:

code
   "quantum"  → decoherence/physics experts
   "def f(x):" → code experts
   "Bonjour"   → French experts

The key asymmetry to hold in your head: compute is sparse, memory is not. Skipped experts cost no FLOPs but still sit in memory. Section 6 is where that bill arrives.

05.The Analogy: The Hospital With a Very Fast Triage Nurse

Carry the hospital through everything below.

  • Experts = specialists' offices. Each gets better at a subset of cases because it only sees its subset (the routing gradient trains them apart).
  • The router = the triage nurse: cheap, quick, sees every patient, points at two doors.
  • The weighted sum = the two specialists' opinions merged into one chart note, weighted by how clearly the nurse was sure.
  • A shared expert (DeepSeek's design) = the general practitioner every patient sees regardless — general medical knowledge never has to be re-learned by each specialist.
  • Expert collapse = the nurse starts sending everyone to the same two doctors because they look busiest; the other specialists' rooms gather dust.
  • Load balancing = the hospital administrator nudging the nurse: keep all doors in use, or idle doctors atrophy (a network with no routed tokens gets no gradient — a dead expert).
  • Aux-loss-free bias = instead of scolding the nurse (a loss term that distorts training), the administrator quietly adjusts the door signs — a bias added only at selection time — so traffic spreads out without changing how anyone is graded.

Everything MoE does — top-k, biases, capacity factors — is hospital logistics around one fact: the more patients a specialist actually sees, the better (and more routable) they become, and routing itself changes who gets seen.

06.Routing Is Fragile: Collapse, Dead Experts, and the Fixes

The router computes one logit per expert (a linear projection of the hidden state), applies softmax, and takes the top-k. Two classic failure modes must be engineered away:

  1. Expert collapse / rich-get-richer: without pressure, a few experts capture all tokens; their gradients keep improving them, they attract even more tokens, and the rest starve. Countermeasures:
    • the load-balancing auxiliary loss (GShard / Shoeybi et al.): an extra term in the total loss that penalizes unequal token distribution across experts;
    • z-loss (ST-MoE): penalizes the squared magnitude of router logits, keeping them from blowing up and making top-k selection unstable;
    • per-token capacity factors: each expert only accepts a quota of tokens; overflow tokens get dropped or rerouted, which bounds the imbalance directly.
  2. Dead experts: an expert never selected receives no gradient, so its weights freeze forever — capacity paid for, never used.

DeepSeek-V3 removed the balancing loss entirely with auxiliary-loss-free balanced routing: each expert keeps a scalar bias (the administrator changing the door sign) updated from observed load, added to router logits only for selection, not for the mixing weights. Overloaded experts get negative bias, balancing throughput without distorting training gradients — a signature 2024-2026 refinement.

The architectural trend is fine-grained experts: instead of 8-64 huge FFNs, use hundreds of small experts and activate more of them (DeepSeek-V2/V3's 256 top-8 + shared expert; Gemma and OLM-oE smaller expert sets). More, smaller experts give better specialization and smoother load balancing per unit of active compute.

07.Systems Reality: The Bill Arrives as Communication and Memory

MoE shifts the bottleneck from FLOPs to communication and memory. Because experts live on different GPUs, every token's "which expert, where?" question becomes a network hop.

Training and serving large MoE models use Expert Parallelism (EP): experts are sharded across GPUs/nodes, and every MoE layer requires an all-to-all dispatch (send each token to the GPUs owning its chosen experts) and a second all-to-all combine (return the expert outputs).

What that means at scale:

  • DeepSeek-V3 trained on 2,048 H800 GPUs using EP with grouped GEMM kernels (batching many small expert matmuls into few big ones), overlapped communication, and DualPipe scheduling that cuts kernel counts — making a 671B MoE trainable for ~$5.6M of H800-hours, roughly an order of magnitude below dense-scaling estimates at that capability.
  • Serving memory: all 671B weights must be resident in HBM or fast tiered memory even though only 37B are active. Kimi K2 (1T) needs ~2TB in BF16, ~1TB in FP8; single-node inference became practical only with aggressive quantization (e.g., Marlin kernels on 8xH200 or large unified-memory machines).
  • Batch effects: with large batches, most experts get touched by some token in the batch, so effective weight reuse is high and MoE serving throughput is bandwidth-bound rather than compute-bound. A batch of one, however, mostly idles the resident weights.
python— Top-2 sparse MoE layer in PyTorch (simplified Mixtral-style FFN)
class MoELayer(nn.Module):
    def __init__(self, d_model, expert_dim, n_experts=8, top_k=2):
        super().__init__()
        self.router = nn.Linear(d_model, n_experts, bias=False)
        self.experts = nn.ModuleList([Expert(d_model, expert_dim)
                                      for _ in range(n_experts)])
        self.top_k = top_k

    def forward(self, x):                      # x: (tokens, d_model)
        scores = self.router(x).softmax(-1)    # (tokens, n_experts)
        weights, idx = scores.topk(self.top_k, dim=-1)
        out = torch.zeros_like(x)
        for k in range(self.top_k):
            for e in range(self.experts.__len__()):
                mask = idx[:, k] == e          # tokens routed to expert e
                if mask.any():
                    out[mask] += weights[mask, k, None] * \
                        self.experts[e](x[mask])
        return out

08.MoE vs Dense in 2024-2026

Frontier practice effectively converged on MoE: DeepSeek-V3/R1, Kimi K2, Qwen3, Llama 4 (Scout: 16 experts; Behemoth rumored), gpt-4-class models widely believed to be MoE, and most open-weights leaders are sparse.

The core argument is token-efficiency: for a fixed training-FLOPs budget, a well-balanced MoE reaches lower loss than the best dense model of equivalent active size. Capacity at a discount is hard to refuse when you are training at the frontier.

Dense models still win on simplicity: deterministic memory footprint, no all-to-all, uniform latency, easy quantization and edge deployment.

A practical rule of thumb for 2026: MoE is a datacenter-scale play — it pays off when you have ≥ hundreds of GB of bandwidth-rich memory and batched traffic. If you need single-stream latency on one consumer GPU, dense (or a small dense distill) is usually the saner deploy.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Total capacity scales near-linearly while per-token FLOPs stay constant (active-parameter pricing).
  • Better loss-per-training-FLOP than dense models at frontier scale — how open models hit GPT-4-class quality cheaply.
  • Per-request compute is low: Mixtral-style serving charges you for 13B, not 47B.

Trade-offs & Constraints

  • Full parameter set must be resident in memory/weight tier: huge weight footprint and cold-cache latency if offloaded.
  • All-to-all expert-parallel communication dominates cost at scale; complex kernels and load-balancing engineering.
  • Router imbalance creates tail-latency and underutilized experts; quantizing routers/experts asymmetrically is subtle.
  • Batch-dependent throughput: single-stream MoE (1 user, 1 token) wastes inactive experts' memory bandwidth.
Production Implementation in Big Tech
DeepSeek• Frontier model trained on 2,048 H800s

DeepSeek-V3 (Dec 2024) uses a fine-grained MoE design — 256 routed experts per layer with top-8 activation plus one shared expert, auxiliary-loss-free load balancing, and Multi-head Latent Attention. Reported total training cost ~$5.6M, with per-token compute of only 37B active parameters despite 671B total weights, matching or beating far costlier dense frontier models on 2024-2025 benchmarks.

Staff+ Engineering Takeaways

  • MoE replaces FFNs with many expert FFNs plus a learned router; only top-k experts run per token, so active parameters ≪ total parameters.
  • Load balancing (aux loss, z-loss, or DeepSeek's aux-loss-free bias) and fine-grained experts are what make modern MoE actually train.
  • Systems costs move to communication and memory: expert parallelism is all-to-all heavy; full weights must stay resident for serving.
  • 2024-2026 frontier practice is MoE: Mixtral 8x7B, DeepSeek-V3 (671B/37B), Qwen3-235B-A22B, Kimi K2 (1T/32B).
  • Dense remains simpler and better for edge/low-batch serving; MoE wins at scale and batched throughput.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

A 47B-parameter MoE like Mixtral 8x7B runs 8 experts per layer and routes each token to the top 2. Approximately how many parameters are active (computed) per token?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Related Concepts & Cross-References

Indexed from curriculum