Mixture of Agents
Mixture of Agents asks several different LLMs the same question, lets each read the others' drafts as references, and has one aggregator fuse a better final answer. It is an ensemble at inference time — not Mixture-of-Experts — and while it buys real quality on open-ended tasks, it costs 10x tokens. Its best 2026 role is offline: generating training data.
01.The Problem: One Expert Always Misses Something
You ask an LLM to plan a two-week trip to Japan.
One model nails the itinerary but books you into a restaurant that closed last year. Another knows every train line but ignores your wheelchair access note. A third is great with budgets and wrong about visas.
You don't have one flaw to fix. You have coverage holes — and each model has different ones.
So the question becomes
What if the model could ask its friends before answering?
That is the entire idea of Mixture of Agents (MoA):
Ask several different models the same question; let each see the others' answers as references; have one model fuse everything into a single improved answer.
It is an ensemble at inference time — no training, just orchestration. Different model families were trained on different data, with different tokenizers and different goals, so their mistakes partially cancel: one recalls the constraint, another knows the API, a third avoids the off-by-one. The fused answer keeps the union.
Carry one analogy through: a newsroom. Three reporters draft the same story from different beats. An editor reads all three drafts and writes ONE better story — not a list of the drafts. The quality comes from different angles plus one deliberate editor. The cost is obvious: three reporters and an editor cost more than one journalist. And if all three reporters copy the same wire story (read Section 7 — same distillation lineage), you paid for one story three times.
MoA: Reference, Then Aggregate, Layer by Layer 🕸️
MoA: Reference, Then Aggregate, Layer by Layer 🕸️
Each layer lets every model read the others' answers as references, then an aggregator fuses them. Depth is the accuracy knob; fan-out is the cost knob.
Unlock Topic #265: Mixture of Agents
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?