Feed-Forward Sublayer: The Per-Token Digestion Step
After attention mixes tokens, a two-layer MLP processes each position alone — and it holds roughly two-thirds of every Transformer's parameters. What the FFN actually computes, why it grew from 4× ReLU to SwiGLU and MoE, and what 2024-2026 research thinks it stores.
01.The Problem: Attention Makes Smoothies; Nobody Tastes Them
Rattle through what attention actually outputs (Topic 119-120):
A weighted average of value vectors.
Averages are blending. Blending is (mostly) linear bookkeeping.
Take the token "it" after attention: its new vector is 0.61·v(street) + 0.22·v(tired) + …. A mixture.
But now the model must do things a mixture cannot do by itself:
- Decide this combination means "an exhausted referent, likely a mouse."
- Apply rules: "after a singular subject, conjugate the verb thus."
- Recall stored facts: "Paris is the capital of France."
Two gaps:
- Nonlinearity gap: averaging gives mixtures; nothing yet interprets the mixture per token. You need a function that can say "if pattern X fires, write effect Y; if not, stay silent."
- Capacity gap: attention has few parameters (four d×d matrices per block). Where does a model store the millions of facts and skills it knows? A mixer is not a warehouse.
The 2017 fix slots between every attention sublayer and the next:
A simple two-layer MLP, run alone on each token, wide in the middle — the "digestion" step.
It looks too modest to matter. It holds ~2/3 of the parameters and, per 2021-2026 evidence, most of the knowledge. This topic shows why.
The MLP Slot: Detect a Pattern, Then Write Its Effect ✍️
The MLP Slot: Detect a Pattern, Then Write Its Effect ✍️
Attention moves tokens around; the FFN transforms each one alone, twice through a wide bottleneck. Its 4× expansion is where most of a model’s parameters — and, evidence says, most of its "knowledge" — live.
Unlock Topic #125: Feed-Forward Sublayer: The Per-Token Digestion Step
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?