TOPIC #125Advanced 14 min read

Feed-Forward Sublayer: The Per-Token Digestion Step

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

After attention mixes tokens, a two-layer MLP processes each position alone — and it holds roughly two-thirds of every Transformer's parameters. What the FFN actually computes, why it grew from 4× ReLU to SwiGLU and MoE, and what 2024-2026 research thinks it stores.

01.The Problem: Attention Makes Smoothies; Nobody Tastes Them

Rattle through what attention actually outputs (Topic 119-120):

Insight

A weighted average of value vectors.

Averages are blending. Blending is (mostly) linear bookkeeping.

Take the token "it" after attention: its new vector is 0.61·v(street) + 0.22·v(tired) + …. A mixture.

But now the model must do things a mixture cannot do by itself:

  • Decide this combination means "an exhausted referent, likely a mouse."
  • Apply rules: "after a singular subject, conjugate the verb thus."
  • Recall stored facts: "Paris is the capital of France."

Two gaps:

  1. Nonlinearity gap: averaging gives mixtures; nothing yet interprets the mixture per token. You need a function that can say "if pattern X fires, write effect Y; if not, stay silent."
  2. Capacity gap: attention has few parameters (four d×d matrices per block). Where does a model store the millions of facts and skills it knows? A mixer is not a warehouse.

The 2017 fix slots between every attention sublayer and the next:

Insight

A simple two-layer MLP, run alone on each token, wide in the middle — the "digestion" step.

It looks too modest to matter. It holds ~2/3 of the parameters and, per 2021-2026 evidence, most of the knowledge. This topic shows why.

The MLP Slot: Detect a Pattern, Then Write Its Effect ✍️

PRO Architecture Blueprint

The MLP Slot: Detect a Pattern, Then Write Its Effect ✍️

Attention moves tokens around; the FFN transforms each one alone, twice through a wide bottleneck. Its 4× expansion is where most of a model’s parameters — and, evidence says, most of its "knowledge" — live.

The MLP Slot: Detect a Pattern, Then Write Its Effect ✍️
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #125: Feed-Forward Sublayer: The Per-Token Digestion Step

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Related Concepts & Cross-References

Indexed from curriculum