Multi-Layer Perceptrons: Depth, Composition, and Universal Approximation
One perceptron draws one straight line and gets stuck on XOR (topic 75). Stack perceptrons into layers with nonlinearities between them, and each layer re-draws the map until the classes fall apart along a straight cut. This page shows the composition math, solves XOR with two hidden units, and explains what the universal approximation theorem does — and does not — promise.
A Two-Hidden-Layer MLP
Every neuron in a layer connects to every neuron in the next (dense/full-connected). Hidden units are intermediate representations — new coordinates learned from data.
01.The Problem: One Line Was Not Enough
Topic 75 in one plain sentence: a perceptron multiplies inputs by weights, adds a bias, and answers yes/no — so it can only ever draw one straight line through its inputs.
That line solves spam, pass/fail, fraud.
But remember the disaster at the end of topic 75:
XOR — "are these two bits different?" — puts its positive points on diagonally opposite corners of a square, and no single straight line can separate them.
A single perceptron fails XOR no matter how long you train it, because the right weights simply do not exist.
So the question becomes
If one line is not enough... what about a team of lines that get to work in stages?
That team is the Multi-Layer Perceptron (MLP) — also called a fully-connected or dense network.
It is the architecture that:
- broke the XOR ceiling in 1986,
- became the backbone of classical deep learning, and
- still hides inside every single layer of every transformer in every LLM you have ever used.
Understanding the MLP is understanding the "thinking with stacked layers" idea that the entire field runs on.
02.The Idea in Plain Words: Layers That Re-Draw the Map
An MLP is simply
Perceptrons arranged in layers, where each layer's outputs become the next layer's inputs — with a nonlinear twist squeezed between every pair of layers.
Written as math:
codeh1 = φ1(W1·x + b1) # hidden layer 1 h2 = φ2(W2·h1 + b2) # hidden layer 2 ŷ = W3·h2 + b3 # output layer
Unpack the symbols:
W1, W2, W3are weight matrices — just many perceptrons' weights stored side by side. Each column is one unit's row of the score sheet.b1, b2, b3are the bias vectors — one grumpiness number per unit.φ(phi) is a nonlinear activation — a tiny per-number twist like ReLU'smax(0, z).h1, h2are the hidden layers: units that never talk to the outside world. They exist purely to invent new coordinates for the layers above.
The crucial trick, in one sentence:
Hidden units perform a learned change of coordinates — they do not answer the question, they restate the problem in a space where the answer becomes obvious.
The final layer is then allowed to be a plain linear readout... but only because the hidden layers moved the data into a friendlier space.
03.A Simple Worked Example: MLP Solves XOR
XOR's four inputs: (0,0) → 0, (1,1) → 0, (1,0) → 1, (0,1) → 1. One line cannot split them. Let's do it with two hidden ReLU units and one output unit — real numbers, no hand-waving.
Hidden unit 1 asks: "is the total above 1?" → h1 = max(0, x1 + x2 − 1)
Hidden unit 2 asks: "is the total below 2?" → h2 = max(0, 2 − x1 − x2)
Compute both for all four corners:
code(x1,x2) h1 = max(0, x1+x2−1) h2 = max(0, 2−x1−x2) (h1,h2) (0,0) max(0, −1) = 0 max(0, 2) = 2 (0, 2) (1,1) max(0, 1) = 1 max(0, 0) = 0 (1, 0) (1,0) max(0, 0) = 0 max(0, 1) = 1 (0, 1) (0,1) max(0, 0) = 0 max(0, 1) = 1 (0, 1)
Look what just happened. The two negative corners became (0,2) and (1,0). The two positive corners became the same point (0,1).
In the original map, the classes were diagonal and inseparable. In the new (h1,h2) map, the negatives both have h1 + h2 = 2 while the positives have h1 + h2 = 1.
So one line in the new space finishes the job:
codeŷ = 1 if −h1 − h2 + 1.5 ≥ 0 (output weights (−1, −1), bias 1.5) check: negatives → −2 + 1.5 = −0.5 → ŷ = 0 ✓ check: positives → −1 + 1.5 = +0.5 → ŷ = 1 ✓
The output layer is a plain linear perceptron — just trained on the coordinates the hidden layer invented.
That is the whole recipe: the hidden layer bends the space; the last layer cuts straight.
04.Visual Intuition: Folding Paper Until One Snip Separates Everything
Picture the input plane as a sheet of paper with colored dots on it — XOR's positives red, negatives blue, one color on each diagonal.
codeBEFORE (input space) AFTER (hidden layer's space) ⊗ · · · · ○ ○ = positive · ╲ ╱ · fold, ● ● ● ● ● ← a single horizontal · ╳ · pull, · · · · · · SCRAPE now separates · ╱ ╲ · crease... ● NEGATIVES everything ○ · · · · ⊗ (all pushed up) no line works positives collapsed together — one straight cut works
Each hidden unit is one crease: w·x + b = 0 is a fold line, and ReLU decides which side of the crease stays flat and which side gets pressed to zero.
A layer of k units makes k creases at once. Stack layers and you fold the paper again on top of itself, creasing the creases.
Deeper means more folds per square inch — and the final linear readout is just one scissor-snip through the folded stack.
The rigorous version of this picture: geometrically, a hidden layer carves input space into polytopes — each ReLU unit contributes a half-space (one side of a fold line), and the network traces paths through intersections of them.
With enough hidden units, an MLP partitions the input space into an arbitrary number of convex polytopes and assigns each a (piecewise-linear) output value.
And depth is not cosmetic: deeper networks can produce an exponential number of linear regions relative to width in some architectures — Montufar et al. (2014) showed piecewise-linear networks with depth d and fixed width can realize exponentially more regions than shallow ones.
That is the modern, rigorous version of "depth helps expressivity": not about whether a function can be represented, but how compactly.
05.The Warning Every Beginner Misses: No Nonlinearity, No Network
Here is the one algebraic fact the entire field of deep learning stands on.
Feed an input through two layers with no activation between them:
codeŷ = W2·(W1·x + b1) + b2 = (W2·W1)·x + (W2·b1 + b2) = W'·x + b'
Two affine maps in sequence collapse into one (W2(W1x) = (W2W1)x).
One hundred such layers? Still one matrix. Five hundred layers of pure matrix multiplies = one expensive linear model that took 500x the time to compute.
So without nonlinearities, depth is useless.
The activation functions between layers are what make composition strictly more expressive.
That one fact is the entire reason deep learning exists as a separate discipline from linear algebra.
Which also previews the punchline: the interesting design questions become
- which nonlinearity to insert (topic 77),
- where the gradient dies during training (topics 85-86),
- and how many folds you can afford.
Sanity check with tiny numbers: y = 3·(2x) + 1 = 6x + 1. Two "layers" (×2, then ×3), one line. You cannot fold 6x + 1 into anything but another line. The max(0, ·) twist is what lets a crease exist at all.
06.The Analogy: A Kitchen Brigade — Prep Cooks, Then the Chef
Carry one analogy through the rest of this topic: a restaurant kitchen brigade.
- The raw ingredients arrive unorganized: whole fish, dirt-covered carrots. That is your raw input vector — pixels, tokens, features.
- The prep cooks are the hidden layers. They do not plate anything. They perform transformations: fillet, dice, reduce. Each cook takes what the previous station produced and restates it in a form the next station can use.
- The head chef is the output layer. She only makes the final call — "this plate is ready / not ready" — and she can make it easily precisely because the prep cooks already did the hard restructuring.
Now watch each technical claim map onto the kitchen:
- "Hidden layers are learned changes of coordinates" → the prep station invents its own cuts; nobody told the cook that dice-5mm beats dice-8mm, the recipes (training data) forced that skill.
- "Linear layers without activation collapse" → ten workers who each only move boxes to the left counter accomplish what one worker could. No transformation, no depth benefit.
- "Depth buys exponential regions" → each extra station can combine previous prep in ways no single cook could reach: julienne-then-blanche-then-fold.
- "MLPs memorize small datasets" → a kitchen with one recipe card and 50 cooks starts inventing reasons to be busy — overfitting is organizational, not personal.
And the punchline connection: in every transformer, the MLP block after attention is exactly this prep station — attention decides which ingredients to gather, the MLP chops them into new representations the next layer consumes.
07.Choosing Shape: Width, Depth, and Parameters
An MLP is a rectangle factory, and every rectangle costs exactly:
n_in × n_outweights plusn_outbiases.
Count the parameters of a 784→256→128→10 MLP on MNIST (784 = 28×28 flattened pixels, 10 = digit classes):
codelayer 1: 784×256 + 256 = 200,960 + 256 = 201,216 layer 2: 256×128 + 128 = 32,768 + 128 = 32,896 layer 3: 128×10 + 10 = 1,280 + 10 = 1,290 total ≈ 235,402 → "about 237k parameters" as usually quoted
One number to internalize: a dense layer's cost is the product of fan-in and fan-out. That is why fully-connected layers are avoided on raw images (see topic 84 for how convolutions exploit structure instead).
Practical shape heuristics in the transformer era:
- Bottlenecks (narrow middle layers) force compression and act as regularizers — the kitchen can't fake a 5mm dice if the basket only fits coarse chops.
- Residual MLP blocks (
x + MLP(x)) — the feedforward blocks inside every transformer — let depth help via identity shortcuts, easing optimization: if a station adds nothing useful, the ingredients can simply walk past it unchanged. - Tabular data: gradient-boosted trees still beat deep MLPs on many mid-size tabular benchmarks; MLPs become competitive at 1M+ rows or when features include high-cardinality categoricals needing embeddings (this is the recipe behind two-tower recommenders).
- Width vs depth: for a fixed parameter budget, deeper-narrower usually folds more regions (Montufar 2014, section 4), but deeper also means more places for gradients to attenuate — the optimizer tax on expressivity.
The PyTorch version of the whole architecture is eight lines:
import torch, torch.nn as nn
model = nn.Sequential(
nn.Flatten(),
nn.Linear(28 * 28, 256), nn.ReLU(), nn.Dropout(0.2),
nn.Linear(256, 128), nn.ReLU(),
nn.Linear(128, 10), # raw logits
)
x = torch.randn(64, 1, 28, 28) # one batch
logits = model(x) # (64, 10)
print(sum(p.numel() for p in model.parameters())) # ≈237k params08.What Hidden Layers Actually Learn — and Why AI Cares
Hidden units are not features you designed — they are learned representations. Nobody codes "detect a vertical stroke"; the loss function forces some unit into that job.
Classic evidence:
- On MNIST, first-layer ReLU units respond to oriented edges; deeper units respond to strokes and digit parts. Layer 1 = "lines", layer 3 = "curves", layer 5 = "loops like in a 6 or an 8" — the classic computer-vision demo where you can literally draw the filters.
- On language, MLP units inside transformers fire for syntactic or semantic patterns — parts of speech, phrase endings, factual associations.
This gives the modern mental model, one line per layer type:
coderaw data → edges/fields → parts/patterns → concepts → linear readout input layer 1 layer 2 layer k output
Each hidden layer is progressively more abstract, more invariant, more task-shaped than the last.
Why every AI engineer should care, concretely:
- On-ramp to LLM internals. The feedforward blocks of GPT-class models are MLPs (with GELU instead of ReLU) sandwiched between attention layers — each transformer block's MLP expands the hidden dimension roughly 4x. If you understand this page, you understand half of a transformer layer already; topic 90 finishes the job.
- Why "representation learning" beat feature engineering. Pre-deep-learning vision pipelines had humans hand-designing SIFT/HOG features. An MLP learns its own creases — better and cheaper at scale.
- Why structure matters as much as size. Dense layers treat every input pair as equally related, burning
n_in × n_outparameters on every pair. CNNs (topic 84) and transformers (topic 89+) win on images/text by assuming locality/sequence; MLPs win once data is already a vector: tabular rows, embeddings, classifier heads on frozen encoders. - Where training still surprises you. The architecture can represent the function (universal approximation) yet gradient descent still fails to find it — bad init, saturation, thin data. Expressivity and trainability are two different books; this topic wrote the first chapter, topics 85-86 write the second.
An MLP is a map folder. Its power is not any single unit — it is that layers of small, local creases compose into global restructurings no hand-designed feature could match.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Breaks the linear-separability ceiling: can approximate any continuous function given enough width.
- Learns its own features — no manual engineering of interactions.
- Uniform structure (matrix multiply + elementwise nonlinearity) maps perfectly onto GPUs and TPUs.
- Inference is a few dense matmuls: fast, simple to deploy, and easy to quantize.
Trade-offs & Constraints
- Parameter count grows as n_in × n_out — dense layers are wasteful for spatial/temporal structure (CNNs and transformers exploit that structure instead).
- Universal approximation gives no training guarantees; optimization can still fail.
- Prone to memorizing small datasets without regularization (dropout, weight decay).
In GPT-class models, each transformer block applies attention over tokens and then a dense MLP that expands the hidden dimension by roughly 4x (e.g., 8k → 35k in large models) with a GELU-like activation. Those MLP layers act as the model's key-value memory for factual associations — mechanistic-interpretability work from 2024-2025 (e.g., Geva et al., "Transformer Feed-Forward Layers Are Key-Value Memories") shows editing them changes stored facts.
Staff+ Engineering Takeaways
- MLPs compose affine maps with nonlinearities; without activations, stacked linear layers collapse into one (W2(W1x) = (W2W1)x).
- Hidden layers are learned coordinate changes that make non-separable classes (XOR) separable — the output layer then cuts with one straight line.
- Universal approximation guarantees representation power, not trainability, compactness, or extrapolation.
- Depth buys exponentially more linear regions than width alone for piecewise-linear networks (Montufar et al., 2014).
- A dense layer costs n_in × n_out weights + n_out biases — products, not sums, which is why raw high-dimensional inputs need structured architectures.
- Modern relevance: the MLP block inside every transformer layer is the network part that stores factual associations.
Topic Knowledge Check
Exercise 1 of 4 • Test your architectural comprehension.
You stack 20 linear layers into one network. Why must nonlinear activation functions still sit between the layers?
How clear and actionable was this distributed systems breakdown?