TOPIC #77Intermediate 12 min read

Activation Functions: The Nonlinearity Engine

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Between every matrix multiply in a neural network sits a tiny per-number rule called an activation. That rule decides whether your "deep" network is a universal function approximator or an expensive straight line. This page builds the selection criteria (saturation, zero-centering, sparsity, smoothness), tours the family from sigmoid to SwiGLU, and names the defaults used in 2026 production models.

Activation Selection Map

Hidden layers default to the ReLU family (GELU/SiLU in transformers); output activations are dictated by the task and loss.

Activation Selection Map
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: A 50-Layer Network That Is Secretly One Layer

Topic 76 in one plain sentence: an MLP stacks layers where each layer is a matrix multiply plus a bias — W·x + b.

Now make a bet with yourself.

Take a network with 50 layers of matrix multiplies and no activation functions between them.

How expressive is it?

Multiply it out for two layers:

code
y = W2·(W1·x + b1) + b2
  = (W2·W1)·x + (W2·b1 + b2)
  = W'·x + b'      ← one affine map. ONE.

Run it with real tiny numbers:

code
layer 1: y = 3·x        (W1 = 3)
layer 2: z = 2·y        (W2 = 2)
together: z = 2·(3·x) = 6·x   ← a "two-layer" network that is one multiply

Fifty layers collapse the same way. An entire GPU cluster grinding through 50 matmuls produces... a straight-line transformation, computed 50 times slower than the one matrix multiply it is equivalent to.

So the question becomes

Insight

Where must something nonlinear enter the stack — and what should that nonlinearity actually look like?

The answer is the subject of this whole topic: the activation function φ, applied element-by-element to the scores between layers.

Insight

Activations are the creases in the paper (topic 76). No creases, no folds, no depth, no deep learning.

02.The Idea in Plain Words: A Tiny Rule on Every Number

An activation function is simply

Insight

A small formula applied to each number individually, squeezing it into a new shape before it is passed on.

Concretely, after a layer computes its scores z = W·x + b (the "pre-activations"), the network applies φ to every entry separately:

code
z = ( 2.0, −5.0, 0.3, −0.1 )
if φ = ReLU:      ( 2.0,  0.0, 0.3,  0.0 )   ← negatives zeroed
if φ = sigmoid:   ( 0.88, 0.007, 0.57, 0.48 ) ← everything squashed to (0,1)
if φ = tanh:      ( 0.96, −1.0, 0.29, −0.10 ) ← squashed to (−1,1)

Same input, four very different networks — and very different training behavior.

The elementwise detail matters: φ never mixes numbers across units. Mixing is the matrix's job. The activation's job is to break linearity at every single wire.

Why so much rides on this one choice? Because every interesting capability of neural networks comes from interleaving affine maps with these nonlinear scalar functions:

  • curved decision boundaries (topic 75's single line becomes any shape)
  • hierarchical features (edges → strokes → digits)
  • universal approximation (topic 76: existence, not trainability)

Think of it this way:

code
matrix layer:  "combine all inputs into new scores"   ← linear, global
activation:    "reshape each score locally"           ← nonlinear, per-number
repeat         ← folds on folds: representation learning

Now the real engineering question: there are dozens of candidate φ's. How do you judge one?

03.The Four Criteria Engineers Actually Check

Every activation choice is judged on the same scorecard of four practical properties.

1. Non-saturation — does φ'(z) stay healthy for large |z|?

When z gets big, some functions flatten out: their slope goes to ~0.

Insight

A unit with slope 0 is a unit that stopped learning: gradients multiplied by ~0 disappear.

Sigmoid saturates hard — gradient max 0.25, vanishes for |z| > 5. ReLU's positive side never saturates. This is criterion #1 because it caps depth (section 7).

2. Zero-centered output — do activations hover around 0 or ride at +0.5?

If most activations share a large positive mean (ReLU: mean ≈ 0.5 for symmetric inputs), gradients bias all weights in one direction, causing zig-zag updates — see the 2015 "Gradient Descent, ReLU Zombies" analysis.

Tiny numeric demo of the zig-zag: if every input to a weight is positive, the gradient of that weight is δ·a with a > 0 always — so the weight can only move with or against the shared error signal δ, never independently. Downstream weights march in lockstep and the loss trajectory sawtooths.

3. Sparsity — does the unit stay quiet most of the time?

Outputs near zero for many inputs reduce interference between patterns (ELU family; k-WTA). A quiet network is like a well-run library: fewer voices talking over each other.

4. Smoothness / monotonicity — is the slope well-behaved?

Differentiable everywhere helps optimizers (kinks like ReLU's at 0 are tolerated but not loved); monotonicity can guarantee convex loss surfaces in single-layer cases.

Scorecard preview (details in the family tour):

code
φ         saturates?  zero-centered?  sparse?  smooth?
sigmoid   YES (both)  no (all >0)     no       yes
tanh      YES (both)  yes             no       yes
ReLU      one side    no (all ≥0)     ~50%     kink at 0
GELU/SiLU mild        no              mild     yes
Insight

No candidate is perfect on all four. Activation design is a trade-off pick, not a lookup.

04.A Simple Worked Example: Watch a Gradient Die

Let's feel saturation with three layers of sigmoid and actual numbers.

Recall (topic 79 has the full story): σ'(z) = σ(z)(1−σ(z)), and its maximum is 0.25, at z = 0.

Backprop through a deep chain multiplies one local derivative per layer. Suppose each of our 3 layers happens to sit at its best possible slope:

code
raw gradient at output:          1.000
after layer 3 (×0.25):           0.250
after layer 2 (×0.25):           0.0625
after layer 1 (×0.25):           0.0156

Already down to 1.5% of the signal — at the optimistic slopes. Now make it realistic: trained units often sit away from 0, where the slope is far smaller.

code
σ'(2)  ≈ 0.105        σ'(5)  ≈ 0.0066
σ'(10) ≈ 0.000045     σ'(20) ≈ 2×10⁻⁹

Three layers at |z| ≈ 5: multiply 0.0066 × 0.0066 × 0.0066 ≈ 3×10⁻⁷.

Insight

Early layers receive a gradient that is seven orders of magnitude smaller than the output's. They effectively never update. The network trains its last layers and pretends the first ones exist.

Now the same 3-layer audit with ReLU, assuming units are active (the common case):

φ'(z) on positive side = 1   →   1 × 1 × 1 = 1

The gradient arrives at layer 1 full strength. That single contrast — 0.0066³ versus 1³ — is why activation choice decided the fate of deep learning between 1990 and 2012.

Run it yourself in two lines:

code
python
import math
sig = lambda z: 1 / (1 + math.exp(-z))
for z in [0, 2, 5, 10]:
    print(z, round(sig(z) * (1 - sig(z)), 6))   # 0.25 → 0.105 → 0.0066 → 0.000045

Watch the slope crater as z drifts from 0. That crater IS the vanishing gradient.

05.Visual Intuition: Flat Hills, Steep Hills, and a Kink

Sketch each φ and its slope tells you everything. The steeper the curve at a point, the bigger the gradient that passes through.

code
sigmoid σ(z)                 ReLU max(0,z)            tanh(z)
 output                        output                   output
 1 ┤            ──── flat     7 ┤          ╱  keeps    1 ┤         ────
   │         ╭── going →0      │       ╱   climbing    │      ╭─── flat
 ½ ┤      ╭─╯ slope ≤0.25      │    ╱                  0 ┤─╱────        →0
   │   ╭──╯ at center          │ ╱ ↕ kink at 0           │╱  ╲
 0 ┤───╯                       └──────► z              −1 ┤    ╲─── flat
            ╲ flat →0                       slope 1 right, 0 left      →0

Read the drawings:

  • Sigmoid/tanh: two flat shoulders. Enter a shoulder, and φ' ≈ 0 — the unit becomes a broken copy machine: it reproduces "1" or "0" no matter what changes upstream.
  • ReLU: a V opening to the right. The right side is a perfect 45° mirror — gradient passes at full strength. The left side is a wall — gradient is exactly zero (both the blessing of sparsity and the curse of dying units; topic 78 is all about ReLU).
  • GELU/SiLU: like ReLU but with the kink sanded smooth — nearly the same slope behavior, friendlier to optimizers.

Rule of thumb from the sketches: flat region = gradient graveyard. Unbounded rising side = gradient highway.

06.The Analogy: A Row of Intercom Boxes (Carried Throughout)

Picture a deep network as a row of intercom boxes down a hallway. Each box listens to the previous box's voice and repeats what it hears into the next box. Training is the engineer standing at the far end shouting corrections backwards down the row (that's backpropagation).

An activation function is the rule each box applies to the voice:

  • Sigmoid box: a polite volume-normalizer — it squeezes any shout into a whisper-to-mumble range (0 to 1). Shout 5 or shout 50, it says the same "≈1". Down the hall, 20 such boxes make the backward-shouted correction inaudible: saturation = the vanishing gradient.
  • Tanh box: same politeness, but centered — quiet is "0", not "0.5" — so corrections don't always sound like "turn everything up".
  • ReLU box: an honest relay for anything above the noise gate; for anything below, stone silence, no slope at all. Cheap hardware (a comparator, no exponential), full-strength echo one way. Sometimes the gate shuts permanently on a speaker — the dying unit (topic 78).
  • Leaky/ELU box: still quiet below the gate, but a small trickle passes so the speaker is never fully dead.
  • GELU/SiLU box: a ReLU box with a soft edge — instead of hard "shout or silence", "shout proportionally to how loud you clearly are".
  • SwiGLU box: two microphones on one speaker — one says "how loud?", the other says "is it real?", multiplied. The current LLM favorite (section 8).

Every criterion from section 3 is a hallway property: saturation = boxes that stop echoing; zero-centering = whether idle boxes hum "0.5" and make the engineer bias every correction; sparsity = how many boxes stay quiet at once.

Keep the hallway in mind while you read the family tree next — each entry is just a different box policy.

07.The Family Tree at a Glance

Historical tour, chronology as storyline (the sigmoid/tanh details get a full topic in 79; ReLU in 78):

  • Step / sign (1958): hard classifier, no gradient — historical only (topic 75's perceptron).
  • Sigmoid σ(z) = 1/(1+e^-z) (used 1986-2010): outputs (0,1), probability-like, but saturates: gradient max 0.25, vanishes for |z| > 5.
  • Tanh tanh(z) = 2σ(2z) − 1: zero-centered, gradient up to 1.0 — better than sigmoid, still saturating.
  • ReLU max(0, z) (Nair & Hinton 2010): unbounded on the positive side, cheap, sparse, but suffers dying units.
  • Leaky ReLU / PReLU / ELU / SELU (2013-2016): small negative slope variants to keep gradients flowing.
  • GELU / SiLU / Swish (2016-2020): smooth, non-monotonic-ish probabilistic gating — the current default in transformers. (GELU ≈ z·Φ(z), the score times "how confidently positive am I".)
  • SwiGLU (2020): gated variant used in LLaMA, PaLM, Mistral — multiply two projections, one through SiLU.
  • Output-only: softmax, identity, sigmoid — chosen by the loss, never by taste (section 3 callout).

The zoo is drop-in in any framework:

python— PyTorch exposes the whole zoo as drop-in modules
import torch.nn as nn

hidden_options = {
    "relu":  nn.ReLU(),
    "gelu":  nn.GELU(),
    "silu":  nn.SiLU(),      # swish with beta=1
    "leaky": nn.LeakyReLU(0.01),
    "elu":   nn.ELU(),
    "tanh":  nn.Tanh(),
}
# output heads
binary_head  = nn.Sigmoid()          # + BCELoss
multi_head   = nn.Softmax(dim=-1)    # or leave logits + CrossEntropyLoss

08.The Empirical Hierarchy (2024-2026) and Saturation as the Common Enemy

You do not have to guess the winners — large-scale comparisons settled much of it. Ramachandran et al. 2018 explored >1000 candidates; Ramachandran/Yogesh 2019 benchmarked GELU/SELU vs ReLU across vision and RL. Production practice converged:

  • Vision CNNs: plain ReLU still wins (fast, hardware-friendly); Leaky ReLU where dying units appear (GAN discriminators).
  • Transformers/LLMs: GELU in encoder-era models (BERT), SiLU/SwiGLU in decoder LLMs (LLaMA, Mistral, Gemma, Claude-class models). SwiGLU consistently beats plain GLU variants on perplexity.
  • Self-normalizing nets: SELU + dropout for shallow MLPs on tabular data — largely superseded by normalization layers.
  • Regression outputs: identity; binary: sigmoid; multiclass: softmax (or raw logits + fused CE loss).

The practical rule, worth tattooing:

Insight

Hidden layers start with the ReLU-family default for your architecture; only tune activations after exhausting learning rate, batch size, and depth.

Now zoom out to the thread that organized all of it — saturation is the common enemy.

Every saturating activation (sigmoid, tanh, softmax at extremes) has regions where φ'(z) ≈ 0. During backprop those regions multiply gradients toward zero as they pass layer to layer — exactly the 0.0066³ arithmetic of section 4 — the vanishing gradient problem that made 1990s sigmoid networks of 10+ layers untrainable.

Unbounded activations (ReLU, GELU, SiLU) pass the positive-side gradient ~intact, letting deep networks train. But they introduce the mirror problem: exploding activations (and, without normalization, exploding gradients) — which is why modern deep training stacks pair ReLU-family activations with normalization layers and careful initialization. Topics 85 and 86 cover both failure modes in detail.

The 60-year arc in one line:

code
1958 step ──► 1986-2010 sigmoid/tanh ──► 2010-2015 ReLU wave ──► 2016-2026 GELU/SiLU/SwiGLU
 no gradient    gradient dies in             gradient highway,     same highway, sanded
                shoulders                    kinked, can die       smooth edges, gating

Each step kept expressivity and fixed the previous gradient-flow bug — which is the whole discipline in one sentence.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Sigmoid/tanh: bounded, stable, interpretable (probabilities); fine for outputs and gates (LSTM/GRU still use them internally).
  • ReLU family: no saturation on one side, O(1) compute, sparse activations, trains very deep nets.
  • GELU/SiLU: smooth gradients, better information flow, empirically strongest in attention-based models.

Trade-offs & Constraints

  • Sigmoid/tanh: gradient vanishing for |z| > ~5; sigmoid outputs are not zero-centered.
  • ReLU: dying units on negative drift; outputs non-zero-centered.
  • Exotic candidates (Maxout, SELU) add compute or fragility and rarely win modern benchmarks.
Production Implementation in Big Tech
Meta (LLaMA) & Mistral AI• SwiGLU feed-forward blocks in open-weights LLMs

LLaMA 1-3 and Mistral replace the classic ReLU/GELU MLP with SwiGLU: hidden = SiLU(x·W1) ⊙ (x·W3), then ·W2, with intermediate size shrunk to 8/3·d_model so the three matrices hold the same parameters as the old two. Ablations in the LLaMA paper explicitly tested ReLU and GLU variants and reported SwiGLU best on perplexity.

Staff+ Engineering Takeaways

  • Activations are the source of all expressivity beyond a single affine map; without them any depth collapses to W'x + b, so choosing them is really an optimization decision.
  • Judge activations by saturation (gradient magnitude), zero-centering, sparsity, and smoothness — no function scores perfect on all four.
  • Sigmoid/tanh killed 1990s deep networks via vanishing gradients (multiply sub-0.25 slopes per layer); unbounded ReLU-family revived them.
  • Current defaults: ReLU for CNNs, GELU for encoder transformers, SiLU/SwiGLU for decoder LLMs.
  • Output activations are set by the loss function, not by taste; softmax never belongs on a hidden layer.

Topic Knowledge Check

Exercise 1 of 4 • Test your architectural comprehension.

Exercise 1 of 40 answered
1

A network has 50 layers and NO activation functions between the linear layers. What is it equivalent to?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?