Sigmoid & Tanh: The Bounded Pioneers
Sigmoid squashes any score into (0,1) — perfect for "what is the probability?" — and tanh is the same S-curve re-centered on zero. Both carried neural networks from 1986 to 2010, and both share one fatal habit: for extreme inputs they flatten out, and a flat function passes on zero gradient. This page walks their formulas, the saturation arithmetic that capped depth, the zig-zag from non-centered outputs, and the three places they still run production AI today.
Sigmoid vs Tanh Geometry
tanh is simply a rescaled sigmoid: same saturating S-shape, but zero-centered with a steeper center slope (1.0 vs 0.25).
01.The Problem: The Score Needs a Volume Knob with Stops
Topic 75's perceptron ends every computation with a raw score z = w·x + b.
Two things feel missing about that score:
First, it is unbounded.
z can be 0.7 or 74,000 depending on the weights. But many questions want an answer between fixed stops:
"What is the probability this email is spam?"
Probabilities must live in [0, 1]. 74000 is not a probability. You need something that squeezes any number into a bounded range, smoothly.
Second, the hard step throws away everything.
Topic 75's step function answers "which side?" but never "by how much?" — and it has zero derivative, so gradients cannot flow (why deep learning left it behind, topic 76).
So the question becomes
Is there a smooth, bounded S-curve we can put in place of the step — one that says "probably yes" instead of "yes", and that calculus can differentiate?
Enter the two pioneers:
- Sigmoid — answers in
[0,1], born as the logistic growth curve from demography (Verhulst, 1838!) before neural networks ever existed. - Tanh — the same shape, re-centered on zero, answering in
[−1, 1].
Between 1986 and 2010, almost every neural network on Earth ran on one of these two. Understanding exactly why they were dethroned (topic 78's ReLU coup) is understanding the vanishing-gradient problem itself — which is why you still meet them everywhere from LSTMs to your final output layer.
02.The Idea in Plain Words: The Squeeze, and Its Mirror
Sigmoid (logistic function): σ(z) = 1 / (1 + e^-z), output in (0,1).
Read it in words: "take the negative exponent, add one, flip it over." Very negative z → e^-z explodes → 1/(huge) → 0. Very positive z → e^-z vanishes → 1/(1+0) → 1. Middle → something between.
Hyperbolic tangent: tanh(z) = (e^z − e^-z)/(e^z + e^-z) = 2σ(2z) − 1, output in (−1,1).
Their derivatives have beautiful closed forms in terms of themselves, which is partly why they dominated the backprop era:
σ'(z) = σ(z)(1 − σ(z))— maximum 0.25 at z = 0tanh'(z) = 1 − tanh²(z)— maximum 1.0 at z = 0
Unpack why "derivative expressed via output" mattered:
- During backprop you already stored the forward output (you need it anyway).
- So the gradient is
out × (1 − out)— one multiplication, zero exponentials. - 1986-era CPUs adored this. Hand-differentiating a whole network was the norm; these formulas made it painless.
Identity check: tanh is exactly a centered, 4x-steeper sigmoid.
codetanh(z) = 2σ(2z) − 1 ↑ ↑ ↑ ↑ same double lift output now (−1,1), shape slope & centered on 0 at 0 shift
If sigmoid was "a squashing function that emits probabilities", tanh was "the same shape fixed for optimization" — the 1998 LeCun/BoardPro era migration from sigmoid to tanh in hidden layers bought roughly one order of magnitude in usable depth.
03.A Simple Worked Example: Squeeze a Number, Then Four
Compute sigmoid by hand at a ladder of inputs. Use e ≈ 2.718.
codez = 0: σ = 1/(1+e^0) = 1/2 = 0.5000 z = 2: σ = 1/(1+e^-2) = 1/(1+0.135) = 0.8808 z = 5: σ = 1/(1+e^-5) = 1/(1+0.0067) = 0.9933 z = 10: σ ≈ 1/(1+0.000045) = 0.99995 z = −5: σ = 1 − σ(5) = 0.0067 (symmetry: σ(−z) = 1−σ(z))
Interpretation: a spam score of 5 means "99.3% confident". A score of 10 means "99.995%".
Now the problem, hidden in plain sight:
The gap between "99.3% sure" and "99.995% sure" is 0.0066 of output... bought by doubling the raw score.
The squeeze is working too well. Extreme inputs are crushed together.
Same story in tanh — it just straddles zero:
codez = 0: tanh = 0.000 z = 1: tanh = 0.762 z = 3: tanh = 0.995 z = 2: tanh = 0.964 z = 2.5: tanh = 0.987 z = 5: tanh = 0.99991
And now the gradient check (the formulas from section 2):
codeσ'(0) = 0.5·(1−0.5) = 0.2500 ← the BEST sigmoid slope that ever exists σ'(2) = 0.881·0.119 = 0.1050 σ'(5) = 0.993·0.0067 = 0.0066 σ'(10) = 0.99995·0.000045 ≈ 0.000045
A unit sitting at z = 10 learns 5,500 times slower than the same unit at z = 0. It is confidently saturated — and confidence is exactly what stops the learning.
04.Visual Intuition: The S-Curve and Its Two Dead Plateaus
Draw both, mark the plateaus, and the whole topic is on the page:
codesigmoid σ(z) tanh(z) 1 ┤ ────────── B 1 ┤ ────────── B │ ╭─── ← flat │ ╭─── ← flat: .5┤ ╭──╯ slope≈0: │ ╭──╯ slope→0 │ ╱ GRADIENT │ ╱ │ ╱ GRAVEYARD │╱ ← slope 1.0 at 0 │ ╱ ← slope 0.25 0╳ (4x sigmoid's best!) │╱ at 0 │╲ 0 ┼───────────────► z −1 ┤ ╲────────── A A ← flat: slope→0 ← flat again, centered on 0
Three zones to memorize:
- Zone A (left plateau,
z ≪ 0): output pinned near 0 (sigmoid) / −1 (tanh). Slope ≈ 0. - Zone B (right plateau,
z ≫ 0): output pinned near 1 / +1. Slope ≈ 0. - The middle (|z| small): the only place with real slope. Sigmoid's whole usable range is roughly
z ∈ [−5, 5]; tanh saturates even earlier, around|z| > 2.5.
Slope ≈ 0 means: push the input harder, output doesn't move.
And no moving output means, during backprop: push the gradient harder, no signal arrives at the weights.
The two functions differ only in where they sit and how tall their middle slope is — the plateaus are shared DNA. That shared plateau is the villain of section 6.
05.The Analogy: A Stereo Volume Knob That Maxes Out
Carry one analogy through the rest of this topic: sigmoid and tanh are stereo volume knobs with hard stops.
- Turn the knob between quiet and medium → the room's loudness tracks the knob beautifully. That's the middle of the S-curve, where
σ'is healthy. - Crank the knob to MAX → the amp is clipping. Moving the knob further changes nothing you can hear. The unit is saturated: output pinned at 1, gradient pinned at ≈ 0.
- A saturated unit is a broken microphone in a chain of relays (the hallway from topic 77): it keeps broadcasting "ONE" — or "ZERO" — no matter what happens upstream.
Now follow the training engineer, who walks backwards down the hallway shouting corrections ("turn that input UP by 0.3!"). Behind a maxed-out microphone, the correction math multiplies by σ' ≈ 0 — so the shout fades to silence before it reaches the front of the network.
The analogy also explains the sigmoid-vs-tanh difference:
- Sigmoid's idle hum is at 0.5 (it idles between its stops at the halfway mark, never at zero) — downstream equipment hears constant background noise on every wire; that hum causes the zig-zag in section 6.
- Tanh re-cut the knob: silence means 0, quiet-to-loud spans negative-to-positive — cleaner for an engineer trying to hear direction of change.
- Both knobs still max out. Re-cutting the center changed the hum, not the clipping. Only ReLU (topic 78) removed the right-hand stop entirely.
Keep "the knob still maxes out" in your pocket — it is the one sentence that separates tanh's historical win from its eventual loss.
06.Why AI Cares: Saturation — the Gradient Killer, in Three Acts
Both functions are saturating: as z → ±∞, outputs pinch against 0/1 (sigmoid) or ±1 (tanh), and the derivatives pinch toward 0 exponentially. Consequences during backprop, act by act:
Act 1 — Small-gradient chain effect (the multiplication catastrophe).
A 20-layer sigmoid net with each layer's derivative ≤ 0.25 multiplies gradients by at most 0.25^20 ≈ 10^-12 — early layers effectively stop learning (the vanishing-gradient crisis of topic 85).
Feel the decay step by step from the output back, at best-case slopes:
codelayer 20: 1.000000 layer 19: 0.250000 layer 18: 0.062500 layer 15: 0.000015 ← one hundred-thousandth, after 5 layers layer 1: ≈ 0.000000000001
Real units rarely sit at z = 0; sitting at z = 5 (σ' = 0.0066) makes the collapse far faster. Exponential decay is merciless: it doesn't fade, it falls off a cliff.
Act 2 — Dead zones are sticky.
Once weights push pre-activations deep into saturation, the unit emits near-constant output and its gradient is near-zero, so it drifts barely at all — a quasi-stable failure analogous to dying ReLU but statistical rather than absolute. A dead ReLU is a corpse (gradient exactly 0, permanent). A saturated sigmoid is a mummy wrapped in bandages: motion is 10^-9-slow, so it stays where it is — and being stuck at max volume is self-reinforcing, because no gradient can pull the knob back.
Act 3 — Not zero-centered (sigmoid's extra sin).
Sigmoid outputs are all positive: for a downstream weight w_ij, the gradient δ_j · a_i is always the sign of δ_j, so all incoming weights to unit j update in lockstep, producing zig-zag loss trajectories (Karpathy's CS231n "sigmoid zombies" analysis).
Watch it with two weights feeding one unit:
codea₁ = 0.8, a₂ = 0.6 (both positive — sigmoid can't help it) δ = +0.1 (error says "push output up") → grad w₁ = +0.1·0.8 = +0.08 (raise w₁) → grad w₂ = +0.1·0.6 = +0.06 (raise w₂) — ALWAYS the same direction as δ next step δ = −0.1 → BOTH flip to decrease together
Every update moves all incoming weights the same way together — even when the ideal move is "raise one, lower the other." The optimizer diagonal-steps toward the answer instead of walking straight: the zig-zag.
tanh mitigates (3) completely and (1) by 4x — but at depth, it too vanishes. The knob was re-centered and made steeper. It still maxes out.
07.In Practice: Where They Still Live in Modern Networks
Sigmoid and tanh were evicted from hidden layers of deep ReLU/GELU networks, but they remain the right tool in three places — all of them "a bounded, interpretable squish is the semantics" jobs:
- Gates. LSTMs (input/forget/output gates), GRUs, and highway networks use sigmoid — you want a bounded 0-1 multiplier that smoothly blends information: "let 0.83 of this memory through". A gate that can't max out isn't a gate. Modern state-space models (Mamba, 2023-2026 selective SSMs) likewise lean on sigmoid-like gating.
- Output heads. Binary classification: sigmoid + BCE loss. Multi-label: independent sigmoids per class (a photo can be "beach" AND "sunset"). And in LLM-era alignment, reward-model heads and calibration layers (Platt scaling / sigmoid temperatures) use them.
- Attention biases & normalization-adjacent squashing. Bounded activations occasionally appear in encoders for numerical hygiene (e.g., diffusion-model outputs scaled by tanh in some 2024 speech codecs).
The rule of thumb, in one line:
Sigmoid where a probability or gate is the semantics; tanh where you need a bounded zero-centered signal; neither where you need depth.
And when you do use sigmoid, use it like a production engineer — the stable piecewise form, and never compose raw sigmoid with a log-loss:
import numpy as np
def sigmoid_stable(z):
out = np.empty_like(z, dtype=float)
pos = z >= 0
out[pos] = 1.0 / (1.0 + np.exp(-z[pos])) # exp of negative arg: safe
e = np.exp(z[~pos]) # z<0 → exp(z) small
out[~pos] = e / (1.0 + e) # algebraically identical
return out
# torch.sigmoid / torch.nn.Sigmoid use equivalent stable CUDA kernels.
# For BCE, never sigmoid+BCELoss — use BCEWithLogitsLoss (log-sum-exp trick).08.Historical Arc: 1986 → 2010 → Now, and Why You Still Study Them
The timeline in three beats:
- 1986. Rumelhart/Hinton/Williams' landmark Nature paper on backpropagation used sigmoid everywhere — its self-referential derivative (
out·(1−out)) made hand-computable gradients practical. - 1998. LeCun's LeNet-5 (the digit recognizer behind real OCR deployments) used tanh hidden layers — one reason it trained where sigmoid nets stalled.
- 2010-2012. The ReLU wave (AlexNet) displaced both from hidden layers within two years, and Hinton famously recounted teaching students sigmoid for decades and then flipping to ReLU overnight.
A function that carried the field for a quarter century was demoted from "default" to "legacy" in about 24 months. This is the pace you are signing up for.
But the pedagogical value endures: sigmoid/tanh are the cleanest way to feel the three concepts every later activation choice is argued on:
- Saturation → the flat plateaus, the
σ' ≤ 0.25ceiling, the0.25^Ldecay — every "why do gradients vanish" story starts here. - Gradient vanishing → depth as exponent: the deeper the net, the more the knob's max-out compounds.
- Centering → the
δ_j · a_isign-lockstep, the zig-zag, and why tanh's re-centered knob beat its predecessor.
Final self-check before moving on — answer all four in plain words:
- Why is
σ'(z)at most 0.25, and where does that maximum sit? - Why does
tanh = 2σ(2z) − 1mean tanh "wins" at the same real estate but still loses at depth? - What does "all-positive outputs bias gradient signs" actually do to a weight matrix update, geometrically? (Zig-zag.)
- What single fact makes sigmoid irreplaceable at an LSTM forget gate, despite everything on this page? (A gate must max out — clipping is the job, not the bug.)
If you can explain why a 10-layer tanh MLP learns slower than a 10-layer ReLU MLP, you understand activation design.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Bounded outputs → activations cannot explode; historically stabilized training.
- Smooth, infinitely differentiable (C∞) — well-behaved gradients near the origin.
- sigmoid is the natural probability/gate emitter; tanh is the natural bounded zero-centered emitter.
- Closed-form derivatives in terms of outputs simplify hand backprop.
Trade-offs & Constraints
- Exponentially vanishing gradients in saturation zones — hard ceiling on depth.
- sigmoid outputs non-zero-centered → zig-zag optimization.
- Exponential compute (minor, but ReLU won partly on this in 2012 GPUs).
Production ASR acoustic models of the late 2010s (DeepSpeech-style and internal assistants) were stacks of tanh hidden-state LSTMs: forget/input/output gates computed by sigmoid decide what the cell retains, while tanh proposes and emits candidate/hidden activations. The bounded activations were essential — unbounded ReLU cells proved far harder to keep stable for multi-second memory horizons.
Staff+ Engineering Takeaways
- tanh = 2σ(2z) − 1: identical saturating shape, but zero-centered with max slope 1.0 vs sigmoid's 0.25.
- Saturation (derivative → 0 at the extremes) is the shared flaw that caps trainable depth — a 20-layer sigmoid net multiplies gradients by ≤ 0.25^20 ≈ 10^-12.
- Sigmoid's positive-only outputs cause correlated gradient signs (δ_j·a_i shares δ_j's sign) and zig-zag updates.
- Sigmoid survives wherever probabilities/gates are the semantics (BCE outputs, LSTM/GRU/SSM gates); tanh wherever bounded zero-centered signals are wanted. Gates WANT clipping — that is the feature.
- Use fused stable losses (BCEWithLogitsLoss) and piecewise stable sigmoid instead of naive exp(-z) forms to avoid overflow/underflow.
Topic Knowledge Check
Exercise 1 of 4 • Test your architectural comprehension.
Compared to sigmoid, tanh in a hidden layer mainly helps because it:
How clear and actionable was this distributed systems breakdown?