TOPIC #79Intermediate 12 min read

Sigmoid & Tanh: The Bounded Pioneers

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Sigmoid squashes any score into (0,1) — perfect for "what is the probability?" — and tanh is the same S-curve re-centered on zero. Both carried neural networks from 1986 to 2010, and both share one fatal habit: for extreme inputs they flatten out, and a flat function passes on zero gradient. This page walks their formulas, the saturation arithmetic that capped depth, the zig-zag from non-centered outputs, and the three places they still run production AI today.

Sigmoid vs Tanh Geometry

tanh is simply a rescaled sigmoid: same saturating S-shape, but zero-centered with a steeper center slope (1.0 vs 0.25).

Sigmoid vs Tanh Geometry
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: The Score Needs a Volume Knob with Stops

Topic 75's perceptron ends every computation with a raw score z = w·x + b.

Two things feel missing about that score:

First, it is unbounded.

z can be 0.7 or 74,000 depending on the weights. But many questions want an answer between fixed stops:

Insight

"What is the probability this email is spam?"

Probabilities must live in [0, 1]. 74000 is not a probability. You need something that squeezes any number into a bounded range, smoothly.

Second, the hard step throws away everything.

Topic 75's step function answers "which side?" but never "by how much?" — and it has zero derivative, so gradients cannot flow (why deep learning left it behind, topic 76).

So the question becomes

Insight

Is there a smooth, bounded S-curve we can put in place of the step — one that says "probably yes" instead of "yes", and that calculus can differentiate?

Enter the two pioneers:

  • Sigmoid — answers in [0,1], born as the logistic growth curve from demography (Verhulst, 1838!) before neural networks ever existed.
  • Tanh — the same shape, re-centered on zero, answering in [−1, 1].

Between 1986 and 2010, almost every neural network on Earth ran on one of these two. Understanding exactly why they were dethroned (topic 78's ReLU coup) is understanding the vanishing-gradient problem itself — which is why you still meet them everywhere from LSTMs to your final output layer.

02.The Idea in Plain Words: The Squeeze, and Its Mirror

Sigmoid (logistic function): σ(z) = 1 / (1 + e^-z), output in (0,1).

Read it in words: "take the negative exponent, add one, flip it over." Very negative z → e^-z explodes → 1/(huge) → 0. Very positive z → e^-z vanishes → 1/(1+0) → 1. Middle → something between.

Hyperbolic tangent: tanh(z) = (e^z − e^-z)/(e^z + e^-z) = 2σ(2z) − 1, output in (−1,1).

Their derivatives have beautiful closed forms in terms of themselves, which is partly why they dominated the backprop era:

  • σ'(z) = σ(z)(1 − σ(z)) — maximum 0.25 at z = 0
  • tanh'(z) = 1 − tanh²(z) — maximum 1.0 at z = 0

Unpack why "derivative expressed via output" mattered:

  • During backprop you already stored the forward output (you need it anyway).
  • So the gradient is out × (1 − out) — one multiplication, zero exponentials.
  • 1986-era CPUs adored this. Hand-differentiating a whole network was the norm; these formulas made it painless.

Identity check: tanh is exactly a centered, 4x-steeper sigmoid.

code
tanh(z) = 2σ(2z) − 1
   ↑        ↑    ↑    ↑
 same    double  lift  output now (−1,1),
 shape   slope   &      centered on 0
         at 0    shift

If sigmoid was "a squashing function that emits probabilities", tanh was "the same shape fixed for optimization" — the 1998 LeCun/BoardPro era migration from sigmoid to tanh in hidden layers bought roughly one order of magnitude in usable depth.

03.A Simple Worked Example: Squeeze a Number, Then Four

Compute sigmoid by hand at a ladder of inputs. Use e ≈ 2.718.

code
z = 0:    σ = 1/(1+e^0)  = 1/2          = 0.5000
z = 2:    σ = 1/(1+e^-2) = 1/(1+0.135)  = 0.8808
z = 5:    σ = 1/(1+e^-5) = 1/(1+0.0067) = 0.9933
z = 10:   σ ≈ 1/(1+0.000045)            = 0.99995
z = −5:   σ = 1 − σ(5)                  = 0.0067   (symmetry: σ(−z) = 1−σ(z))

Interpretation: a spam score of 5 means "99.3% confident". A score of 10 means "99.995%".

Now the problem, hidden in plain sight:

Insight

The gap between "99.3% sure" and "99.995% sure" is 0.0066 of output... bought by doubling the raw score.

The squeeze is working too well. Extreme inputs are crushed together.

Same story in tanh — it just straddles zero:

code
z = 0:    tanh = 0.000     z = 1:  tanh = 0.762     z = 3:  tanh = 0.995
z = 2:    tanh = 0.964     z = 2.5: tanh = 0.987    z = 5:  tanh = 0.99991

And now the gradient check (the formulas from section 2):

code
σ'(0)  = 0.5·(1−0.5)   = 0.2500   ← the BEST sigmoid slope that ever exists
σ'(2)  = 0.881·0.119   = 0.1050
σ'(5)  = 0.993·0.0067  = 0.0066
σ'(10) = 0.99995·0.000045 ≈ 0.000045

A unit sitting at z = 10 learns 5,500 times slower than the same unit at z = 0. It is confidently saturated — and confidence is exactly what stops the learning.

04.Visual Intuition: The S-Curve and Its Two Dead Plateaus

Draw both, mark the plateaus, and the whole topic is on the page:

code
 sigmoid σ(z)                        tanh(z)
  1 ┤              ────────── B       1 ┤            ────────── B
    │          ╭───             ← flat  │        ╭───             ← flat:
  .5┤       ╭──╯    slope≈0:            │     ╭──╯                 slope→0
    │     ╱       GRADIENT              │   ╱
    │   ╱       GRAVEYARD               │╱ ← slope 1.0 at 0
    │ ╱ ← slope 0.25                    0╳  (4x sigmoid's best!)
    │╱    at 0                          │╲
  0 ┼───────────────► z               −1 ┤ ╲────────── A
    A ← flat: slope→0                     ← flat again, centered on 0

Three zones to memorize:

  • Zone A (left plateau, z ≪ 0): output pinned near 0 (sigmoid) / −1 (tanh). Slope ≈ 0.
  • Zone B (right plateau, z ≫ 0): output pinned near 1 / +1. Slope ≈ 0.
  • The middle (|z| small): the only place with real slope. Sigmoid's whole usable range is roughly z ∈ [−5, 5]; tanh saturates even earlier, around |z| > 2.5.

Slope ≈ 0 means: push the input harder, output doesn't move.

And no moving output means, during backprop: push the gradient harder, no signal arrives at the weights.

The two functions differ only in where they sit and how tall their middle slope is — the plateaus are shared DNA. That shared plateau is the villain of section 6.

05.The Analogy: A Stereo Volume Knob That Maxes Out

Carry one analogy through the rest of this topic: sigmoid and tanh are stereo volume knobs with hard stops.

  • Turn the knob between quiet and medium → the room's loudness tracks the knob beautifully. That's the middle of the S-curve, where σ' is healthy.
  • Crank the knob to MAX → the amp is clipping. Moving the knob further changes nothing you can hear. The unit is saturated: output pinned at 1, gradient pinned at ≈ 0.
  • A saturated unit is a broken microphone in a chain of relays (the hallway from topic 77): it keeps broadcasting "ONE" — or "ZERO" — no matter what happens upstream.

Now follow the training engineer, who walks backwards down the hallway shouting corrections ("turn that input UP by 0.3!"). Behind a maxed-out microphone, the correction math multiplies by σ' ≈ 0 — so the shout fades to silence before it reaches the front of the network.

The analogy also explains the sigmoid-vs-tanh difference:

  • Sigmoid's idle hum is at 0.5 (it idles between its stops at the halfway mark, never at zero) — downstream equipment hears constant background noise on every wire; that hum causes the zig-zag in section 6.
  • Tanh re-cut the knob: silence means 0, quiet-to-loud spans negative-to-positive — cleaner for an engineer trying to hear direction of change.
  • Both knobs still max out. Re-cutting the center changed the hum, not the clipping. Only ReLU (topic 78) removed the right-hand stop entirely.

Keep "the knob still maxes out" in your pocket — it is the one sentence that separates tanh's historical win from its eventual loss.

06.Why AI Cares: Saturation — the Gradient Killer, in Three Acts

Both functions are saturating: as z → ±∞, outputs pinch against 0/1 (sigmoid) or ±1 (tanh), and the derivatives pinch toward 0 exponentially. Consequences during backprop, act by act:

Act 1 — Small-gradient chain effect (the multiplication catastrophe).

A 20-layer sigmoid net with each layer's derivative ≤ 0.25 multiplies gradients by at most 0.25^20 ≈ 10^-12 — early layers effectively stop learning (the vanishing-gradient crisis of topic 85).

Feel the decay step by step from the output back, at best-case slopes:

code
layer 20: 1.000000
layer 19: 0.250000
layer 18: 0.062500
layer 15: 0.000015      ← one hundred-thousandth, after 5 layers
layer 1:  ≈ 0.000000000001

Real units rarely sit at z = 0; sitting at z = 5 (σ' = 0.0066) makes the collapse far faster. Exponential decay is merciless: it doesn't fade, it falls off a cliff.

Act 2 — Dead zones are sticky.

Once weights push pre-activations deep into saturation, the unit emits near-constant output and its gradient is near-zero, so it drifts barely at all — a quasi-stable failure analogous to dying ReLU but statistical rather than absolute. A dead ReLU is a corpse (gradient exactly 0, permanent). A saturated sigmoid is a mummy wrapped in bandages: motion is 10^-9-slow, so it stays where it is — and being stuck at max volume is self-reinforcing, because no gradient can pull the knob back.

Act 3 — Not zero-centered (sigmoid's extra sin).

Sigmoid outputs are all positive: for a downstream weight w_ij, the gradient δ_j · a_i is always the sign of δ_j, so all incoming weights to unit j update in lockstep, producing zig-zag loss trajectories (Karpathy's CS231n "sigmoid zombies" analysis).

Watch it with two weights feeding one unit:

code
a₁ = 0.8, a₂ = 0.6      (both positive — sigmoid can't help it)
δ = +0.1                (error says "push output up")
→ grad w₁ = +0.1·0.8 = +0.08   (raise w₁)
→ grad w₂ = +0.1·0.6 = +0.06   (raise w₂) — ALWAYS the same direction as δ
next step δ = −0.1      → BOTH flip to decrease together

Every update moves all incoming weights the same way together — even when the ideal move is "raise one, lower the other." The optimizer diagonal-steps toward the answer instead of walking straight: the zig-zag.

tanh mitigates (3) completely and (1) by 4x — but at depth, it too vanishes. The knob was re-centered and made steeper. It still maxes out.

07.In Practice: Where They Still Live in Modern Networks

Sigmoid and tanh were evicted from hidden layers of deep ReLU/GELU networks, but they remain the right tool in three places — all of them "a bounded, interpretable squish is the semantics" jobs:

  1. Gates. LSTMs (input/forget/output gates), GRUs, and highway networks use sigmoid — you want a bounded 0-1 multiplier that smoothly blends information: "let 0.83 of this memory through". A gate that can't max out isn't a gate. Modern state-space models (Mamba, 2023-2026 selective SSMs) likewise lean on sigmoid-like gating.
  2. Output heads. Binary classification: sigmoid + BCE loss. Multi-label: independent sigmoids per class (a photo can be "beach" AND "sunset"). And in LLM-era alignment, reward-model heads and calibration layers (Platt scaling / sigmoid temperatures) use them.
  3. Attention biases & normalization-adjacent squashing. Bounded activations occasionally appear in encoders for numerical hygiene (e.g., diffusion-model outputs scaled by tanh in some 2024 speech codecs).

The rule of thumb, in one line:

Insight

Sigmoid where a probability or gate is the semantics; tanh where you need a bounded zero-centered signal; neither where you need depth.

And when you do use sigmoid, use it like a production engineer — the stable piecewise form, and never compose raw sigmoid with a log-loss:

python— Stable sigmoid — the textbook piecewise implementation
import numpy as np

def sigmoid_stable(z):
    out = np.empty_like(z, dtype=float)
    pos = z >= 0
    out[pos] = 1.0 / (1.0 + np.exp(-z[pos]))      # exp of negative arg: safe
    e = np.exp(z[~pos])                            # z<0 → exp(z) small
    out[~pos] = e / (1.0 + e)                      # algebraically identical
    return out

# torch.sigmoid / torch.nn.Sigmoid use equivalent stable CUDA kernels.
# For BCE, never sigmoid+BCELoss — use BCEWithLogitsLoss (log-sum-exp trick).

08.Historical Arc: 1986 → 2010 → Now, and Why You Still Study Them

The timeline in three beats:

  • 1986. Rumelhart/Hinton/Williams' landmark Nature paper on backpropagation used sigmoid everywhere — its self-referential derivative (out·(1−out)) made hand-computable gradients practical.
  • 1998. LeCun's LeNet-5 (the digit recognizer behind real OCR deployments) used tanh hidden layers — one reason it trained where sigmoid nets stalled.
  • 2010-2012. The ReLU wave (AlexNet) displaced both from hidden layers within two years, and Hinton famously recounted teaching students sigmoid for decades and then flipping to ReLU overnight.
Insight

A function that carried the field for a quarter century was demoted from "default" to "legacy" in about 24 months. This is the pace you are signing up for.

But the pedagogical value endures: sigmoid/tanh are the cleanest way to feel the three concepts every later activation choice is argued on:

  1. Saturation → the flat plateaus, the σ' ≤ 0.25 ceiling, the 0.25^L decay — every "why do gradients vanish" story starts here.
  2. Gradient vanishing → depth as exponent: the deeper the net, the more the knob's max-out compounds.
  3. Centering → the δ_j · a_i sign-lockstep, the zig-zag, and why tanh's re-centered knob beat its predecessor.

Final self-check before moving on — answer all four in plain words:

  • Why is σ'(z) at most 0.25, and where does that maximum sit?
  • Why does tanh = 2σ(2z) − 1 mean tanh "wins" at the same real estate but still loses at depth?
  • What does "all-positive outputs bias gradient signs" actually do to a weight matrix update, geometrically? (Zig-zag.)
  • What single fact makes sigmoid irreplaceable at an LSTM forget gate, despite everything on this page? (A gate must max out — clipping is the job, not the bug.)

If you can explain why a 10-layer tanh MLP learns slower than a 10-layer ReLU MLP, you understand activation design.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Bounded outputs → activations cannot explode; historically stabilized training.
  • Smooth, infinitely differentiable (C∞) — well-behaved gradients near the origin.
  • sigmoid is the natural probability/gate emitter; tanh is the natural bounded zero-centered emitter.
  • Closed-form derivatives in terms of outputs simplify hand backprop.

Trade-offs & Constraints

  • Exponentially vanishing gradients in saturation zones — hard ceiling on depth.
  • sigmoid outputs non-zero-centered → zig-zag optimization.
  • Exponential compute (minor, but ReLU won partly on this in 2012 GPUs).
Production Implementation in Big Tech
LSTM-based speech recognition (Apple Siri, Amazon Alexa, 2016-2020 era)• Sigmoid gates steering memory

Production ASR acoustic models of the late 2010s (DeepSpeech-style and internal assistants) were stacks of tanh hidden-state LSTMs: forget/input/output gates computed by sigmoid decide what the cell retains, while tanh proposes and emits candidate/hidden activations. The bounded activations were essential — unbounded ReLU cells proved far harder to keep stable for multi-second memory horizons.

Staff+ Engineering Takeaways

  • tanh = 2σ(2z) − 1: identical saturating shape, but zero-centered with max slope 1.0 vs sigmoid's 0.25.
  • Saturation (derivative → 0 at the extremes) is the shared flaw that caps trainable depth — a 20-layer sigmoid net multiplies gradients by ≤ 0.25^20 ≈ 10^-12.
  • Sigmoid's positive-only outputs cause correlated gradient signs (δ_j·a_i shares δ_j's sign) and zig-zag updates.
  • Sigmoid survives wherever probabilities/gates are the semantics (BCE outputs, LSTM/GRU/SSM gates); tanh wherever bounded zero-centered signals are wanted. Gates WANT clipping — that is the feature.
  • Use fused stable losses (BCEWithLogitsLoss) and piecewise stable sigmoid instead of naive exp(-z) forms to avoid overflow/underflow.

Topic Knowledge Check

Exercise 1 of 4 • Test your architectural comprehension.

Exercise 1 of 40 answered
1

Compared to sigmoid, tanh in a hidden layer mainly helps because it:

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?