Softmax: From Scores to Probability Distributions
Your model spits out raw scores. Softmax turns any list of scores into probabilities that add up to 1 — the output layer of almost every classifier and the heart of attention. Learn its shift-invariance, the temperature knob, the subtract-the-max stability trick, and its famously clean cross-entropy gradient (p − y).
Softmax Pipeline and Temperature
Exponentiate to enforce positivity, normalize to enforce sum 1. The temperature T divides logits before this map and interpolates from argmax to uniform.
01.The Problem: Raw Scores Are Not Probabilities
A model looks at a photo. You ask it: "what is this?"
Inside, the last layer spits out three raw numbers.
cat: 2.4 dog: −0.7 fish: 1.1
These numbers are called scores, or fancier, logits. One per class. That's all they are — unbounded, arbitrary real numbers produced by a matrix multiply.
So the question becomes
How do I turn scores into probabilities?
Because "the probability it's a cat" should be something like 0.66 — not 2.4. A probability needs three things a raw score does not have:
- Never negative.
−0.7cannot be a probability. - Never above 1.
2.4cannot be a probability. - Everything adds up to exactly 1. Here
2.4 − 0.7 + 1.1 = 2.8. Nope.
You could clamp negatives to 0 and divide by the sum — but that hack has zero gradient below 0 (like a ReLU bolted onto your output), and training needs a smooth gradient everywhere, for every class.
You want one smooth, differentiable machine that takes any list of real numbers and outputs one valid probability distribution.
That machine is softmax.
It is the output layer of virtually every classifier you will ever build — and, as you will see in section 8, the heart of attention in every LLM.
02.The Idea in Plain Words: Exponentiate, Then Split One Pot
Softmax is simply
Turn every score into a positive number by exponentiating it, then divide by the total so everyone shares one pot.
For a logit vector z ∈ R^K (one raw score per class):
softmax(z)_i = e^{z_i} / Σ_{j=1..K} e^{z_j}
Let's unpack it piece by piece.
z_iis classi's raw score. Any real number.e^{z_i}is that score pushed through the exponential. Exponentials are always positive — problem 1 solved instantly.Σ_{j} e^{z_j}is the sum over all classes. Dividing by it makes the outputs sum to exactly 1 — problem 3 solved.- Each output
p_i = softmax(z)_iis classi's probability: strictly between 0 and 1 — problem 2 solved. - The exponential also does something opinionated: it amplifies differences. A score 2 points higher becomes
e² ≈ 7.4times bigger, not 2 times bigger. Softmax is a winner-takes-most distributor.
Three properties make it the canonical classifier output:
- Shift invariance:
softmax(z + c) = softmax(z)for any constant c. Add 5 to every score — the probabilities do not move. Only differences between logits matter. That fact is why "logit" is the right name, and why (section 7) you can subtract the max for free. - Scale sensitivity:
softmax(z/T)is the temperature-controlled version. Dividing the scores before exponentiating compresses or stretches their gaps. T→0 sharpens toward a one-hot argmax (winner takes everything); T→∞ flattens toward uniform1/K(everyone ties). LLM sampling, knowledge distillation (Hinton 2015), and calibration all turn this one dial. - Differentiability + log-prob conjugacy: combined with cross-entropy loss, the gradient takes the simplest possible form (
p − y, section 6). That cleanliness is why the pairing is universal.
03.A Tiny Worked Example: Logits [2, 1, 0]
Three scores: z = [2, 1, 0] for classes [cat, dog, fish]. No negatives, no weirdness — let's just apply the recipe.
Step 1 — exponentiate each score.
e² ≈ 7.389
e¹ ≈ 2.718
e⁰ = 1.000
All positive. ✓
Step 2 — sum them.
7.389 + 2.718 + 1.000 = 11.107
Step 3 — divide each by the total.
p_cat = 7.389/11.107 ≈ 0.665
p_dog = 2.718/11.107 ≈ 0.245
p_fish = 1.000/11.107 ≈ 0.090
Check: 0.665 + 0.245 + 0.090 = 1. A valid distribution, from arbitrary scores.
Now feel the shift invariance. Subtract the max (2) from every score: z − 2 = [0, −1, −2].
e⁰ = 1, e^{−1} ≈ 0.368, e^{−2} ≈ 0.135 — total 1.503.
Divide: 1/1.503 ≈ 0.665, 0.368/1.503 ≈ 0.245, 0.135/1.503 ≈ 0.090.
Identical probabilities. The exponential of the shift, e^{−c}, appears in every numerator and in the denominator, and cancels. You just proved the trick that keeps softmax from overflowing (section 7).
Now turn the temperature knob. With T = 0.5 the scores are divided by 0.5 (i.e., doubled): [4, 2, 0] → e⁴ ≈ 54.6, e² ≈ 7.39, 1 → p ≈ [0.867, 0.117, 0.016]. Sharper: the favorite takes 87%.
With T = 2 scores are halved: [1, 0.5, 0] → p ≈ [0.506, 0.307, 0.187]. Flatter: the underdog fish gets almost 19%.
Same ranking, different confidence — the whole of LLM "temperature" in two arithmetic lines.
04.Visual Intuition: The Megaphone and the Pie
Picture the scores as people shouting their bid for a share of one pie. The exponential is a megaphone: loud voices become much louder, quiet voices barely get heard. Then the pie is cut strictly in proportion to the amplified shouting.
codescores (logits) after e^z (megaphone) shares of the pie (softmax) cat +2.0 ████ → 7.39 ██████████████ → 66.5% ████████████ dog +1.0 ██ → 2.72 █████ → 24.5% ████ fish 0.0 · → 1.00 ██ → 9.0% █ gap of 1.0 in scores → 2.7x in exponentials → 2.7x in probabilities gap of 2.0 in scores → 7.4x in exponentials → 7.4x in probabilities
And temperature is the volume dial on the megaphone:
codeT → 0 T = 1 T → ∞ ┌──┐ ┌──┐── ┌──┐──┐ │██│ winner │██│░░│ softmax │██│██│ everyone ties │ │ takes all │░░│ as-is │██│██│ at 1/K └──┘ └──┘── └──┘──┘ one-hot argmax uniform
Two geometric notes worth keeping in your head:
- The output
palways lives on the probability simplex: the flat triangle of all vectors withp_i > 0andΣp = 1. Softmax maps the entire infinite space of logits onto this tiny triangle (never touching its edges — no probability is ever exactly 0 or 1). - Because only differences matter, shifting all logits moves you along a direction softmax ignores — the direction perpendicular to the simplex.
05.The Analogy: A Bookmaker Turning Stakes Into Odds
Carry one analogy through the rest of this topic: softmax is a bookmaker at the race track.
Horse bettors hand the bookie their raw stakes — any amount, any direction of enthusiasm. That's your logits: [7.39, 2.72, 1.00].
The bookie must publish shares of the total pot: each payout as a fraction of the whole. That's exactly what dividing by Σe^{z_j} does — no matter how wild the stakes, published shares always sum to 1.
Now the temperature:
- A cold bookie (T→0) simply says "the favorite takes the whole pot" — the argmax, one-hot.
- A hot bookie (T→∞) says "it's everyone's race" — uniform
1/K. - The standard bookie (T = 1) splits proportionally to the exponentiated stakes.
And the beautiful part: the bookie never argues about absolute amounts. If every bettor doubles or shifts their stake identically, the published shares move only with relative enthusiasm — shift invariance is "only the gaps between bets matter".
Hold this picture: scores are bets, the exponential is the bookie's preference for amplifying gaps, and the probability vector is the published share of one pot. Everything below — the gradient, the max-subtraction, attention — is the same bookie doing different jobs.
06.Why AI Cares: The Cross-Entropy Gradient Is Just ŷ − y
This is the single most satisfying derivation in deep learning, and it explains why softmax is the output layer.
Setup: true label as a one-hot vector y (all zeros, 1 at the correct class), prediction p = softmax(z). Cross-entropy loss (the negative log of the probability you gave the right answer):
L = −Σ y_i log p_i = −log p_target
Now differentiate L with respect to each logit z_i. The result collapses to:
∂L/∂z_i = p_i − y_i
The gradient of the logits is simply "prediction minus label". Nothing else. Over-predicted classes get pushed down, under-predicted classes pushed up, all in exact proportion to their current probability.
Run our example. Truth is cat, so y = [1, 0, 0] and p = [0.665, 0.245, 0.090]:
∂L/∂z = [−0.335, +0.245, +0.090]
Read the update z ← z − lr·∂L/∂z: the cat logit rises (subtracting a negative), dog and fish logits fall. And notice the sizes: dog, at 24.5%, gets a bigger correction than fish at 9% — the model is told to fix its biggest overconfidence first.
Why does it collapse so nicely? The Jacobian (matrix of partial derivatives) of softmax itself is ∂p_i/∂z_j = p_i(δ_ij − p_j) — not obviously simple. But when the chain rule contracts it against the one-hot label and the log from the loss, every messy term cancels into p − y.
Two consequences:
- The loss's gradient is this clean and the loss is convex in the logits, so softmax + cross-entropy is the standard multiclass head everywhere — from 10-class MNIST MLPs to 150k-token-vocabulary LLMs, where the final matmul produces logits and the training loss gradient is exactly this residual.
- This is the same
p − ybeauty you met in the gradient topics: classification training is, layer by final layer, "move the numbers toward the label".
07.In Practice: Never Let an Exponential Overflow
Naive softmax crashes on real numbers. For z_i = 1000 in float32, e^1000 = inf, and inf/inf = NaN. Your whole training run dies on one line.
Shift invariance hands you the fix for free:
softmax(z) = softmax(z − max(z))
Subtract the largest score from all of them. After the shift the biggest exponent is e^0 = 1 — no overflow is possible; the remaining terms underflow harmlessly toward 0 (they should be ~0 anyway). You watched this exact computation work in section 3.
The stable denominator log(Σe^z) deserves its own name: logsumexp. Frameworks implement it as one fused max + sum-of-shifted-exponentials pass.
Production rule: never compose Softmax() + CrossEntropyLoss() by hand. Use nn.CrossEntropyLoss / F.cross_entropy, which fuse the operations and compute the loss via logsumexp directly on the logits — the softmax probabilities are never even formed. Beyond stability this saves enormous memory: at 100k+ vocabulary, the softmax matrix over (batch × sequence × vocab) would dominate your activation memory, which is why fused kernels like Liger and Cut Cross Entropy avoid the logits-to-probs tensor entirely.
import numpy as np
def softmax_stable(z): # z: (..., K)
m = z.max(axis=-1, keepdims=True)
e = np.exp(z - m) # largest exponent is now 0
return e / e.sum(axis=-1, keepdims=True)
def logsumexp(z): # stable log(Σ e^z)
m = z.max(axis=-1, keepdims=True)
return (m + np.log(np.exp(z - m).sum(-1, keepdims=True))).squeeze(-1)
# torch equivalents (what you actually use in training):
# loss = F.cross_entropy(logits, targets) # fused, stable, no softmax call
# probs = F.softmax(logits / temperature, dim=-1) # only at inference/sampling08.Softmax Beyond Output Layers: The Universal "Soft Choice" Operator
Remember the bookie. Any time you need "turn a list of arbitrary preference scores into shares of one pot", the same operation applies. Softmax is a soft-argmax / soft-selection mechanism, and that is its second great use:
- Attention:
Attention(Q,K) = softmax(QK^T/√d)V. The dot productsQK^Tare raw "how relevant is position j to position i?" scores — logits, in the hundreds-of-tokens count. The softmax turns them into a probability distribution over context positions, so each output becomes a weighted average of value vectors. Multi-head, grouped-query, and multi-quant attention in every 2024–2026 LLM keep this exact softmax step. (Plain words: attention is the bookie splitting "100% of my looking" across the tokens in the sentence.) - Mixture-of-Experts routing: router logits → softmax → top-k weighted combination of experts (Mixtral, Switch-style). The shares decide how much each expert contributes.
- Attention over time / alignment (seq2seq, 2015): the original Bahdanau usage — soft alignment over source words.
- Sampling: a temperature-scaled softmax on LLM logits is the token distribution at generation time (nucleus/top-p sampling then post-processes this same vector).
One sharp caveat for hidden layers: never use softmax as a hidden activation. Its outputs must sum to 1, so it couples every unit's value to every other unit — a neuron can only get louder by making its neighbors quieter. That wastes degrees of freedom and starves "losing" classes of gradient (they saturate toward 0). If you genuinely want a bounded-selection hidden signal, use sparse alternatives (top-k, k-WTA) instead.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Calibrated-ish probability output with a clean, minimal (p − y) cross-entropy gradient.
- Shift invariance → numerically stabilizable at zero cost by subtracting the max.
- Temperature parameterization gives a principled sharpness knob (sampling, distillation, calibration).
- Universal: the same "shares of one pot" op serves classification heads, attention, and expert routing.
Trade-offs & Constraints
- All-out competition: increasing one class probability necessarily decreases others — unsuitable for multi-label tasks (use per-class sigmoids instead).
- O(K) exponentials and a global normalization term; at 100k+ vocab it dominates head compute/memory without fused kernels.
- Overconfidence: deep nets' softmax outputs are poorly calibrated without temperature scaling or label smoothing.
Inside every transformer layer, softmax(QK^T/√d_k) defines attention over the context — the mechanism that lets a model "look at" relevant tokens. At the output, a final softmax over the full vocabulary (~32k-256k logits) yields the next-token distribution; generation temperature/top-p operate on exactly this vector. Training uses fused cross-entropy on raw logits (the p−y residual) so the giant softmax never materializes during backprop.
Staff+ Engineering Takeaways
- softmax(z)_i = e^z_i / Σe^z_j: exponentials for positivity, the sum for normalization; only differences between logits matter (shift-invariant).
- With cross-entropy, ∂L/∂logits = p − y — the cleanest gradient in deep learning and the reason the pairing is universal.
- Stable implementation always subtracts the max (or works in logspace via logsumexp); use fused framework losses on logits, never softmax+CE manually.
- Temperature T divides logits before softmax: T→0 argmax, T→∞ uniform — the lever for sampling, distillation, and calibration.
- Softmax is also a soft-selection operator: attention weights and MoE routing are softmax over similarities, not classes. Never use it as a hidden activation.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Why does subtracting max(z) from every logit before computing softmax leave the answer exactly the same?
How clear and actionable was this distributed systems breakdown?