Cross-Entropy
Cross-entropy is the surprise bill: your model sets the odds, reality picks the outcome, and you pay −log(probability you assigned to what actually happened). Minimizing it means being honestly surprised as little as possible — and with softmax the gradient collapses to the beautiful error signal q − y.
01.The Problem: You Need a Number for "How Wrong Are My Odds?"
Your model looks at a photo and says:
- 70% cat
- 25% dog
- 5% bird
The truth is: dog.
How wrong is that? And how do you turn "wrongness" into a smooth number a gradient can chase (Topic 11)?
The obvious candidate — plain accuracy — fails immediately:
Accuracy is a count of right/wrong. It has zero gradient almost everywhere. You cannot train on it.
You need a loss that:
- Punishes confident wrongness hard.
- Rewards correct confidence gently.
- Is smooth, so optimization works.
So the question becomes
Can we price each prediction by how surprised the model should be at reality?
Topic 21 already built the price tag: surprisal I = −log p. Cross-entropy just aims it at models.
02.The Idea in Plain Words: Surprise Under Someone Else's Code
Cross-entropy is simply
H(p, q) = − Σₓ p(x) log q(x) — the average surprisal of reality p, when your model q sets the prices.
Unpack the cast:
pis reality — the true distribution. In classification it is usually one-hot: the correct class has p = 1, everything else 0.qis the model — its predicted probabilities.−log q(x)is the price your model's code assigns to event x (Topic 21's telegraph rates, but your rates, not the optimal ones).
The three-way identity ties Topics 21 and 23 together:
H(p, q) = H(p) + KL(p ‖ q)
true cost + a non-negative penalty for q disagreeing with p. Consequences worth memorizing:
H(p, q) ≥ H(p)always; equality exactly when q = p.- Asymmetric:
H(p, q) ≠ H(q, p)— coding p with q's code is not the same transaction as the reverse. - When p is a fixed one-hot label, everything cancels except the true class:
H(p, q) = −log q(correct class). The loss is literally the surprisal of the right answer under the model. Minimizing it = raising the true label's probability.
03.A Simple Worked Example: Paying the Bill
Your model predicts: 70% cat, 25% dog, 5% bird. Reality collects: it was a dog.
Your bill:
L = −ln 0.25 ≈ 1.386 nats
Now change only the model's confidence, same truth:
codeP(dog) assigned Bill −ln P(dog) 0.99 0.010 ← nearly free 0.25 1.386 ← hedged, moderate 0.05 3.000 ← wrong AND bold 0.001 6.908 ← ruinously confident nonsense 0.0 ∞ ← −log 0: the model is deleted from reality
See the shape of the punishment:
- Being wrong is expensive; being confidently wrong is unboundedly expensive.
- Being right quietly (0.25 on the truth) already beats being confidently wrong on it.
- That log-shaped squeeze is exactly why gradient descent pushes probabilities of true labels up fast at first, then carefully.
04.Visual Intuition: The −log Curve
The entire loss, drawn:
codeloss ▲ │╲ │ ╲ −log q : infinite punishment │ ╲ at q = 0, free at q = 1 │ ╲ │ ╲__ │ ╲____ │ ╲─────___ └──┬──────┬──────┬────╱──► q = probability 0.0 0.25 0.5 1.0 the model gave the truth
Two features do all the work:
- The wall at q → 0. Any real chance the truth gets probability ~0 costs ∞. So the model can never "safely" exclude an outcome it might face.
- The gentle floor at q → 1. Near-perfect predictions earn near-zero loss — diminishing returns on already-good confidence.
Compare the naive alternative, MSE on the probability: it treats "predicted 0.0 for the truth" as a bounded error of 1.0. The log says: that error is unbounded, and it is right — overconfident nonsense must be punished harder than ordinary mistake.
05.The Analogy: The Bookie Who Must Quote Every Outcome
Carry one analogy through the rest: your model is a bookie setting odds on tomorrow.
- Before the race, the bookie publishes odds for every horse: those are the probabilities q. Rules: all positive, must sum to 1 (softmax guarantees this).
- After the race, reality names the winner, and the bookie pays out
−log q(winner). - A cautious bookie who spreads odds loses slowly. A bookie who quotes 1000:1 against the actual winner is bankrupt.
Training = a bookie who has run the same race millions of times and keeps adjusting odds to minimize the average payout. And the "true" optimal bookie quotes exactly the real frequencies: then average payout = H(p), entropy, the irreducible cost of the race itself (Topic 21). Anything worse is the KL penalty for a wrong belief (Topic 23).
This is why the identity H(p,q) = H(p) + KL(p‖q) is the bookie's ledger: fair price + ignorance surcharge.
06.The Workhorse Forms: BCE and CCE
Binary cross-entropy (BCE / log loss). p = Bernoulli(y), q = Bernoulli(ŷ):
L = −[ y log ŷ + (1 − y) log(1 − ŷ) ]
with y in {0, 1}, ŷ = sigmoid(z) the predicted probability. One term dies per example: the truth pays −log ŷ if y = 1, −log(1−ŷ) if y = 0. Multi-label problems apply BCE independently per output.
Categorical cross-entropy (CCE). p = one-hot target class c, q = softmax(z):
L = −log q_c = −log( exp(z_c) / Σⱼ exp(zⱼ) ) = −z_c + logsumexp(z)
Multi-class with soft targets (distillation, mixup labels): L = −Σ pⱼ log qⱼ.
Both forms are negative log-likelihoods: BCE is the Bernoulli MLE objective, CCE the Categorical MLE objective (Topic 19). Classification losses are just MLE with the right output distribution. sklearn's log_loss and PyTorch's CrossEntropyLoss compute exactly this — note the trap that PyTorch's CrossEntropyLoss expects raw logits, not probabilities, and applies log-softmax internally with the numerically stable log-sum-exp trick (Topic 24).
For language models, average token CCE in nats converts to perplexity = exp(mean CCE) — the standard "effective branching factor" metric of next-token prediction.
import numpy as np
def softmax(z):
e = np.exp(z - z.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
z = np.array([2.0, 0.5, -1.0]) # logits for 3 classes
y = np.array([0, 1, 0]) # true class = 1
q = softmax(z)
loss = -(y * np.log(q)).sum()
print("probs:", q.round(4), " loss:", round(loss, 4))
# gradient of CE w.r.t. logits is (q - y) — the famous clean result
print("analytic grad:", (q - y).round(4))
eps = 1e-5
num = np.array([( -(y * np.log(softmax(z + eps * np.eye(3)[i]))).sum()
- -(y * np.log(softmax(z - eps * np.eye(3)[i]))).sum() ) / (2 * eps)
for i in range(3)])
print("numeric grad:", num.round(4))07.Why Softmax + Cross-Entropy Trains So Well
Differentiate CCE with respect to the logits and everything collapses:
∂L/∂z = q − y
"predicted probability vector minus one-hot target." That's it. No mess.
Derivation in one line: the log in the loss cancels the exponential in softmax, and the log-sum-exp derivative cancels the normalization — the softmax Jacobian (Topic 13) and CE's gradient annihilate each other's ugliness.
Sanity-check the bookie reading: if you quoted 0.78 on the winner (true class) but the winner needed 1.0, the gradient entry for that class is q − y = 0.78 − 1 = −0.22 → negative → gradient descent raises that logit. Errors in the odds get corrections proportional to their size.
Compare with MSE on softmax outputs: the gradient carries a factor q(1 − q) and becomes flat when the model is confidently wrong — a well-known pathological plateau. That is why:
- Binary classification: sigmoid + BCE, never MSE.
- Multi-class: softmax + CCE, never MSE.
- Gradient of the loss with respect to the last-layer weights is
(q − y)·features— a bounded, interpretable error signal.
The property extends: for any exponential-family output head (Bernoulli, Categorical, Gaussian, Poisson), NLL + natural parameterization gives gradient = (predictions − sufficient statistics of the target), which is why the "output unit = distribution" view generalizes cleanly.
08.In Practice: Variants and Failure Modes
The textbook loss meets production. Five things every practitioner knows:
- Label smoothing: replace one-hot p with
(1 − α)p + α/K uniform. Bounds the maximum achievable loss, discourages logit blow-up, and improves calibration and translation/RLHF performance (Inception's "label-smoothing regularized" training made it famous). - Class imbalance: weighted CE scales per-class terms; when that is not enough, go to focal loss (down-weights easy examples by
(1 − q_c)^γ) or resampled batches. - Numerical safety: never log a raw probability from a sigmoid/softmax that could round to 0 — fused log-softmax kernels and ε-clipping prevent
−log 0 = ∞loss. - Miscalibration: CE rewards pushing probabilities to extremes; post-hoc temperature scaling tunes one scalar on a validation set to match confidence to accuracy — which matters when CE-trained scores feed thresholds or cost-sensitive decisions.
- Per-sample CE is unbounded: a single label-noise example with high confidence can dominate a batch — the motivation behind robust losses and noise-aware objectives.
Back to the bookie: smoothing means "never post true 0-odds," clipping means "the ledger can't print ∞," and temperature means "re-mark your odds after the season so 0.9 really wins 90% of the time."
Pretraining loss is mean per-token categorical cross-entropy over a 100k-entry vocabulary softmax (often factorized or fused for speed). Everything downstream reports in this currency: loss curves in nats, validation perplexity via exp, and distillation objectives transfer teacher soft distributions as CE targets for the student to regress onto.
Staff+ Engineering Takeaways
- Cross-entropy H(p, q) is expected surprisal of p under q's code, and equals H(p) + KL(p ‖ q) — entropy plus disagreement penalty.
- It is asymmetric and minimized (for fixed p) exactly when q = p.
- BCE is Bernoulli NLL, categorical CE is Categorical NLL: classification losses are MLE with the right output distribution.
- Softmax + CE yields the razor gradient ∂L/∂z = q − y, while MSE-on-probabilities saturates — the reason this pairing is standard.
- Label smoothing, class weighting, fused log-softmax numerics, and temperature calibration are the practical engineering layer over the same objective.
- For language models, perplexity = exp(average token cross-entropy) — loss in nats directly reports effective choice size.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
With a fixed true distribution p, minimizing H(p, q) over q gives:
How clear and actionable was this distributed systems breakdown?