TOPIC #22Intermediate 12 min read

Cross-Entropy

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Cross-entropy is the surprise bill: your model sets the odds, reality picks the outcome, and you pay −log(probability you assigned to what actually happened). Minimizing it means being honestly surprised as little as possible — and with softmax the gradient collapses to the beautiful error signal q − y.

01.The Problem: You Need a Number for "How Wrong Are My Odds?"

Your model looks at a photo and says:

  • 70% cat
  • 25% dog
  • 5% bird

The truth is: dog.

How wrong is that? And how do you turn "wrongness" into a smooth number a gradient can chase (Topic 11)?

The obvious candidate — plain accuracy — fails immediately:

Insight

Accuracy is a count of right/wrong. It has zero gradient almost everywhere. You cannot train on it.

You need a loss that:

  • Punishes confident wrongness hard.
  • Rewards correct confidence gently.
  • Is smooth, so optimization works.

So the question becomes

Insight

Can we price each prediction by how surprised the model should be at reality?

Topic 21 already built the price tag: surprisal I = −log p. Cross-entropy just aims it at models.

02.The Idea in Plain Words: Surprise Under Someone Else's Code

Cross-entropy is simply

Insight

H(p, q) = − Σₓ p(x) log q(x) — the average surprisal of reality p, when your model q sets the prices.

Unpack the cast:

  • p is reality — the true distribution. In classification it is usually one-hot: the correct class has p = 1, everything else 0.
  • q is the model — its predicted probabilities.
  • −log q(x) is the price your model's code assigns to event x (Topic 21's telegraph rates, but your rates, not the optimal ones).

The three-way identity ties Topics 21 and 23 together:

H(p, q) = H(p) + KL(p ‖ q)

true cost + a non-negative penalty for q disagreeing with p. Consequences worth memorizing:

  • H(p, q) ≥ H(p) always; equality exactly when q = p.
  • Asymmetric: H(p, q) ≠ H(q, p) — coding p with q's code is not the same transaction as the reverse.
  • When p is a fixed one-hot label, everything cancels except the true class: H(p, q) = −log q(correct class). The loss is literally the surprisal of the right answer under the model. Minimizing it = raising the true label's probability.

03.A Simple Worked Example: Paying the Bill

Your model predicts: 70% cat, 25% dog, 5% bird. Reality collects: it was a dog.

Your bill:

L = −ln 0.25 ≈ 1.386 nats

Now change only the model's confidence, same truth:

code
P(dog) assigned   Bill −ln P(dog)
   0.99           0.010   ← nearly free
   0.25           1.386   ← hedged, moderate
   0.05           3.000   ← wrong AND bold
   0.001          6.908   ← ruinously confident nonsense
   0.0            ∞       ← −log 0: the model is deleted from reality

See the shape of the punishment:

  • Being wrong is expensive; being confidently wrong is unboundedly expensive.
  • Being right quietly (0.25 on the truth) already beats being confidently wrong on it.
  • That log-shaped squeeze is exactly why gradient descent pushes probabilities of true labels up fast at first, then carefully.

04.Visual Intuition: The −log Curve

The entire loss, drawn:

code
 loss
  ▲
  │╲
  │ ╲   −log q : infinite punishment
  │  ╲  at q = 0, free at q = 1
  │   ╲
  │    ╲__
  │       ╲____
  │            ╲─────___
  └──┬──────┬──────┬────╱──►  q = probability
  0.0     0.25   0.5    1.0    the model gave the truth

Two features do all the work:

  • The wall at q → 0. Any real chance the truth gets probability ~0 costs ∞. So the model can never "safely" exclude an outcome it might face.
  • The gentle floor at q → 1. Near-perfect predictions earn near-zero loss — diminishing returns on already-good confidence.

Compare the naive alternative, MSE on the probability: it treats "predicted 0.0 for the truth" as a bounded error of 1.0. The log says: that error is unbounded, and it is right — overconfident nonsense must be punished harder than ordinary mistake.

05.The Analogy: The Bookie Who Must Quote Every Outcome

Carry one analogy through the rest: your model is a bookie setting odds on tomorrow.

  • Before the race, the bookie publishes odds for every horse: those are the probabilities q. Rules: all positive, must sum to 1 (softmax guarantees this).
  • After the race, reality names the winner, and the bookie pays out −log q(winner).
  • A cautious bookie who spreads odds loses slowly. A bookie who quotes 1000:1 against the actual winner is bankrupt.

Training = a bookie who has run the same race millions of times and keeps adjusting odds to minimize the average payout. And the "true" optimal bookie quotes exactly the real frequencies: then average payout = H(p), entropy, the irreducible cost of the race itself (Topic 21). Anything worse is the KL penalty for a wrong belief (Topic 23).

This is why the identity H(p,q) = H(p) + KL(p‖q) is the bookie's ledger: fair price + ignorance surcharge.

06.The Workhorse Forms: BCE and CCE

Binary cross-entropy (BCE / log loss). p = Bernoulli(y), q = Bernoulli(ŷ):

L = −[ y log ŷ + (1 − y) log(1 − ŷ) ]

with y in {0, 1}, ŷ = sigmoid(z) the predicted probability. One term dies per example: the truth pays −log ŷ if y = 1, −log(1−ŷ) if y = 0. Multi-label problems apply BCE independently per output.

Categorical cross-entropy (CCE). p = one-hot target class c, q = softmax(z):

L = −log q_c = −log( exp(z_c) / Σⱼ exp(zⱼ) ) = −z_c + logsumexp(z)

Multi-class with soft targets (distillation, mixup labels): L = −Σ pⱼ log qⱼ.

Both forms are negative log-likelihoods: BCE is the Bernoulli MLE objective, CCE the Categorical MLE objective (Topic 19). Classification losses are just MLE with the right output distribution. sklearn's log_loss and PyTorch's CrossEntropyLoss compute exactly this — note the trap that PyTorch's CrossEntropyLoss expects raw logits, not probabilities, and applies log-softmax internally with the numerically stable log-sum-exp trick (Topic 24).

For language models, average token CCE in nats converts to perplexity = exp(mean CCE) — the standard "effective branching factor" metric of next-token prediction.

python— Softmax + categorical cross-entropy from scratch, with gradient check
import numpy as np

def softmax(z):
    e = np.exp(z - z.max(axis=-1, keepdims=True))
    return e / e.sum(axis=-1, keepdims=True)

z = np.array([2.0, 0.5, -1.0])          # logits for 3 classes
y = np.array([0, 1, 0])                 # true class = 1
q = softmax(z)
loss = -(y * np.log(q)).sum()
print("probs:", q.round(4), " loss:", round(loss, 4))

# gradient of CE w.r.t. logits is (q - y) — the famous clean result
print("analytic grad:", (q - y).round(4))
eps = 1e-5
num = np.array([( -(y * np.log(softmax(z + eps * np.eye(3)[i]))).sum()
                 - -(y * np.log(softmax(z - eps * np.eye(3)[i]))).sum() ) / (2 * eps)
                for i in range(3)])
print("numeric grad:", num.round(4))

07.Why Softmax + Cross-Entropy Trains So Well

Differentiate CCE with respect to the logits and everything collapses:

∂L/∂z = q − y

"predicted probability vector minus one-hot target." That's it. No mess.

Derivation in one line: the log in the loss cancels the exponential in softmax, and the log-sum-exp derivative cancels the normalization — the softmax Jacobian (Topic 13) and CE's gradient annihilate each other's ugliness.

Sanity-check the bookie reading: if you quoted 0.78 on the winner (true class) but the winner needed 1.0, the gradient entry for that class is q − y = 0.78 − 1 = −0.22 → negative → gradient descent raises that logit. Errors in the odds get corrections proportional to their size.

Compare with MSE on softmax outputs: the gradient carries a factor q(1 − q) and becomes flat when the model is confidently wrong — a well-known pathological plateau. That is why:

  • Binary classification: sigmoid + BCE, never MSE.
  • Multi-class: softmax + CCE, never MSE.
  • Gradient of the loss with respect to the last-layer weights is (q − y)·features — a bounded, interpretable error signal.

The property extends: for any exponential-family output head (Bernoulli, Categorical, Gaussian, Poisson), NLL + natural parameterization gives gradient = (predictions − sufficient statistics of the target), which is why the "output unit = distribution" view generalizes cleanly.

08.In Practice: Variants and Failure Modes

The textbook loss meets production. Five things every practitioner knows:

  • Label smoothing: replace one-hot p with (1 − α)p + α/K uniform. Bounds the maximum achievable loss, discourages logit blow-up, and improves calibration and translation/RLHF performance (Inception's "label-smoothing regularized" training made it famous).
  • Class imbalance: weighted CE scales per-class terms; when that is not enough, go to focal loss (down-weights easy examples by (1 − q_c)^γ) or resampled batches.
  • Numerical safety: never log a raw probability from a sigmoid/softmax that could round to 0 — fused log-softmax kernels and ε-clipping prevent −log 0 = ∞ loss.
  • Miscalibration: CE rewards pushing probabilities to extremes; post-hoc temperature scaling tunes one scalar on a validation set to match confidence to accuracy — which matters when CE-trained scores feed thresholds or cost-sensitive decisions.
  • Per-sample CE is unbounded: a single label-noise example with high confidence can dominate a batch — the motivation behind robust losses and noise-aware objectives.

Back to the bookie: smoothing means "never post true 0-odds," clipping means "the ledger can't print ∞," and temperature means "re-mark your odds after the season so 0.9 really wins 90% of the time."

Production Implementation in Big Tech
LLM pretraining at scale (GPT-class models)• Next-token cross-entropy as the universal training signal

Pretraining loss is mean per-token categorical cross-entropy over a 100k-entry vocabulary softmax (often factorized or fused for speed). Everything downstream reports in this currency: loss curves in nats, validation perplexity via exp, and distillation objectives transfer teacher soft distributions as CE targets for the student to regress onto.

Staff+ Engineering Takeaways

  • Cross-entropy H(p, q) is expected surprisal of p under q's code, and equals H(p) + KL(p ‖ q) — entropy plus disagreement penalty.
  • It is asymmetric and minimized (for fixed p) exactly when q = p.
  • BCE is Bernoulli NLL, categorical CE is Categorical NLL: classification losses are MLE with the right output distribution.
  • Softmax + CE yields the razor gradient ∂L/∂z = q − y, while MSE-on-probabilities saturates — the reason this pairing is standard.
  • Label smoothing, class weighting, fused log-softmax numerics, and temperature calibration are the practical engineering layer over the same objective.
  • For language models, perplexity = exp(average token cross-entropy) — loss in nats directly reports effective choice size.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

With a fixed true distribution p, minimizing H(p, q) over q gives:

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?