Logarithms & Exponentials
A computer cannot hold the product of hundreds of tiny probabilities — it silently rounds to zero. Logarithms fix this by turning those products into sums, and the log-sum-exp trick is the one maneuver that makes softmax and likelihoods overflow-proof. Almost all of ML numerics is built on this pair.
The Log-Space Survival Path
Probabilities multiply and underflow; logits add and stay in range. The log-sum-exp trick is the one maneuver that turns the naive computation into the fused kernels every ML library actually ships.
01.The Problem: Tiny Numbers Multiply Into Silent Zero
Your language model scores a sentence, token by token.
Each token gets a probability, say around 0.001.
A 300-token sentence? Its likelihood is the product of 300 such numbers:
0.001 × 0.001 × ... × 0.001 (300 times)
That is 10^(−900). The computer's biggest negative number is only about 10^(−308).
So the hardware does something terrifying: it returns exactly 0. No error. No warning you can act on. Two different sentences now both score 0, and your ranking logic collapses.
So the question becomes
How do we do arithmetic with numbers this small without the machine lying to us?
Answer: don't multiply them. Add their logs instead. This topic is the whole toolkit.
02.The Idea in Plain Words: A Log Undoes an Exponent
A logarithm answers one question:
log_b(x) = y means "what exponent do I put on b to produce x?" — i.e. b^y = x.
- The natural log ln uses base e ≈ 2.71828, and logₑ x is just written log x in ML code.
- Why base e? The unique base whose exponential has itself as derivative — softmax derivatives and ODE-based models (normalizing flows, continuous-time nets) all prefer it.
The rewrite rules are the entire daily toolkit:
- Product to sum:
log(ab) = log a + log b— the reason likelihoods (huge products) become sums. - Quotient to difference:
log(a/b) = log a − log b. - Power to multiple:
log(aᵏ) = k·log a— exponents inside densities (Gaussian squares!) become scalars. - Inverse functions:
exp(log x) = xandlog(exp z) = z. - Change of base:
log_b x = ln x / ln b— bits ↔ nats conversion (divide bits by ln 2… multiply nats by 1/ln 2). - Shape facts:
log(1) = 0, log grows without bound but slowly,x → 0⁺gives log → −∞, and log is monotone increasing — so argmax is preserved under log while sums of products become sums of logs.
Exponentials are the mirror image: exp(a + b) = exp(a)·exp(b), strictly positive, convex, and growing faster than any polynomial.
03.A Simple Worked Example: 200 Coins, One Ledger
Take the disaster from section 1 and shrink it: 200 probabilities of 0.01 each.
- Linear world:
0.01^200 = 10^(−400)→ the float64 register bottoms out near 2.2e−308 → machine says 0. Both this sentence and any longer one score "0.0". Information destroyed. - Log world:
log 0.01 = −4.605, so the whole sentence is200 × (−4.605) = −921. An ordinary, comparable number. A sentence scoring −900 beats one scoring −921 — the ranking survives.
Two more hand-calculations to own the idea:
log₂ 8 = 3because 2³ = 8 — "8 needs three bits." (This is Topic 21's surprisal in action.)- Compare
p₁ = 0.002vsp₂ = 0.001: logs give −6.21 vs −6.91. p₁ is bigger, and its log is bigger too. Order preserved — that is the monotone fact from section 2, and it is why argmax works in log space.
codelinear space: 0.01 · 0.01 · ... · 0.01 = 0 (lie) ↓ take log of each, add ↓ log space: (−4.61) + (−4.61) + ... = −921 (truth)
04.Visual Intuition and the Analogy: A Ruler for Astronomical Ratios
Logs compress multiplicative distance into additive distance — like the Richter scale for earthquakes:
codelinear ruler: 1 10 1000 10^900 |---------|--------------|----------------------X (off the page) log ruler: 0 1 2 ... 900 |----|----|----|----|----| ... |----| | every ONE mark = one factor of TEN in reality
Carry one analogy through the rest: a ledger that charges interest.
- Multiplying probabilities = compounding. Painful to track, and the balance hits zero.
- Logs = converting the ledger to "interest units," where every event is a small negative number you simply add.
- Back to reality whenever needed = one exp() call.
So the daily ML loop is: live in the ledger (log space), publish in reality (probabilities) only at the end. Everything below — logits, log-likelihood, log-sum-exp, log-partition — is bookkeeping on that ledger.
05.Why Everything in ML Lives in Log Space
Floating point has brutal finite ranges: float32 holds about 1.2e−38 to 3.4e38, float64 bottoms out near 2.2e−308. A sentence-level likelihood from an LM is a product of hundreds of token probabilities — each around 1e−3 — so 1e−300 territory underflows to exactly zero, silently.
Logarithms rescue every part of the pipeline:
- Products become sums: the log-likelihood of m samples is
Σ log p(xᵢ)— no giant product, no zeros. - Order preservation:
argmax p = argmax log p, so softmax over logits gives the same decisions as probabilities, and comparisons work on log-odds. - Bounded-looking parameters:
sigmoid(z) = 1/(1 + e^(−z))maps the whole real line to (0, 1); its inverse is the logit/log-oddslog(p/(1−p)). Logistic regression, attention gates, and probability heads are all defined this way: models emit unconstrained scores, the exponential interprets them as evidence ratios. - Temperature and scaling:
p ∝ exp(z/T)— softmax with temperature is literally exponentials of scaled logits (Boltzmann distributions in statistical physics are the same construction). - Loss curves, learning-rate grids, and multiplicative averaging all read naturally in logs; the geometric mean = exp(mean of logs).
Standard vocabulary, in ledger terms:
- "log probability" — the balance in interest units: log p ∈ (−∞, 0].
- "logit" — an unconstrained real score, before normalization.
- "log-likelihood" — the sum of log probs over the data.
06.The Log-Sum-Exp Trick: The One Hard Maneuver
One thing the ledger cannot do naively: add. Softmax's normalizer needs log(Σ exp(zⱼ)) — you must sum exponentials, but the exponentials are the very numbers that overflow. exp(1000) = inf, and inf/inf returns NaN even though the true answer is finite.
The trick: factor out the largest term m = maxⱼ zⱼ:
log Σⱼ exp(zⱼ) = m + log Σⱼ exp(zⱼ − m)
Why it is bulletproof, in three short facts:
- Every shifted exponent is ≤ 0, so each exp ≤ 1 → no overflow.
- At least one term (the max itself) equals exactly 1, so the inner sum ≥ 1 → no log-of-zero.
- The identity is exact algebra, not an approximation — pulling exp(m) out of the sum and moving it outside the log.
Work it: z = (1000, 1000.5, 999). m = 1000.5. Shifted: (−0.5, 0, −1.5) → exps (0.607, 1, 0.223) → sum 1.830 → log ≈ 0.604 → answer 1001.104. The naive route gives inf.
This is why:
torch.logsumexpandscipy.special.logsumexpexist as primitives.- Frameworks ship fused log-softmax kernels computing
z − logsumexp(z)in one pass. - Cross-entropy is computed from logits, never from probabilities (the PyTorch trap that
CrossEntropyLossexpects raw scores — Topic 22). - Mixture-model weights, HMM forward algorithms, and VAE ELBO normalizers are all logsumexp pipelines.
import numpy as np
from scipy.special import logsumexp
def stable_logsumexp(z):
m = np.max(z)
return m + np.log(np.sum(np.exp(z - m)))
z = np.array([1000.0, 1000.5, 999.0])
naive = np.log(np.sum(np.exp(z))) # inf / RuntimeWarning
print("naive:", naive)
print("stable:", stable_logsumexp(z)) # 1001.1044 = max(1000.5) + log(sum exp(z - m))
print("scipy :", logsumexp(z)) # identical
# log-softmax in one line, safe at any scale
log_q = z - logsumexp(z)
print("log-softmax:", log_q.round(4))
print("probs sum:", np.exp(log_q).sum().round(6))07.In Practice: Exponentials on the Output Side
Where logs read evidence, exponentials manufacture structure. The ledger in reverse:
- Softmax / Boltzmann: exp turns real scores into positive numbers; the normalization makes them sum to one — the canonical construction of a distribution from unconstrained energy-style logits.
- Positive parameters: variances, rates, and scales are parameterized as
exp(θ)so an optimizer can wander the whole real line while the model stays legal (σ = exp(log σ)). - Growth and decay: exponential moving averages (the β coefficients in Adam, batch-norm running stats, and weight-averaging) are discrete solutions of decay equations; a half-life maps directly to a log base:
steps = ln(0.5)/ln(β). - Exponential family:
log p(x|θ) = ⟨T(x), θ⟩ − A(θ)— every distribution in Topic 16's list fits this pattern, with A(θ) (the logsumexp of the continuous world) generating all moments. - Continuous-time and flow models: normalizing flows and neural ODEs use log|det J| accumulation (Topic 13 meets exponentials) to transform densities multiplicatively.
Numerical hygiene to keep, as ledger rules:
- Prefer log-space APIs (logsumexp, log-softmax, Bernoulli log-probs).
- Clip only log quantities, and only with a reason.
- When you must convert back, remember exp grows fast — a "recovered probability" above 1 is always a lost logsumexp somewhere.
Decoder beam search keeps every hypothesis score as a sum of log-probabilities rather than a product of probabilities: hundreds of chained 0.001-scale token scores would underflow to zero in linear space, collapsing ranking ties. Comparisons, pruning, and rescoring run on log-probs; final confidence is recovered with exp only at the end — and confusion-matrix-style scoring systems add explicit logsumexp merges whenever a hypothesis can be reached through multiple paths.
Staff+ Engineering Takeaways
- Logs convert products to sums, powers to multiples, and preserve argmax — the reason likelihoods, entropies, and losses are defined in log space.
- Sigmoid/exponentials map unconstrained logits to legal probabilities; models emit scores, exp + normalization interprets them.
- float32/float64 ranges force log space: products of hundreds of small probabilities underflow to silent zeros.
- logsumexp(z) = max(z) + log Σ exp(z − max(z)) is exact, overflow-proof, and the core of fused log-softmax kernels.
- exp parameterizes positive quantities (σ, rates), drives EMAs and decay schedules, and defines softmax/Boltzmann distributions.
- Exponential-family log-densities are linear in natural parameters minus a logsumexp-style log-partition function — the deep unification of Topic 16.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Why can you replace argmax p with argmax log p in any classifier or beam search?
How clear and actionable was this distributed systems breakdown?