Maximum Likelihood Estimation (MLE)
You have data and a model with knobs. MLE says: turn the knobs to the setting that makes the data you actually saw most plausible. That one principle produces the frequency estimate, MSE, cross-entropy — and every neural-net loss — while carrying a few famous failure modes.
The MLE Pipeline
Model → likelihood over the dataset → log → maximize. The negative log-likelihood of the chosen output distribution is exactly the loss function a trainer minimizes.
01.The Problem: Your Data Showed Up — Which Setting Caused It?
You borrow a coin from a stranger.
You flip it 10 times. It lands heads 9 times.
The coin has some true heads-probability p — a hidden knob between 0 and 1. You'll never see p directly. You only have the flips.
Which value of p makes "9 heads out of 10" the least surprising story?
Your gut already answers: "probably around 0.9, not 0.5."
That gut move — pick the knob setting under which the observed data looks most plausible — has a name: maximum likelihood estimation, or MLE.
And it is not just party-trick statistics:
- Fitting the mean of your error distribution? MLE.
- Training a classifier with cross-entropy? MLE.
- Pretraining ChatGPT-style LLMs token by token? MLE at trillion-token scale.
This topic makes the gut move precise, derives its famous closed forms, and lists where it betrays you.
02.The Idea in Plain Words: Likelihood Is Probability with the Roles Swapped
Given a model p(x | θ) and an observation, probability asks:
"How likely is this data, under fixed parameters?"
The likelihood function L(θ) = p(x | θ) keeps the data frozen and asks the reverse:
"Which parameters make this data most plausible?"
Same formula, different variable. Probability varies the data; likelihood varies the knobs.
But the consequences differ:
- Probability is normalized over data — it integrates to 1.
- Likelihood is not normalized over θ. So "the maximum likelihood value" has no probability meaning — only comparisons between settings of the same knobs, on the same data, matter.
For m i.i.d. samples (Topic 15: independent, identical draws), the joint likelihood factorizes:
L(θ) = Π p(xᵢ | θ)
MLE (θ̂) = the θ that maximizes this product. That's the entire method.
03.A Tiny Worked Example: The 9-Heads Coin, Scored
Back to the coin. Data: k = 9 heads, 1 tail, n = 10. Ignoring the ordering count (it doesn't depend on p), the likelihood of a candidate p is:
L(p) = p⁹ · (1−p)¹
Score a few suspects:
| candidate p | L(p) = p⁹(1−p) | verdict |
|---|---|---|
| 0.5 (fair coin) | 0.5¹⁰ ≈ 0.001 | looks bad |
| 0.6 | ≈ 0.004 | better |
| 0.9 | ≈ 0.039 | best |
| 0.99 | ≈ 0.009 | too greedy — that one tail protests |
The winner is p̂ = 0.9 = k/n. Satisfying — and later we'll derive that the MLE of a Bernoulli is always the observed frequency.
Why we take the log instead. Two reasons, both practical:
- Products of probabilities underflow — a real 1000-token sentence with per-token p ≈ 0.01 gives 0.01¹⁰⁰⁰ = 10⁻²⁰⁰⁰, which a computer stores as plain zero. Logs turn it into a sum:
−4.6 × 1000 ≈ −4600. Fine. - Products differentiate badly; sums don't. And log is monotone, so:
ℓ(θ) = Σ log p(xᵢ | θ)
Maximizing the log-likelihood ℓ is identical to maximizing L — same winner, friendlier math. Divide by m and it becomes the negative of the empirical average log-density — an expectation estimated from data (Topic 17). That is the exact bridge from statistics to empirical risk minimization: your training loss is an average log-likelihood with a minus sign.
04.The Analogy: The Detective's Lineup
Carry one analogy (and note it's the mirror of Topic 18's detective — there the evidence updated suspects' weights; here the evidence ranks the suspects once):
A crime happened (= your dataset). The station holds a lineup of suspects (= candidate parameter values θ).
- Probability asks each suspect: "would you have produced this crime scene?" and scores them:
p(data | θ). - The detective cannot ask "how likely is the suspect to exist" (no prior over θ — that's the Bayesian's job, and Topic 20 blends the two views).
- So MLE does the humble thing: pick the suspect whose story fits the evidence best. No humility about error bars, no runner-up credit — just the argmax.
Three consequences the detective's case file warns about, developed in section 7:
- With a tiny amount of evidence, the best-fitting suspect can be wildly wrong (overfitting: memorize 3 flips of 3 heads and you accuse p = 1.0).
- A suspect can fake perfect fit by shrinking their story to a razor's width around the evidence (degenerate likelihoods: σ → 0).
- The lineup itself may contain no real culprit (misspecification: you'll still pick someone).
05.Three Derivations You Should Own
Bernoulli (biased coin). Data: k heads in n flips. ℓ(p) = k log p + (n−k) log(1−p). Setting the derivative k/p − (n−k)/(1−p) = 0 gives the MLE p̂ = k/n — the observed frequency. Our 9/10 example, now proven.
Gaussian. Data: x₁ … xₘ. Differentiating ℓ(μ, σ²) yields μ̂ = (1/m) Σ xᵢ (the sample mean) and σ̂² = (1/m) Σ (xᵢ − μ̂)² — the biased variance with denominator m, not m−1. The MLE deliberately trades unbiasedness for likelihood optimality (the code block below shows both numbers side by side).
Linear regression. Assume y = wᵀx + ε with ε ~ N(0, σ²). The log-likelihood contains the term −Σ (yᵢ − wᵀxᵢ)² / (2σ²), so maximizing likelihood over w is exactly minimizing the sum of squared errors. Normal equations = MLE in closed form. Same logic: logistic regression's MLE minimizes binary cross-entropy; a softmax head's MLE minimizes categorical cross-entropy.
So the "loss zoo" isn't a zoo. Every standard loss is one likelihood calculation with the furniture rearranged.
import numpy as np
rng = np.random.default_rng(42)
x = rng.normal(loc=3.0, scale=2.0, size=200)
mu_hat = x.mean()
sigma2_mle = x.var() # denominator m: the MLE
sigma2_unbiased = x.var(ddof=1) # denominator m-1
print(f"true sigma^2 = 4.0")
print(f"MLE sigma^2 = {sigma2_mle:.3f} (biased low by factor {(len(x)-1)/len(x):.3f})")
print(f"Bessel sigma^2 = {sigma2_unbiased:.3f}")06.The Deepest View: MLE = Minimum KL Divergence to the Data
There is a way to see what MLE is really doing. Divide the log-likelihood by m and add a constant to see
(1/m) Σ log p(xᵢ | θ) ≈ E_{p̂_data}[log p_θ] = H(p̂_data) − KL(p̂_data ‖ p_θ)
where p̂_data is the empirical distribution of the sample (put one grain of probability mass on each observed point).
Since the data's own entropy does not depend on θ, maximizing likelihood is minimizing the KL divergence (Topic 23) from the empirical distribution to the model.
Plain words: training is fitting the model distribution to the data distribution — which is precisely the objective of every generative model.
Large-sample theory then makes MLE the reference estimator (as m → ∞, with identifiable, well-specified models):
- Consistency: θ̂ → true θ as data grows.
- Asymptotic normality:
√m(θ̂ − θ) → N(0, I(θ)⁻¹), where I(θ) is the Fisher information — the curvature of the expected log-likelihood. - Efficiency (Cramér–Rao): no unbiased estimator has asymptotic variance lower than
I(θ)⁻¹. MLE hits the theoretical noise floor: you cannot extract more information from the data than MLE extracts.
07.In Practice: Failure Modes — What MLE Gets Wrong
The detective picks the best story, not the true story. Five named ways that bites:
- Overfitting / extreme estimates. With zero smoothing, MLE assigns probability 1 to every label seen in a small training set and 0 to unseen ones — the discrete MLE memorizes. (Classic: observing 0 days with incidents in m days gives hazard-rate MLE = 0. No accident ever? The model says impossible.)
- Degenerate likelihoods. In Gaussian mixture models, a component collapsing (σ → 0) onto a single data point sends the likelihood to infinity; unbounded likelihood requires priors or constraints.
- Bias. The Gaussian variance MLE is biased by (m−1)/m; bias vanishes asymptotically but matters in small samples.
- Misspecification. If no θ matches the true distribution, MLE converges to the KL projection of reality onto the model family — a defensible answer, but only to the best available wrong model.
- No uncertainty output. MLE returns a point; confidence comes separately from Fisher information, bootstrap, or the posterior (Topic 20).
The repair kit has become daily practice: add a prior or penalty and MLE becomes MAP / regularized training (weight decay is exactly this); add smoothing for categorical heads; check likelihood pathologies when training loss spikes.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Unified derivation principle: MSE, BCE, categorical CE, Poisson NLL are all MLE under declared noise models.
- Optimally efficient asymptotically: variance cannot be beaten by any unbiased estimator (Cramér–Rao).
- Invariance: the MLE of a function of θ is that function of the MLE.
Trade-offs & Constraints
- No regularization built in — small data produces overconfident or degenerate estimates.
- Sensitive to model misspecification and outliers (Gaussian-MLE mean is dragged by extremes).
- Point estimate only; uncertainty requires Fisher information, bootstrap, or Bayesian machinery.
A transformer LM defines a Categorical distribution over the vocabulary for each next token. Pretraining minimizes the average next-token cross-entropy, which is exactly the negative log-likelihood of the corpus under the model — i.e., MLE at trillions-of-tokens scale. Loss spikes in training are likelihood pathologies in real time, and temperature sampling manipulates the fitted distribution at inference.
Staff+ Engineering Takeaways
- Likelihood treats data as fixed and parameters as variables; it is not normalized, so always work with the log-likelihood.
- Classic closed forms: Bernoulli MLE = observed frequency; Gaussian MLE = sample mean and the m-denominator (biased) variance.
- MSE, binary cross-entropy, and categorical cross-entropy are negative log-likelihoods of Gaussian, Bernoulli, and Categorical heads — MLE is why they are the losses.
- Maximizing likelihood equals minimizing KL(empirical data distribution ‖ model), the objective of every generative model.
- MLE is consistent, asymptotically normal and efficient, but overfits small data and gives no uncertainty by itself — the motivation for MAP.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Training linear regression by minimizing MSE is equivalent to maximum likelihood under which assumption?
How clear and actionable was this distributed systems breakdown?