Bayes' Theorem
You usually know the probability of the evidence given a hypothesis; you want the probability of the hypothesis given the evidence. Bayes' theorem flips one into the other — prior × likelihood over evidence — and the medical-test example shows why ignoring base rates fools almost everyone.
The Bayesian Update Loop
Bayes' theorem converts the probability of data given a hypothesis into the probability of the hypothesis given data. Each posterior becomes tomorrow's prior, producing a continuous learning cycle.
01.The Problem: The Number You Know Is Not the Number You Want
A doctor says:
"This test is 99% accurate for the disease."
Translation: if you have the disease, the test says "positive" 99% of the time. Written as a conditional probability (Topic 15: conditioning means "given that this already happened"):
P(positive | sick) = 0.99
But you, holding a positive result, want a completely different number:
P(sick | positive) = ???
Those two look similar. They are not similar — and confusing them has ruined diagnoses, convictions, and fraud-alert systems.
The world hands us evidence-shaped numbers:
- P(symptom | disease)
- P(evidence | suspect)
- P(clicks | user is in variant B)
We need hypothesis-shaped numbers back:
- P(disease | symptom)
- P(suspect | evidence)
- P(variant B | clicks)
How do you legally walk from "data given the story" to "the story given the data"?
That walk is Bayes' theorem — a one-line piece of algebra with a philosophy attached.
02.The Idea in Plain Words: The Belief Inversion Machine
Start from the definition of conditional probability: P(A | B) = P(A ∩ B) / P(B) (read "A given B equals the probability of both happening, divided by the probability of the given part").
Write P(A ∩ B) — "both happen" — two ways: as P(B|A)·P(A) and as P(A|B)·P(B). Set them equal and rearrange:
Bayes' theorem:
P(A | B) = P(B | A) · P(A) / P(B)
In inference language, A is a hypothesis (a parameter, a class, a cause) and B is the observed data:
- Prior
P(hypothesis): how plausible you found it before the evidence. - Likelihood
P(data | hypothesis): how expected the evidence is if the hypothesis were true. (The number you usually have!) - Evidence
P(data) = Σover all hypotheses ofP(data | h)·P(h): how often the data shows up at all — a normalizer that makes the posterior sum to 1. - Posterior
P(hypothesis | data): the updated belief. (The number you wanted!)
Read it as a recipe: posterior ∝ likelihood × prior, tidy up with the evidence.
The theorem is pure algebra, but its power is directional: data collectors reason hypothesis → data; decision makers need data → hypothesis. Bayes is the toll bridge between the two.
03.A Tiny Worked Example: The 99% Test That Means Only 33%
A disease affects 0.5% of people (prior = 0.005). A test has 99% sensitivity (P(positive | sick) = 0.99) and a 1% false-positive rate (P(positive | healthy) = 0.01).
You test positive. Are you ~99% likely to be sick?
Do it the human way first — imagine 1000 people:
| group | size | test positive |
|---|---|---|
| sick | 5 | 0.99 × 5 ≈ 4.95 |
| healthy | 995 | 0.01 × 995 ≈ 9.95 |
Total positives ≈ 14.9, and only 4.95 of them are sick:
P(sick | positive) ≈ 4.95 / 14.9 ≈ 33%
Now the formula, same numbers:
- Numerator for "sick":
0.99 × 0.005 = 0.00495. - Numerator for "healthy":
0.01 × 0.995 = 0.00995. - Evidence:
P(positive) = 0.00495 + 0.00995 = 0.0149. - Posterior:
P(sick | positive) = 0.00495 / 0.0149 ≈ **33%**.
Even a highly accurate test leaves most positives healthy, simply because the healthy population is so much larger. The 1% false-positive rate applies to 995 people; the 99% sensitivity applies to only 5.
Rare-disease screening, fraud detection, and anomaly alerts all live in this regime: precision depends on the base rate, which is why alert triage in production systems uses multiple signals (and thus repeated Bayesian updating) before escalating.
prevalence = 0.005
sensitivity = 0.99 # P(positive | sick)
false_positive = 0.01 # P(positive | healthy)
p_pos = sensitivity * prevalence + false_positive * (1 - prevalence)
p_sick_given_pos = sensitivity * prevalence / p_pos
print(f"P(positive) = {p_pos:.4f}")
print(f"P(sick | positive) = {p_sick_given_pos:.2%}") # ~33%, not 99%
# Second independent positive test updates again: prior becomes 1/3
prior2 = p_sick_given_pos
p_pos2 = sensitivity * prior2 + false_positive * (1 - prior2)
print(f"After 2nd positive: {sensitivity * prior2 / p_pos2:.2%}")04.Visual Intuition: Two Streams Joining
Picture P(positive) as a river fed by two streams:
codesick (0.5% of people) healthy (99.5%) │ │ 99% test positive 1% test positive │ │ 0.00495 ──────┐ ┌────── 0.00995 ▼ ▼ P(positive) = 0.0149 │ "You are a random drop in THIS river." Chance your drop came from the sick stream = 0.00495 / 0.0149 ≈ 33%
The trick of the whole picture:
- The left stream is strong (99% of its tiny flow passes) but the flow is small (only 0.5% of people).
- The right stream is weak (1% leakage) but enormous (99.5% of people).
- Weak × huge can beat strong × tiny. Always. That's the base rate speaking.
A positive result doesn't tell you "sick". It tells you "you're a drop in this river" — and Bayes just measures where river drops usually come from.
05.The Analogy: The Detective and the Suspect Board
Carry one analogy through the rest: a detective with a corkboard of suspects.
- Behind each photo, the detective keeps a weight — how likely this person did it right now. Those weights are the priors. They start from base rates: in a village of 100 plausible suspects, nobody walks in at 99%.
- A clue arrives — say, the culprit left muddy size-45 footprints. The detective asks each suspect one question: "if you were the culprit, how expected would this clue be?" That is the likelihood,
P(clue | suspect). - Each weight gets multiplied by its likelihood; the weights are renormalized to sum to 1 (the evidence step). The updated weights are the posterior.
- Then — and this is the beautiful part — the next clue uses the new weights as its starting point. Today's posterior is tomorrow's prior. The board keeps updating; that loop IS learning.
Watch the base rate in action on the board: if 99 villagers wear size 45, the footprint barely moves any weight — the clue was expected anyway. If only one does, weights swing hard. Evidence only teaches you something when it discriminates.
Everything else in this topic — conjugate priors, naive Bayes, Bayesian neural networks — is just bookkeeping for one detective, or for millions of suspects at once.
06.Conjugacy: When Updating Is Just Pencil Work
Sometimes the math collapses deliciously: when the prior and posterior land in the same family, the prior is conjugate and updating becomes parameter bookkeeping. The flagship case is Beta–Binomial:
Prior: p ~ Beta(a, b). Observe k successes in n Bernoulli trials. Posterior: p ~ Beta(a + k, b + n − k).
That's it — count heads into the first parameter, tails into the second. No integrals.
The prior parameters act like pseudo-counts: imaginary flips you bring to the table.
Beta(1, 1)is uniform — "no data yet" — and adding it is exactly Laplace smoothing.- A
Beta(2, 2)prior that says "probabilities near 0.5 are a bit more likely" nudges a coin that came up 10/10 heads away from claiming p = 1: the posterior is Beta(12, 2), not Beta(10, 0).
Other common conjugate pairs: Gaussian-Gaussian (mean known/unknown), Dirichlet-Categorical (topic models), Gamma-Poisson (rates), Normal-Inverse-Gamma (full Gaussian parameters).
Conjugacy is why Bayesian toolkits for counts and rates are closed-form and fast enough for online ranking systems — the detective updating the board in microseconds per request.
07.Why AI Cares: Bayes Across Machine Learning
- Naive Bayes classifiers: for text and tabular spam/sentiment work, the class posterior is proportional to prior × product of per-feature conditional probabilities (the "naive" conditional-independence assumption: the detective treats each word as an independent clue). Training is counting; inference is a log-sum of probabilities — and it still beats more complex models on tiny datasets.
- MAP estimation (Topic 20): maximizing the posterior instead of the likelihood; regularization is the prior in disguise.
- Bayesian neural networks: put priors over weights and reason about posterior predictive uncertainty — frameworks like Laplace approximations and ensemble distillation bring this to production at low cost.
- Generative vs discriminative modeling: generative models learn
P(x, y)and deriveP(y | x)with Bayes; discriminative models learnP(y | x)directly. - Sequential decision systems: A/B-test analyzers (posterior probability of uplift), Thompson-sampling bandits, and recalibrated alerting all run the update loop in real time — the corkboard, spinning.
In practice: every time your model outputs a probability and your threshold fires an alert, a base rate is hiding behind that number. Ask what prior produced it, and what evidence would move it.
Classic spam filters learn P(word | spam) and P(word | ham) from counted corpora, then classify each incoming message by summing log-likelihood ratios against the spam prior. Production fraud systems do the same at request time: a device-risk signal raises the fraud posterior only after being combined with the ~0.1% base rate of fraud per transaction, preventing alert storms from high false-positive rates on huge traffic.
Staff+ Engineering Takeaways
- Bayes' theorem inverts conditioning: posterior = likelihood × prior / evidence, and the evidence normalizes over all hypotheses.
- Base rates matter: with rare conditions or rare fraud, even accurate tests produce mostly false positives (precision collapses without prevalence).
- Posterior from one step becomes the prior for the next — Bayesian inference is a loop, not a one-shot formula.
- Conjugate priors (Beta-Binomial, Dirichlet-Categorical, Gaussian-Gaussian) make updating closed-form; prior parameters are pseudo-counts.
- Naive Bayes, MAP estimation, Bayesian neural networks, and generative models are all direct applications of the same inversion.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
A test is 99% sensitive and 1% false-positive for a disease affecting 0.5% of people. P(sick | positive) is closest to:
How clear and actionable was this distributed systems breakdown?