TOPIC #16Beginner 11 min read

Probability Distributions

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

A probability distribution is a named shape that matches a data story: counting rare arrivals? Poisson. Averaged noise? Gaussian. Uncertainty about a probability itself? Beta. This tour covers the distributions every ML practitioner must know — and the deep fact that choosing an output distribution is choosing a loss function.

Choosing a Distribution

A practical selection map: the structure of your data (discrete vs continuous, positive-only, count vs duration) strongly constrains which family belongs in your model.

Choosing a Distribution
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: Your Data Has a Shape — Do You Have a Name for It?

Imagine three boxes of numbers sitting in front of you.

  • Box 1: whether each visitor clicked a banner — lots of 0s, a few 1s.
  • Box 2: support tickets that arrived each hour — 2, 5, 1, 4, 0, 3…
  • Box 3: exact page-load times in milliseconds — 203.1, 188.7, 900.2…

So far you know random variables (Topic 15 — the scorekeeper rules that turn messy outcomes into numbers).

But a variable alone doesn't tell you what's likely.

Insight

Which shapes of data are possible? Which are rare?

Answering that is the job of a probability distribution: the official map of "what can happen, and how often."

Statisticians noticed something amazing centuries ago: reality keeps reusing the same handful of shapes.

Clicks, tickets, latencies, heights, errors, word frequencies — almost all of them fit one of a few named distributions.

Learning the names is like learning tool shapes in a workshop: once you can recognize the job, the right tool jumps out.

02.The Idea in Plain Words: A Distribution Is a Data Story with a Formula

Insight

A probability distribution says how probability is spread over the possible values of a random variable.

For discrete variables it's the PMF list of masses (they sum to 1); for continuous ones it's the PDF whose areas are probabilities — the sand piles vs sand beds from Topic 15.

But "distribution" in practice means something more useful:

Insight

It's a story about how the data was generated, compressed into a formula with a couple of knobs (parameters).

Each story has a signature:

  • "One yes/no event with probability p" → a Bernoulli story.
  • "Count of successes in n repeats of that event" → Binomial.
  • "Rare arrivals in a fixed window" → Poisson.
  • "Many tiny independent effects added up" → Gaussian.
  • "Wait until the next random arrival" → Exponential.
  • "Uncertainty about a probability itself" → Beta.

The whole skill is pattern-matching the story to the data. Once matched, you get probabilities, expectations, surprises (variance), and a training objective for free.

03.A Tiny Worked Example: Three Stories, Three Formulas

Story 1 — one click. Your banner gets clicked with probability p = 0.2.

Bernoulli: P(click) = 0.2, P(no click) = 0.8. Two numbers, done.

Story 2 — ten visitors. Show 10 independent visitors, each clicks with p = 0.2. How often do exactly 3 click?

Binomial: P(k) = C(10,k) · 0.2ᵏ · 0.8^(10−k).

  • P(0) = 0.8¹⁰ ≈ 0.107
  • P(3) = 120 · 0.008 · 0.134 ≈ 0.201
  • Expected clicks: np = 10 × 0.2 = 2. Most runs land near 2. ✓

Story 3 — tickets per hour. On average λ = 3 tickets arrive per hour, independently and rarely (relative to the sea of potential minutes).

Poisson: P(k) = e^(−3) · 3ᵏ / k!

  • P(0) = e⁻³ ≈ 0.05 — a dead hour is rare but happens.
  • P(3) = e⁻³·27/6 ≈ 0.22 — the most likely count.
  • Fingerprint: mean = variance = 3.

Same "count of 3s" question, different stories, different formulas. That's why you match the story first and the formula second.

04.The Discrete Workhorses (Meet the Animals)

  • Bernoulli(p): one trial, success probability p. Values {0, 1}. The native distribution of a binary classifier's output; its negative log-likelihood is binary cross-entropy.
  • Binomial(n, p): number of successes in n independent Bernoulli trials; mean np, variance np(1−p). Models click counts out of n impressions.
  • Categorical(π): one draw from k classes with probability vector π — the softmax layer's home turf, and the distribution behind a language model's next token.
  • Multinomial(n, π): counts across k categories in n draws (word counts in a document; the naive Bayes and topic-model building block).
  • Poisson(λ): counts of rare events arriving independently in a fixed window — support tickets per hour, defects per wafer. Its defining trait: mean = variance = λ. Under-dispersion or over-dispersion in real count data signals you need a quasi-Poisson or negative binomial instead.
  • Geometric(p): trials until the first success; the discrete analogue of the Exponential.

05.The Continuous Zoo

  • Uniform(a, b): constant density on an interval; every sub-interval of equal length is equally likely. Used for random initialization and Monte-Carlo sampling.
  • Normal(μ, σ²): the bell curve; mean μ, variance σ²; everything standardized by z = (x − μ)/σ against the standard normal N(0, 1). Sixty-eight–ninety-five–99.7 within 1, 2, 3 standard deviations.
  • Exponential(λ): waiting time until the next event of a Poisson process; memoryless — P(X > s + t | X > s) = P(X > t) (having waited s changes nothing). Models session churn and is the maximum-entropy distribution on positive reals with fixed mean.
  • Beta(a, b): a distribution over probabilities in [0, 1]; with Bernoulli likelihood it is the conjugate prior (Topic 18), shaping everything from A/B-test posteriors to Thompson sampling.
  • Laplace(μ, b): like a Gaussian with heavier tails and a sharp cusp at the peak. Gaussian noise assumption gives L2 loss; Laplace noise assumption gives L1 loss — which is why robust regressors and sparse priors use it.
  • Log-Normal(μ, σ²): a positive quantity whose log is Gaussian — salaries, file sizes, response times. Multiplicative noise produces it naturally.
  • Gamma(a, rate λ): sums of exponentials; positive, skewed; priors on precisions and inter-arrival totals.
python— Same mean, different tails: Gaussian vs Laplace survival beyond 3 units
import numpy as np
from math import erf

x = np.arange(0, 7)                          # thresholds 0..6
gauss_tail = 0.5 * (1 - erf(x / np.sqrt(2)))   # N(0,1) survival function
laplace_tail = 0.5 * np.exp(-x)                # Laplace(0, 1) survival

for xi, g, l in zip(x, gauss_tail, laplace_tail):
    print(f"P(X > {xi:.0f})  Gaussian={g:.2e}   Laplace={l:.2e}")
# At x = 6 the Laplace tail is orders of magnitude heavier: outliers are not rare here.

06.Visual Intuition: Bell, Cusp, and the Long Tail

Sketch the shapes and the habitat of each animal becomes obvious:

code
  Gaussian (bell, smooth tails)     Laplace (cusp, fat tails)
     ╭─╮                              ╲
    ╱   ╲        round peak            ╲╎  sharp point
   ╱     ╲       −−> averages, noise    ╱│╲  −−> occasional
  ╱───────╲    (errors, heights)      ╱ │ ╲    big outliers
            ╲___ decay fast                   ╲___ decay slowly

  Exponential (waiting times)       Poisson (counts, λ=3)
  │╲                                 ▊
  │ ╲  starts high, only falls        ▊  ▊
  │  ╲───  = "nothing happened yet"   ▊▊ ▊ ▊
  └───────► x                     0 1 2 3 4 5 → k

Two pairs to burn in:

  • Exponential ↔ Poisson are the same process wearing two hats: Poisson counts how many arrive in a window; Exponential times how long until the next one.
  • Gaussian ↔ Laplace differ only in the tails — but that difference decides whether your loss is MSE or MAE (Section 8).

07.The Analogy: You Are the Zookeeper

Carry this through: you're a zookeeper, and every distribution is an animal with a habitat it refuses to leave.

  • You don't ask "which animal do I like?" You ask "what's the habitat?" — is the data discrete or continuous? Positive-only? Counts or durations? Bounded probabilities or unbounded noise?
  • Drop a penguin (Gaussian) into the desert (heavy-tailed salaries) and it dies: it will predict negative salaries and call 6-sigma events impossible.
  • Put the right animal in the right habitat and it feeds you for free: mean, variance, tail probabilities, and a training loss all fall out of the species name.

The decision tree in the diagram above is literally the zookeeper's intake form: habitat questions in, species out.

And when no animal fits, you don't force it — you get a bigger cage (negative binomial for over-dispersed counts, mixtures, non-parametric models). Recognizing "none of the above" is a skill too.

08.Why the Gaussian Dominates: the Central Limit Theorem

One animal keeps winning: the bell curve. Here's why.

Add many independent, finite-variance random variables together and, after standardizing, the sum approaches a Normal distribution — regardless of the pieces' original shapes. This Central Limit Theorem explains why measurement noise, averaged errors, and aggregate demand look Gaussian, and why the Normal is the default "unknown noise" assumption.

Two supporting properties make the Gaussian the ML favorite:

  1. Maximum entropy: among all continuous distributions with a fixed mean and variance, the Gaussian has the highest entropy — the most honest "least information" assumption. Choosing it is admitting ignorance in the least biased way possible.
  2. Analytical tractability: closed-form marginals, conditionals, and linear transformations; products and convolutions stay Gaussian. The algebra never fights you.

Caveat: the CLT needs finite variance and independence-ish behavior. Financial returns, network latencies, city sizes, and word frequencies are heavy-tailed; Gaussian assumptions there understate catastrophic outliers by orders of magnitude.

09.In Practice: Choosing an Output Distribution = Choosing a Loss

Here is the payoff that turns zoo trivia into ML engineering.

A generative view of modeling: pick the distribution you believe generates y from x, then train by maximum likelihood (Topic 19 — choose the parameters that make your observed data most probable). The negative log-likelihood (NLL) loss of each familiar family reduces to a loss you already know:

  • Bernoulli head → binary cross-entropy with sigmoid.
  • Categorical head → categorical cross-entropy with softmax.
  • Gaussian head with constant variance → mean squared error.
  • Gaussian head with learned variance → a heteroscedastic uncertainty model (predict μ and σ per input).
  • Laplace head → mean absolute error (robust to outliers).
  • Poisson head → Poisson NLL for count data (traffic forecasting, insurance claims).
  • Negative-log-likelihood weighting is exactly how multi-task learning balances heads with different units.

So "which loss function?" is really "what story do you believe about your data?" — the zookeeper's intake form again.

All of the above families share one structure: they are members of the exponential family, whose log-density splits into a natural parameter times a sufficient statistic plus log-normalizer. This unification (generalized linear models, natural gradients, conjugate priors) is one of the deepest regularities in statistical ML.

Production Implementation in Big Tech
Ride-hailing and delivery platforms• ETA and demand modeling

Trip durations are modeled as right-skewed positive variables (log-normal or gamma), demand per zone per interval as Poisson/negative-binomial counts, and no-show probabilities as Bernoulli heads. Quantile predictions come from the fitted distributions inverses, letting dispatch systems reason in P50/P90 terms instead of raw means.

Staff+ Engineering Takeaways

  • Match the distribution to the data-generating story — the habitat: Bernoulli/Binomial/Poisson/Categorical for counts and classes; Normal/Exponential/Beta/Laplace/Log-Normal for continuous quantities.
  • Poisson has mean = variance; real count data that over-disperses needs negative binomial or quasi-Poisson.
  • The Central Limit Theorem and maximum-entropy optimality explain why Gaussian noise is the default assumption — and heavy-tailed data shows where that default breaks.
  • Laplace noise assumptions produce L1 loss, Gaussian produce L2: the assumed distribution determines the loss function.
  • Classification heads (Bernoulli, Categorical) turn cross-entropy into maximum-likelihood training; all these families are exponential-family members.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

You model "support tickets arriving per hour," assuming arrivals are independent and rare. Which distribution fits, and what signature property does it have?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?