Random Variables
A random variable is not a variable — it is a function that converts messy outcomes into numbers so math can do probability with them. Here are the three descriptors that capture its behavior — the PMF (discrete), the PDF (continuous), and the CDF (both) — and why the datasets, labels and batches in ML are random variables.
From Experiment to Distribution
A random variable turns messy outcomes into numbers; a PMF or PDF assigns probabilities to those numbers, and the CDF works for every kind of random variable.
01.The Problem: You Can't Do Arithmetic with "Heads"
The world is uncertain.
Dice roll. Users click. Latency spikes. Emails land.
You want to calculate with all that uncertainty: averages, risks, predictions.
But there's a silly obstacle:
You can't add "heads" + "tails". You can't average "spam, not spam, spam".
Math doesn't speak outcomes. Math speaks numbers.
So the fix is almost embarrassingly simple:
First turn every outcome into a number.
- Coin flips → "number of heads."
- An email → "0 if ham, 1 if spam."
- A user visit → "number of clicks in one minute."
The turning-into-numbers rule has a name: a random variable.
The name is famous for misleading people. It is not a variable in the programming sense — nothing "varies randomly" inside it. Keep reading.
02.The Idea in Plain Words: A Function, Not a Variable
A random variable is a function that maps each outcome of a random experiment to a real number.
Let's define the pieces.
Start with a random experiment — rolling dice, sampling a user from traffic, drawing a batch of images.
Its set of all possible results is the sample space Ω (the Greek letter "omega" — just the complete list of what could happen).
A random variable X is a deterministic function X: Ω → R that assigns a real number to each outcome.
That's the whole definition. Three things to notice:
- The randomness lives in the experiment — in which outcome actually happens.
- X itself is a boring, fixed measurement rule. Same outcome in → same number out, every time.
- Because X is just a function, probability can be studied with the same mathematical machinery as functions and spaces. That's the entire reason this definition exists.
Convention: random variables get capital letters (X, Y); the values they take get lowercase. So P(X = 2) reads "the probability that the measurement equals 2" — a statement about the experiment, not about the function.
03.A Tiny Worked Example: Three Coin Flips
Flip 3 coins. An outcome is the whole ordered sequence — there are exactly 8 of them:
HHH, HHT, HTH, THH, HTT, THT, TTH, TTT
That list is Ω.
Now define the random variable: X = number of heads.
X is just a translation table:
HHH→ 3HHT,HTH,THH→ 2HTT,THT,TTH→ 1TTT→ 0
Since the 8 outcomes are equally likely (each 1/8), counting gives:
| value k | outcomes | P(X = k) |
|---|---|---|
| 0 | 1 | 1/8 |
| 1 | 3 | 3/8 |
| 2 | 3 | 3/8 |
| 3 | 1 | 1/8 |
Four probabilities, and they sum to 1/8 + 3/8 + 3/8 + 1/8 = 1. ✓
That little table has a name — the probability mass function (PMF) — and it's the next section.
One thing to feel before moving on: the coins never contained the number 2. "2 heads" is what your scorekeeping rule X calls the outcome HHT. The measurement is in the notebook, not in the coins.
04.Two Worlds: PMFs, PDFs, and the Trusty CDF
Random variables come in two flavors, and the split causes more beginner confusion than everything else combined.
Discrete random variables take countably many values — number of clicks, word IDs, bit flips, heads in 3 flips. Their distribution is a probability mass function (PMF):
p(x) = P(X = x), with p(x) ≥ 0 and all the values summing to 1.
Each value gets a little "mass" of probability. Like the coin table above.
Continuous random variables take values in intervals — latency in ms, a pixel intensity, Gaussian noise. Here exact points are a problem:
What is P(Xexactly 0.5)?
Zero. A single point has no room for probability mass — think of dividing a finite amount of sand over infinitely many razor-thin slices. So instead of masses we use a probability density function (PDF):
- Probability of landing in an interval [a, b] = the area under f over that interval (the integral).
- The density itself can exceed 1 — only areas are probabilities, and the total area must be 1.
The bridge that works for both worlds is the cumulative distribution function:
F(x) = P(X ≤ x)
- Always non-decreasing, with
F(−∞) = 0andF(+∞) = 1. - For discrete X it's a staircase (one step per value).
- For smooth continuous X, the PDF is just its derivative:
f = F′. - Quantiles are inverse-CDF lookups — the median is
F⁻¹(0.5).
import numpy as np
rng = np.random.default_rng(7)
rolls = rng.integers(1, 7, size=(100_000, 2))
s = rolls.sum(axis=1) # sum of two dice, values 2..12
pmf = np.bincount(s, minlength=13)[2:] / s.size
cdf = np.cumsum(pmf)
print("P(S = 7) approx", pmf[5].round(3), "| exact", round(6/36, 3))
print("P(S <= 7) approx", cdf[5].round(3), "| exact", round(21/36, 3))05.Visual Intuition: Sand Piles vs Sand Beds
Picture the difference:
codeDISCRETE: PMF = separate sand piles CONTINUOUS: PDF = one sand bed ▊ ▊ ▊ ▊ ▊▊ ▊ ▊ ▊ area ▊▊▊ ▊ ▊ ▊ ▊ under the ▊▊▊▊▊ ──┐ ▊ ▊ ▊ ▊ ▊ curve = 1 ▊▊▊▊▊▊▊▊▊ ────┴─ a slice's [0][1][2][3][4] pile height _/_/_/_/_/_/_/ area = its values = its probability values probability
- Discrete: probability sits in lumps you can point at (P(X = 2) = the height of one pile).
- Continuous: probability is spread over a bed — a single razor-thin slice has no sand in it (hence P(X = exact point) = 0), and you ask for the sand between two marks (an interval's area).
And the CDF? It's a walking counter: walk left to right and keep a running total of probability collected.
codeF(x) 1 ┤ ┌──────── │ ┌────┘ ← flat between values: ½ ┤ ┌───── no new sand collected │ ───── (staircase = discrete) 0 └─┴─────┴─────┴─────┴────► x
For a smooth variable the staircase's steps are infinitely small, so the counter looks like a smooth ramp — and its slope at each point is exactly the PDF.
06.The Analogy: The Scorekeeper at a Noisy Tournament
Carry one analogy through the rest of probability: X is a scorekeeper at a chaotic tournament.
The games themselves are messy — players scramble, crowd noise, chaos.
The scorekeeper sits at the side with a clipboard and speaks only numbers.
- The experiment is the match being played (this is where the randomness is).
- The scorekeeper's rule (count goals, time the lap) is fixed and deterministic — that's the function
X: Ω → R. - The PMF is the final scoreboard: each possible score with how often it occurred.
- The CDF is the running tally: "how many matches ended with score ≤ x?"
- A continuous quantity — exact finish time — is a scorekeeper with infinite decimal precision: no two matches ever land on the exact same number, so you report "how many finished between 10.0 and 10.5 seconds" (an interval), never an exact point.
Whenever a definition feels abstract, ask: "what does the scorekeeper write down?"
Datasets, metrics, model outputs — behind every one of them stands a scorekeeper turning chaos into numbers.
07.Making Variables Talk: Joint, Marginal, and Conditional
Real models juggle many random variables at once — traffic AND weather AND price. The joint distribution p(X, Y) describes them together. From a joint, three moves unlock everything:
- Marginalize (sum or integrate out) unwanted variables:
p(X) = Σ_y p(X, y). Marginalizing is how you go from the joint "probability of rain AND traffic" back down to the single question "probability of traffic" — rain summed away, left out in the cold. - Condition:
p(X | Y) = p(X, Y) / p(Y). Conditioning = "I learned Y happened; now what should I believe about X?" This is the foundation of inference — Bayes' theorem (the next topic) is just two conditioning orders related by the marginals. - Independence means the joint factors:
p(X, Y) = p(X)·p(Y)— knowledge of one changes nothing about the other. Rarely exactly true in the real world, and enormously convenient when approximated.
And the master identity:
- The chain rule:
p(X₁, …, Xₙ) = Π p(Xᵢ | X₁, …, Xᵢ₋₁)— always holds, no assumptions. Any big joint probability splits into a product of conditionals.
That last line is not book trivia. Autoregressive language models (ChatGPT-style LLMs that generate text one token at a time) literally factor a sentence's distribution this way: each new token is one p(Xᵢ | past) term.
08.Why AI Cares: ML Is Built on Random Variables
Once you see inputs, labels, parameters, and noise as random variables (each one a scorekeeper rule), the whole ML pipeline becomes probability language:
- A dataset is a sample of m i.i.d. (independent, identically distributed) draws from the data distribution
p(x, y)— every training batch is one realization of a random variable, which is why two runs with different shuffles end up slightly different. - The training loss is an empirical estimate of an expected loss over the data distribution; optimization targets the expectation, not the sample (Topic 17 makes expectation precise).
- Model outputs are often predictive distributions
p(y | x; θ): classification is a Categorical random variable, regression a Gaussian (see Topic 16). - Dropout samples Bernoulli masks (random 0/1 switches over neurons); BatchNorm estimates means/variances of activation random variables; Bayesian networks and Markov decision processes are structured families of random variables with conditional-independence rules.
In practice: whenever you define a metric, split data, or sanity-check a model's outputs, you're deciding what the scorekeeper writes down — so decide it before training, and the probabilities will mean what you think they mean.
Every metric in an A/B test — click-through rate, revenue per session, latency p99 — is modeled as a random variable with an unknown distribution. The platform simulates or resamples (bootstrap) its distribution to compute confidence intervals and p-values, and stops the experiment early when the posterior of the treatment effect clears a decision threshold.
Staff+ Engineering Takeaways
- A random variable is a function from the sample space of a random experiment to real numbers — the randomness is in the experiment, not the variable.
- Discrete variables use a PMF (masses that sum to 1); continuous variables use a PDF (density whose areas are probabilities).
- The CDF F(x) = P(X ≤ x) is the universal description: non-decreasing from 0 to 1, with the PDF as its derivative when smooth.
- Joint distributions combine variables; marginalization sums them out; conditioning divides by the evidence; independence is a factorization.
- ML training is estimating expectations over random variables: batches, labels, dropout masks, and model outputs are all random variables.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Which statement best describes a random variable?
How clear and actionable was this distributed systems breakdown?