DDPM: Denoising Diffusion Probabilistic Models
DDPM (2020) generates images by learning to reverse a simple process: gradually bury a real photo in Gaussian noise until it is pure static, then train one small skill — "given this noisy image, tell me what noise is in it" — and undo the damage one tiny step at a time. The whole method reduces to a noise-prediction MSE, and it launched the diffusion revolution behind Stable Diffusion and Sora.
The DDPM Forward and Reverse Markov Chains
Forward: a fixed chain q gradually adds Gaussian noise until data becomes pure noise. Reverse: a learned chain p_theta removes noise one step at a time, each transition modeled as a Gaussian whose mean is predicted by the denoising network.
01.The Problem: How Do You Create an Image From Nothing?
Imagine asking a computer to draw "a red fox in the snow."
The computer starts with nothing. No template. No stencil. Just random numbers. And it has to turn those numbers into millions of pixels that look like a fox.
One older approach was the GAN (topic 184): two networks play a fake-vs-real game against each other. It works, but the game is unstable — if one player gets strong, training collapses.
Then, in 2020, a different idea appeared that threw the game away. In plain words:
Destroying an image is easy. Just sprinkle noise on it. So learn to play the destruction backward.
That is the whole seed of DDPM — Denoising Diffusion Probabilistic Models (Ho, Jain, Abbeel, 2020), the paper that started the diffusion revolution.
Carry one picture through this whole topic: a scoop of ice cream on a hot day.
- Melting is easy to describe and impossible to stop. Warmth + time = soup.
- Playing the melting backward — soup reassembling into a perfect scoop — looks like magic.
- But if you know the melting rule exactly, then "un-melting" is just one question:
Given this slightly-melted state, what did it look like one degree ago?
DDPM answers that question over and over, about a thousand times, until static becomes a fox.
02.The Idea in Plain Words: Break It for Free, Learn the Repair
DDPM splits the job into two chains of steps.
Chain 1 — the forward process (the melting). Take a real image x_0 and add a tiny bit of Gaussian noise. Then a bit more. Then a bit more, for T steps (the paper uses T = 1000), until the image is indistinguishable from TV static. This chain is fixed: no learning, no choices, just a published recipe of how much noise to add at each step.
Chain 2 — the reverse process (the un-melting). Start from pure static and remove noise one small step at a time, until an image emerges. This chain is the thing we train.
Both chains are Markov chains. In one plain sentence: a Markov chain is a sequence of steps where the next state depends only on the current one — like dominoes falling; each domino only cares about its neighbor, not the whole row.
So the strategy becomes:
- Forward:
x_0 → x_1 → x_2 → … → x_T, each arrow adds noise. Destroying is free, so we hard-code it. - Reverse:
x_T → x_(T-1) → … → x_0, each arrow removes noise. Repairing is hard, so we learn it.
Why bother with such a roundabout plan? Because the hard task ("draw a fox from nothing") becomes a easy, well-defined task:
Given a noisy image, say what noise is inside it.
Spotting noise in a picture is a simple regression problem — exactly the kind neural networks are good at. The grand trick of the paper is that this humble skill, repeated a thousand times, is image generation.
03.The Forward (Noising) Chain: A Recipe, Not a Model
DDPM defines a fixed, non-learned Markov chain q over T steps that corrupts x_0 toward noise. Each transition adds a little Gaussian noise:
q(x_t | x_{t-1}) = N( x_t ; sqrt(1 - beta_t) * x_{t-1}, beta_t * I )
Read that slowly. All it says is: to get step t, take the image from step t-1 and
- shrink it slightly by
sqrt(1 - beta_t), and - add a pinch of fresh random noise with strength
beta_t.
(N(mean ; variance) is the normal/Gaussian distribution — the familiar bell curve of random jitter. I means each pixel gets its own independent jitter.)
The knob beta_t is called the noise schedule, and it grows over time: early steps barely touch the image, later steps bury it. The sqrt(1 - beta_t) shrinkage guarantees the signal is steadily destroyed while the total "volume" of the picture stays around 1, so x_T ends up looking exactly like standard normal noise.
Writing alpha_t = 1 - beta_t and alpha_bar_t = prod_{i<=t} alpha_i (the running product of all the shrink factors so far), a key Gaussian-composition result lets you skip straight to any timestep from the clean image:
codeq(x_t | x_0) = N( x_t ; sqrt(alpha_bar_t) * x_0, (1 - alpha_bar_t) * I ) x_t = sqrt(alpha_bar_t) * x_0 + sqrt(1 - alpha_bar_t) * epsilon, epsilon ~ N(0, I)
"Gaussian composition" just means: noise added twice equals noise added once, with combined strength — because the sum of two bell curves is another bell curve. So you do not have to melt the ice cream 400 times to reach step 400. One scoop of the right amount of noise gets you there.
That closed form is what makes training cheap: pick a random t, add the known amount of noise in one shot, and regress.
04.A Simple Worked Example: T = 3 Steps
Let tiny numbers make the formulas feel real. Suppose only 3 steps, with schedule:
beta_1 = 0.1, beta_2 = 0.2, beta_3 = 0.3
Then alpha_t = 1 - beta_t gives 0.9, 0.8, 0.7, and the running products are:
codealpha_bar_1 = 0.9 alpha_bar_2 = 0.9 × 0.8 = 0.72 alpha_bar_3 = 0.72 × 0.7 = 0.504
Now stand at t = 2 and use the skip-ahead formula:
codex_2 = sqrt(0.72) * x_0 + sqrt(0.28) * epsilon ≈ 0.85 * x_0 + 0.53 * epsilon
Translation: at step 2 the picture is about 85% original signal, 53% fresh noise. The fox is still clearly a fox, but with visible grain — a slightly melted scoop.
At t = 3: x_3 ≈ 0.71 * x_0 + 0.70 * epsilon — half photo, half static.
Keep extending the chain and alpha_bar_T shrinks toward zero. At alpha_bar ≈ 0:
x_T ≈ 0 * x_0 + 1 * epsilon = pure noise
Every pixel is now standard normal random jitter. The ice cream is fully soup.
05.The Reverse Chain: Learning to Undo One Step
Here is the part we actually train.
In the ice-cream video, every tiny forward step is melt a little. The reverse step is un-melt a little. So we want a model for:
p_theta(x_{t-1} | x_t)— "given this mostly-melted state, what was it one step ago?"
DDPM makes the same Gaussian guess the forward chain uses: each reverse transition is also Gaussian,
p_theta(x_{t-1} | x_t) = N( mu_theta(x_t, t), sigma_t^2 * I )
A Gaussian needs two numbers: a center (mu, where to land) and a spread (sigma^2, how jittery). But there is a catch:
- The true spread of the reverse step (the posterior variance) is intractable — in plain words, computing the exact right amount of jitter would require summing over every image that ever existed. Impossible.
- Ho et al. did something brazen: they fixed
sigma_t^2to a constant (beta_t, or the tighter boundbeta_tilde_t) and refused to learn it at all. The network only has to predict the meanmu_theta— the direction, not the wiggle room.
It worked. And the mean has a beautiful reparameterization. Since x_t = sqrt(alpha_bar_t) x_0 + sqrt(1-alpha_bar_t) epsilon, knowing the noise epsilon is equivalent to knowing the clean image — so instead of predicting mu directly, predict the noise that got in there, and algebra hands you the mean.
Sanity-check it with the t = 2 numbers from Section 4. If x_2 ≈ 0.85 x_0 + 0.53 ε and the network guesses ε, then:
x_0 estimate ≈ (x_2 - 0.53 * eps_hat) / 0.85
Scrape off the estimated noise, stretch back the shrink — instant cleaner image estimate. Then add the small fixed jitter and step to x_1. Repeat.
The training loss is a reweighted variational lower bound (ELBO — a proxy for "how likely is this image under my model", used because the true likelihood is the same intractable sum; topic from the VAE family). Empirically, the weights only hurt: the shockingly simple objective works best.
L_simple = E_{t, x_0, epsilon} [ || epsilon - epsilon_theta( sqrt(alpha_bar_t)x_0 + sqrt(1-alpha_bar_t)epsilon, t ) ||^2 ]
Read it out loud: pick a real image, pick a random step, bury it in known noise, and train the network to name the noise. Plain squared error. No adversary, no game, no instability — just MSE.
So the model epsilon_theta is literally a noise predictor: given x_t and t, guess the epsilon that created it. Subtracting that estimate recovers a better x_{t-1}.
def p_sample(model, x_t, t):
eps = model(x_t, t) # predicted noise
a, abar = alpha[t], alpha_bar[t]
beta_t = 1 - a
# recover x_{t-1} mean and add posterior noise
mean = (1 / sqrt(a)) * (x_t - beta_t / sqrt(1 - abar) * eps)
if t > 0:
z = torch.randn_like(x_t)
return mean + sqrt(beta_t) * z
return mean06.Visual Intuition: The Fox Dissolving Into Static
Picture the forward chain as a film of a photo dissolving:
codet=0 t=250 t=500 t=750 t=1000 🦊 🦊≈ 🌫️🦊? 🌫️≈ ▒▒▒▒▒ sharp light grain ghost of barely pure photo a fox anything static forward →→→→→→→→→ add noise (fixed recipe, the melting) reverse ←←←←←←←←← remove noise (learned, the un-melting)
And here is the loop one training run actually performs:
codereal image x_0 ──► add ε (one shot, any t) ──► noisy x_t │ network sees (x_t, t) ▼ must name ε ◄── ε_theta(x_t, t) │ MSE loss pulls ε_hat → ε
At sampling time the same pieces run in the other direction:
codez_T (pure noise) │ predict ε_hat → subtract → x_(T-1) ▼ x_999 → x_998 → … → x_1 → x_0 (a fox) ▲ one tiny un-melt per network call, ~1000 calls
Notice what the network never sees: a whole image appearing from nothing. Each call is a humble question — "how much grain of this type is in this picture?" — and the miracle is that a thousand humble answers stack into a photograph.
07.What the Denoiser Actually Is, and the 2021 Upgrades
The denoiser epsilon_theta in the original paper is a U-Net — a convolutional network that downsamples then upsamples, keeping skip connections — with self-attention added at the 16x16 and 8x8 feature maps, GroupNorm instead of batch norm, and sinusoidal timestep embeddings injected through residual blocks (the network is told which step of the melt it is looking at, because the right amount of noise to scrape off depends on t).
The original noise schedule was linear: beta_1 = 1e-4 rising to beta_T = 0.02. Good, but not great near t = 0, where the signal-to-noise ratio falls off a cliff and the hardest denoising (barely any noise!) happens. Nichol & Dhariwal's Improved DDPM (2021) fixed that and more:
- Cosine schedule: let
alpha_bar_tfollow a cosine curve so the signal-to-noise ratio changes smoothly neart = 0, reducing sample-quality loss at low noise and improving likelihood. - Learned reverse variance: instead of freezing
sigma_t^2between its two fixed bounds, interpolate the bounds and let the network predict the blend. - Importance sampling of
t: low-tterms dominate the ELBO, so sample timesteps non-uniformly during training and re-weight to stay unbiased — spend compute where the loss lives. - Discretized log-likelihood for pixels: pixels are quantized integers, so model the data likelihood as a bin around each pixel value rather than a continuous Gaussian.
Log-likelihood itself is estimated with an importance-weighted bound, because the exact marginal q(x_0) is a high-dimensional integral you cannot compute directly.
These changes were the difference between "diffusion works" and "diffusion beats everything on FID / likelihood."
08.Why AI Cares: The DDPM Legacy in 2024-2026
DDPM is now a foundation, not the shipping frontier. Product models (Stable Diffusion, Imagen, Sora-class systems) moved diffusion into a VAE latent (the "compress first" idea — topic 189) and switched to deterministic/ODE samplers like DDIM, v-prediction or rectified-flow parameterizations, and transformer backbones. Yet the DDPM DNA survives everywhere:
- the
epsilon_thetanoise-predictor network, - the
alpha_bar_tclosed-form one-shot noising, - cosine schedules,
- timestep conditioning (the network always knows how melted it is looking),
are all in every modern scheduler's source code.
In research, the "improved DDPM" ideas — learned variance, non-uniform timestep sampling, better bounds — directly fed the score-based SDE unification (viewing diffusion as following a field of "which way is noisier" arrows) and the design of distillation objectives (consistency models) that later collapsed sampling to one step.
And back to the hiker's cousin, the restorer's rule of thumb: everything added since 2020 — samplers, schedulers, guidance, distillation — is only about how big a scrape to make per step, and whether you can scrape twice as much and skip frames. The un-melting idea itself never changed. Understanding DDPM cold is the fastest way to read any 2024-2026 diffusion paper.
Architectural Trade-offs & Production Realities
Architectural Advantages
- A single, stable MSE objective with a clear probabilistic derivation - easy to train, no adversarial game.
- Strong mode coverage and high-quality samples; became the reference formulation for the whole field.
- Cheap training via the closed-form q(x_t|x_0) jump - no iterative noising per sample.
Trade-offs & Constraints
- Ancestral sampling needs ~1000 sequential network calls (slow inference).
- Pixel-space DDPM at high resolution is memory- and compute-hungry - hence latent diffusion.
- Fixed-variance, noise-prediction form is suboptimal; v-prediction, learned variance, and flow-matching improved on it.
The paper demonstrated unconditional image synthesis (CIFAR-10, LSUN bedrooms at 256px) and class-conditional generation with an attention U-Net. Its forward/reverse chains and DDIM sampler ship today as the DDPMScheduler and cosine-schedule helpers inside the diffusers library that powers Stable Diffusion apps.
Staff+ Engineering Takeaways
- DDPM defines a fixed forward Gaussian Markov chain (the melting) and a learned Gaussian reverse chain (the un-melting) that inverts it.
- The jump q(x_t|x_0) = N(sqrt(abar_t)x_0, (1-abar_t)I) lets you noise to any timestep in one operation, because two Gaussians compose into one.
- Training reduces to the simplified loss ||eps - eps_theta(x_t, t)||^2: the network just predicts the added noise, and subtracting it steps toward a cleaner sample.
- Fixing the reverse variance to a constant (beta_t or beta_tilde_t) sidesteps the intractable posterior - the mean, carrying the signal, is all the network must learn.
- Improved DDPM added a cosine schedule, learned variance, and importance sampling of timesteps - each a real quality win.
- DDPM is the theoretical bedrock; modern systems port its epsilon-parameterization into latent space with ODE samplers and transformers.
Topic Knowledge Check
Exercise 1 of 4 • Test your architectural comprehension.
Why can DDPM training add noise to any timestep t in a single operation, instead of stepping through every earlier timestep?
How clear and actionable was this distributed systems breakdown?