Concepts

Diffusion models

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

A diffusion model learns to reverse a process that gradually adds noise to data.

1 · What it is

A fixed Markov chain gradually adds Gaussian noise to data according to a variance schedule. Training learns transitions that reverse the diffusion process.

The reverse-time SDE transforms a known prior back into the data distribution by slowly removing noise. It depends on a time-dependent gradient field called the score. A numerical solver provides an approximate SDE trajectory.

Sampling time scales linearly with trajectory length in the DDIM analysis. DDIM reported samples produced 10 to 50 times faster than DDPM in wall-clock comparisons. One score-matching method uses a sequence of noise-perturbed distributions.

2 · Why it exists

A complex data distribution can be connected to a simple known distribution through many small transitions.

Known startThe forward process transforms data toward a known prior distribution by slowly injecting noise.
Learned returnA reverse-time process can transform that prior back toward the data distribution by removing noise.
Many stepsSampling time grows with the length of the sampling trajectory.
3 · How it works

Follow one example through training and one sample through generation.

Forward noise process and learned reverse process The top row shows a clean example becoming noisy through a fixed schedule and a model learning a reverse transition. The bottom row shows sampling from a Gaussian prior through repeated reverse transitions into a generated sample.
Training learns transitions that reverse diffusion. A reverse-time process transforms the prior distribution back into the data distribution.
  1. 1 · corruptA fixed Markov chain gradually adds Gaussian noise according to a variance schedule.
  2. 2 · learnModel transitions are trained to reverse the diffusion process.
  3. 3 · startSampling begins from the known prior distribution.
  4. 4 · solveNumerical solvers provide approximate trajectories from SDEs.

DDPM says the forward variances can be held constant as hyperparameters.

4 · Where it's used
WhoWhat they askWhat it works with
Image-generation team“Can a sample be generated from a noise prior?”Learned reverse transitions
Sampling researcher“Can fewer trajectory steps still produce useful samples?”DDIM sampling trajectories
Score-model researcher“Which vector field points toward more likely data?”Scores across multiple noise levels
5 · What it solves, and what it doesn't
solves
  • The original diffusion formulation slowly destroys structure through an iterative forward process.
  • Its learned reverse process restores structure in data.
  • Model transitions are learned to reverse a diffusion process.
  • Score-based SDEs express generation as a reverse-time stochastic differential equation.
  • DDIM reported samples produced 10 to 50 times faster than DDPM in wall-clock comparisons.
doesn't solve
  • A numerical solver provides an approximate trajectory rather than an exact continuous path.
  • The time required for a sample scales linearly with the trajectory length in the DDIM analysis.
  • One score-matching method uses a sequence of noise-perturbed distributions.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperDeep Unsupervised Learning using Nonequilibrium Thermodynamics, Sohl-Dickstein et al. · read 28 Sept 2026
  2. paperDenoising Diffusion Probabilistic Models, Ho, Jain and Abbeel · read 28 Sept 2026
  3. paperScore-Based Generative Modeling through Stochastic Differential Equations, Song et al. · read 28 Sept 2026
  4. paperDenoising Diffusion Implicit Models, Song, Meng and Ermon · read 28 Sept 2026
  5. paperGenerative Modeling by Estimating Gradients of the Data Distribution, Song and Ermon · read 28 Sept 2026