The Idea: Destroy, Then Reconstruct
A diffusion model learns to generate data by learning to reverse a destruction process. Gradually add noise until data becomes pure Gaussian noise, then train a neural network to undo each step of corruption.
The forward process is trivial (just add noise). The reverse process is learned. The insight is that learning to denoise is far easier than learning to generate from scratch.
1. The Forward Process: A Prescribed Destruction
Define a forward SDE that smoothly transforms data into noise:
The standard choice (Variance Preserving SDE):
where is a noise schedule. At , we have data. At , we have (approximately) .
The key property: the marginal at any time has a closed-form Gaussian expression:
where . This means we can sample directly from without simulating the SDE step by step—essential for efficient training.
2. The Score Function: Gradient of Log-Density
The score of a distribution is the gradient of its log-density:
Why this matters: the score tells you the direction of steepest increase in probability density at any point. It points from low-density regions toward high-density regions—from noise toward data.
Anderson’s Theorem (1982): the forward SDE has a time-reversed counterpart:
The reverse SDE has the same diffusion coefficient but a modified drift that involves the score . If we know the score at every noise level, we can exactly reverse the destruction process.
3. Denoising Score Matching: Tweedie’s Identity
We don’t know —that’s the whole problem. But we can learn the score via a beautiful connection to denoising.
Tweedie’s formula: For the Gaussian perturbation q(x_t | x_0) = \mathcal{N}(\sqrt{\bar{\alpha}_t} x_0, (1 - \bar{\alpha}_t} )I):
where is the noise that was added. So the score is proportional to the negative noise. Predicting the score is equivalent to predicting the noise—which is a standard denoising problem.
The training loss (denoising score matching):
Sample a data point , a noise level , add noise to get , and train the network to predict . The learned noise predictor gives the score via .
4. Why the Noise Schedule Matters
The schedule controls how quickly information is destroyed. It profoundly affects both training and generation:
Too fast (aggressive schedule): The transition from data to noise happens in a few steps. The intermediate distributions change rapidly, making the score function highly nonlinear—harder for the network to approximate.
Too slow (conservative schedule): Requires many reverse steps to generate, increasing inference cost. But the score varies smoothly, making it easier to learn.
The SNR perspective: At time , the signal-to-noise ratio is . The training loss at each is weighted by the SNR. High-SNR timesteps (early, low noise) focus on large-scale structure. Low-SNR timesteps (late, high noise) focus on fine details.
The noise schedule implicitly defines a curriculum: the model learns coarse structure first (from high-noise denoising) and fine details later (from low-noise denoising). This coarse-to-fine hierarchy is a key reason why diffusion models produce such coherent outputs.
5. Diffusion vs. Flow Matching: The Fundamental Difference
Both learn a time-dependent transformation from noise to data. The difference is in the path:
| Aspect | Diffusion | Flow Matching |
|---|---|---|
| Forward process | Stochastic (SDE) | Deterministic (ODE) |
| Paths | Curved, noisy | Straight lines |
| Training target | Noise (score) | Velocity |
| Inference steps | 50-1000 | 10-50 |
| Theoretical framework | Score matching | Optimal transport |
The diffusion framework is richer theoretically (connections to statistical physics, Langevin dynamics, thermodynamics). Flow matching is simpler practically (straight paths, fewer steps, no noise schedule).
Both converge to the same learned distribution. The paths between noise and data differ, but the endpoints are the same. Flow matching can be seen as the deterministic limit of diffusion (the “probability flow ODE”), with the added innovation of straight-line conditional paths.
6. The Score Is All You Need
The score function is the central object. From it, you can:
- Generate (reverse the SDE/ODE)
- Compute likelihoods (via the probability flow ODE and the instantaneous change-of-variables formula)
- Inpaint (condition on partial observations by modifying the score)
- Guide (classifier-free guidance adds a conditional score term)
This is why diffusion models are so flexible: the score is a local quantity (a gradient at a point), and local modifications to the score produce global changes in the generated distribution. Classifier-free guidance, for instance, simply interpolates between the conditional and unconditional score—a one-line modification that dramatically improves sample quality.
The Takeaway
Diffusion models convert the hard problem of generation into the easy problem of denoising, mediated by the score function. The forward process destroys structure; the learned reverse process reconstructs it. The noise schedule determines the curriculum. And the score—the gradient of log-density—is the thread that connects destruction to reconstruction, training to inference, and probability theory to practical generative modeling.