# Implementing Diffusion Models (DDPM) from Raw Math: A Step-by-Step Coding Guide

> Learn to implement diffusion models (DDPM) from raw math. Code a forward noising process, train a neural network, and reverse corruption to synthesize data. Perfect for AI engineering.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-25

---

**Implementing diffusion models (DDPM) from raw math involves coding a forward Gaussian noising process with a closed-form transition, training a neural network to predict the injected noise using a simple MSE loss, and iteratively reversing the corruption to synthesize data from pure Gaussian noise.**

The repository `rohitg00/ai-engineering-from-scratch` provides a comprehensive educational resource for building **Denoising Diffusion Probabilistic Models (DDPM)** from first principles. This guide examines the mathematical derivations and reference implementation found in the diffusion lesson, translating theoretical concepts into executable Python code.

## The Forward Noising Process

The forward process defines a fixed **Markov chain** \(q(x_t|x_{t-1})\) that gradually corrupts data by adding Gaussian noise across \(T\) timesteps. Because each transition is Gaussian, the entire chain admits a **closed-form expression** allowing direct sampling of \(x_t\) from \(x_0\):

\[
q(x_t \mid x_0)=\mathcal{N}\bigl(\sqrt{\bar\alpha_t}\,x_0,\;(1-\bar\alpha_t)\mathbf I\bigr)
\]

Here, \(\bar\alpha_t=\prod_{s=1}^{t}(1-\beta_s)\) represents the cumulative product of noise variances, and \(\beta_t\) is the per-step variance schedule. According to the lesson documentation in [`phases/08-generative-ai/06-diffusion-ddpm-from-scratch/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/06-diffusion-ddpm-from-scratch/docs/en.md) ([lines 20-24](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/06-diffusion-ddpm-from-scratch/docs/en.md#L20-L24)), you typically implement \(\beta_t\) as either a **linear schedule** or a **cosine schedule**, with the latter often yielding superior sample quality. The reference code builds this schedule using a simple list comprehension as shown in the repository ([lines 66-73](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/06-diffusion-ddpm-from-scratch/docs/en.md#L66-L73)).

## The Reverse Denoising Process

To generate data, you learn the reverse chain \(p_\theta(x_{t-1}\mid x_t)\) parameterized by a neural network \(\varepsilon_\theta(x_t,t)\) that **predicts the noise** added at timestep \(t\). Given a noisy sample \(x_t\), the reverse step computes:

\[
x_{t-1}= \frac{1}{\sqrt{\alpha_t}}\Bigl(x_t-\frac{\beta_t}{\sqrt{1-\bar\alpha_t}}\;\varepsilon_\theta(x_t,t)\Bigr)+\sigma_t\,z
\]

where \(\sigma_t\) equals \(\sqrt{\beta_t}\) (or a learned variance), and \(z\sim\mathcal N(0,\mathbf I)\) injects stochasticity. This formula appears in the lesson markdown ([lines 30-33](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/06-diffusion-ddpm-from-scratch/docs/en.md#L30-L33)).

## Training Objective: Simple MSE Loss

Training reduces to a **single-line MSE loss** that regresses the true noise \(\varepsilon\) against the network's prediction. Unlike GANs, DDPMs require no adversarial training or complex variational bounds:

\[
\mathcal L_{\text{simple}} = \mathbb{E}_{x_0,t,\varepsilon}\bigl\|\varepsilon - \varepsilon_\theta\bigl(\sqrt{\bar\alpha_t}\,x_0+\sqrt{1-\bar\alpha_t}\,\varepsilon,\,t\bigr)\bigr\|^2
\]

This **epsilon prediction** objective appears in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) ([lines 38-40](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/06-diffusion-ddpm-from-scratch/docs/en.md#L38-L40)), providing a stable training signal that minimizes the variational lower bound.

## Sampling from the Learned Distribution

During inference, you initialize \(x_T\) from pure Gaussian noise \(\mathcal N(0,\mathbf I)\) and **iterate the reverse step** from \(t=T\) down to \(1\). The final output \(x_0\) represents a sample from the learned data distribution. The lesson documentation describes this sampling loop in ([lines 44-45](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/06-diffusion-ddpm-from-scratch/docs/en.md#L44-L45)).

## Complete 1-D Implementation

The file [`phases/08-generative-ai/06-diffusion-ddpm-from-scratch/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/06-diffusion-ddpm-from-scratch/code/main.py) contains a minimal yet complete implementation. Below are the essential components distilled from the source:

```python
import math, random

# 1. Build the noise schedule (linear default)

T = 1000
betas = [1e-4 + (0.02 - 1e-4) * t / (T - 1) for t in range(T)]
alphas = [1 - b for b in betas]
alpha_bars = []
cum = 1.0
for a in alphas:
    cum *= a
    alpha_bars.append(cum)

# 2. Forward sampling (closed-form)

def forward_sample(x0, t, rng):
    a_bar = alpha_bars[t]
    eps = rng.gauss(0, 1)               # ε ~ N(0, I)

    x_t = math.sqrt(a_bar) * x0 + math.sqrt(1 - a_bar) * eps
    return x_t, eps

# 3. Training step (ε-prediction)

def train_step(x0, model, rng):
    t = rng.randrange(T)
    x_t, eps = forward_sample(x0, t, rng)
    eps_hat = model_forward(model, x_t, t)   # neural network forward pass

    loss = (eps - eps_hat) ** 2
    # ... backpropagation omitted ...

    return loss

# 4. Reverse sampling (generation)

def sample(model, rng):
    x = rng.gauss(0, 1)               # start from pure noise x_T

    for t in range(T - 1, -1, -1):
        eps_hat = model_forward(model, x, t)
        beta_t = 1 - alphas[t]
        # Reparameterized reverse step

        x = (x - beta_t / math.sqrt(1 - alpha_bars[t]) * eps_hat) / math.sqrt(alphas[t])
        if t > 0:                     # add noise except at final step

            x += math.sqrt(beta_t) * rng.gauss(0, 1)
    return x

```

To run the demonstration, execute the script from the repository root:

```bash
python3 phases/08-generative-ai/06-diffusion-ddpm-from-scratch/code/main.py

```

The script trains a tiny MLP on a 1-D bimodal Gaussian mixture and outputs a sampled value after 200 epochs, validating that your **implementation of the diffusion process** correctly learns the data distribution.

## Practical Tips and Pitfalls

When **implementing diffusion models from scratch**, consider these implementation details from the lesson's "Pitfalls" section:

- **Noise schedule selection** significantly impacts quality. While the code uses a linear schedule by default, cosine schedules often improve FID scores and training stability.
- **Timestep embedding** requires careful encoding. The reference uses sinusoidal embeddings for the toy 1-D model; production systems typically employ **FiLM conditioning** or transformer-based time embeddings.
- **Prediction targets** vary by implementation. While this guide covers \(\varepsilon\)-prediction (noise), modern pipelines often predict **velocity** \(v\) or the direct denoised \(x_0\) for improved numerical stability.
- **Sampling speed** remains a bottleneck. The naive 1000-step loop shown above can be accelerated using **DDIM**, **DPM-Solver**, or distillation techniques that reduce steps to 50 or fewer.

These considerations appear in the lesson documentation ([lines 22-32](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/06-diffusion-ddpm-from-scratch/docs/en.md#L22-L32)).

## Summary

- **DDPMs** work by learning to reverse a fixed forward corruption process that adds Gaussian noise to data.
- The **forward process** has a closed-form solution allowing direct sampling of any timestep \(x_t\) from clean data \(x_0\).
- Training requires only **MSE loss** between the true noise and the network's prediction, eliminating the need for adversarial training.
- **Reverse sampling** starts from pure noise and iteratively applies the learned denoising step for \(T\) iterations.
- The reference implementation in [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py) demonstrates these concepts on a 1-D dataset, providing a foundation for scaling to high-dimensional image generation.

## Frequently Asked Questions

### What is the closed-form forward process in DDPM?

The closed-form allows you to sample \(x_t\) directly from \(x_0\) using \(q(x_t \mid x_0)=\mathcal{N}(\sqrt{\bar\alpha_t}\,x_0,\;(1-\bar\alpha_t)\mathbf I)\), where \(\bar\alpha_t\) is the cumulative product of \((1-\beta_s)\). This eliminates the need to iteratively apply noise for \(t\) steps during training, enabling efficient random timestep sampling.

### Why does DDPM use epsilon prediction instead of predicting the denoised image directly?

Predicting the noise \(\varepsilon\) provides a **stable training target** with consistent scale across timesteps. Directly predicting \(x_0\) would require the network to handle vastly different scales as \(t\) varies from clean data to pure noise. The epsilon parameterization simplifies the optimization objective and aligns with the variational lower bound derivation.

### How do you choose between linear and cosine noise schedules?

**Linear schedules** (beta linearly interpolated from \(10^{-4}\) to \(0.02\)) work adequately for simple datasets but often place too much noise weight in early timesteps. **Cosine schedules** distribute noise more evenly, typically yielding better FID scores on complex distributions like natural images. The choice depends on your data complexity and can be experimented with in the schedule construction code.

### Can this 1-D implementation scale to images?

While the mathematical framework remains identical, scaling to images requires architectural changes: replacing the MLP with a **U-Net** containing attention blocks, using 2-D convolutions instead of scalar operations, and implementing **classifier-free guidance** for conditional generation. The core sampling loop and loss function from [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py) remain valid, but the `model_forward` function must handle spatial structure.