Difference Between DDPM and DDIM Sampling in Diffusion Models: Implementation Guide

DDPM sampling uses a fully stochastic reverse Markov chain with Gaussian noise injected at every timestep, while DDIM sampling provides a deterministic or partially stochastic implicit method that enables high-quality generation with as few as 10–50 steps.

The difference between DDPM and DDIM sampling in diffusion models fundamentally changes how latent variables traverse the reverse diffusion trajectory during inference. Both methods share the same training objective and forward diffusion process, but they diverge in their reverse step formulation, stochasticity control, and computational efficiency. In the labmlai/annotated_deep_learning_paper_implementations repository, these strategies are implemented as interchangeable samplers within the Stable Diffusion framework, allowing researchers to switch between stochastic fidelity and deterministic speed.

Core Architectural Differences

Understanding the distinction between these sampling methods requires examining how each computes the transition from a noisy latent x_t to a less noisy x_{t-1}.

Reverse Process Formulation

DDPM (Denoising Diffusion Probabilistic Models) implements the exact reverse Markov chain derived from the variational lower bound. In labml_nn/diffusion/stable_diffusion/sampler/ddpm.py, the DDPMSampler class computes the posterior p_θ(x_{t-1}|x_t) by sampling from a Gaussian distribution characterized by a learned mean μ_t and variance β̃_t. This approach explicitly models the probabilistic nature of the reverse diffusion.

DDIM (Denoising Diffusion Implicit Models) treats the reverse process as a non-Markovian implicit sampler. The DDIMSampler class in labml_nn/diffusion/stable_diffusion/sampler/ddim.py computes x_{τ_{i-1}} directly from x_{τ_i} using a deterministic update formula that predicts the clean image x̂₀ and adjusts the trajectory without relying on the learned posterior variance. This formulation eliminates the Markov chain assumption, allowing the sampler to jump between arbitrary timesteps in the diffusion schedule.

Stochasticity and the Eta Parameter

The DDPM sampler is inherently stochastic. During each call to p_sample, the implementation draws random noise using torch.randn(...) and scales it by exp(0.5 * log_var), ensuring every generation follows a unique random path through the latent space.

DDIM introduces controllable stochasticity through the ddim_eta parameter (η). When η = 0, the get_x_prev_and_pred_x0 method produces fully deterministic outputs by setting the noise scale σ_{τ_i} = 0. For values η > 0, DDIM adds scaled random noise σ_{τ_i} ε_{τ_i}, allowing interpolation between deterministic and stochastic sampling. This flexibility is implemented in lines 88–94 of ddim.py, where ddim_sigma is computed based on the eta value and the alpha schedule.

Implementation Deep Dive

Both samplers inherit from the DiffusionSampler base class defined in labml_nn/diffusion/stable_diffusion/sampler/__init__.py, which provides shared utilities such as get_eps for unconditional guidance and abstract sample method signatures.

DDPM Sampler Implementation

The DDPMSampler class pre-computes diffusion schedule parameters during initialization (lines 67–84 in ddpm.py):

  • sqrt_alpha_bar and sqrt_1m_alpha_bar for reconstructing x₀ from noise predictions
  • mean_x0_coef and mean_xt_coef for computing the deterministic component of the mean
  • log_var representing the logarithm of the reverse process variance

During the p_sample method, the sampler calculates the posterior mean using these coefficients, then adds stochastic noise scaled by the variance term. This ensures the reverse process maintains the Gaussian properties required by the variational bound.

DDIM Sampler Implementation

The DDIMSampler constructs a sub-schedule (self.time_steps) based on the desired number of inference steps, enabling uniform or quadratic discretization of the original training timestep range. Key pre-computed tensors include:

  • ddim_alpha and ddim_alpha_prev representing α values at selected timesteps
  • ddim_alpha_sqrt for directional weighting
  • ddim_sigma derived from η to control the noise injection scale

The get_x_prev_and_pred_x0 method (lines 31–48 and 84–106) implements the implicit update formula that reconstructs the predicted clean image and interpolates along the diffusion trajectory without sampling from the full posterior distribution.

Mathematical Foundations

The mathematical distinction between these methods determines their practical behavior in generation pipelines.

DDPM Reverse Markov Chain

DDPM sampling follows the equation:


x_{t-1} = μ_t(x_t) + √β̃_t · ε,  where ε ~ N(0, I)

Here, μ_t(x_t) represents the predicted mean, and β̃_t is the variance of the reverse process. The noise term ε ensures the sampler explores the full distribution of possible previous states, maintaining the probabilistic integrity of the diffusion model at the cost of requiring many sequential steps (typically 1000) for high-quality results.

DDIM Implicit Sampling Formula

DDIM computes the previous latent using:


x_{τ_{i-1}} = √α_{τ_{i-1}} · ((x_{τ_i} - √(1-α_{τ_i})·ε_θ(x_{τ_i})) / √α_{τ_i})
              + √(1-α_{τ_{i-1}}-σ_{τ_i}²) · ε_θ(x_{τ_i})
              + σ_{τ_i} · ε_{τ_i}

When the ddim_eta parameter is set to zero, σ_{τ_i} = 0, eliminating the final noise term and producing deterministic outputs. This formulation allows the sampler to skip intermediate timesteps (sub-sampling the schedule to 10–50 steps) while maintaining trajectory coherence because the path to the predicted clean image x̂₀ is computed directly rather than accumulated through Markov transitions.

Practical Performance Trade-offs

Selecting between DDPM and DDIM involves balancing generation quality, inference speed, and output diversity.

Sampling Speed and Step Reduction

DDPM requires the full training schedule (typically 1000 steps) to maintain distributional fidelity. Reducing steps significantly degrades output quality because the stochastic transitions accumulate error when spaced too far apart.

DDIM enables arbitrary sub-sampling of the original schedule. By setting n_steps=50 and ddim_eta=0.0, the sampler achieves comparable visual fidelity to 1000-step DDPM in a fraction of the inference time. The deterministic trajectory follows a more direct path through latent space, making each forward pass through the U-Net more information-dense.

Quality and Consistency

While DDPM provides the theoretical guarantee of sampling from the true data distribution given enough steps, the stochastic nature introduces variance between runs with the same seed and prompt. DDIM with η = 0 produces identical outputs for identical inputs, enabling reproducible generation and deterministic editing workflows. When η > 0, DDIM introduces controlled variation that can be tuned for creative diversity without requiring the full computational cost of DDPM.

Code Examples

The following examples demonstrate how to instantiate and use both samplers with a Stable Diffusion model from the repository.

Using DDPMSampler for Stochastic Generation

from labml_nn.diffusion.stable_diffusion.sampler.ddpm import DDPMSampler
from labml_nn.diffusion.stable_diffusion.latent_diffusion import LatentDiffusion

# Load pretrained model

model = LatentDiffusion.load('sd-v1-4')  # Placeholder loader

# Initialize DDPM sampler (stochastic, full schedule)

ddpm = DDPMSampler(model)

# Generate samples

samples = ddpm.sample(
    shape=[1, 4, 64, 64],
    cond=text_embedding,
    temperature=1.0,
    uncond_scale=7.5,  # Classifier-free guidance

)

# Decode to RGB

images = model.decode(samples)

Using DDIMSampler for Fast Deterministic Generation

from labml_nn.diffusion.stable_diffusion.sampler.ddim import DDIMSampler

# Initialize DDIM with 50 deterministic steps

ddim = DDIMSampler(model, n_steps=50, ddim_eta=0.0)

# Generate samples with optional stochasticity (eta > 0)

samples = ddim.sample(
    shape=[1, 4, 64, 64],
    cond=text_embedding,
    temperature=1.0,
    uncond_scale=7.5,
)

images = model.decode(samples)

Both samplers accept identical conditioning inputs and follow the same interface defined in DiffusionSampler, ensuring seamless interchangeability in production pipelines.

Summary

  • DDPM sampling implements the full stochastic reverse Markov chain with Gaussian noise at every step, requiring the complete diffusion schedule (typically 1000 steps) for optimal quality.
  • DDIM sampling uses an implicit non-Markovian formulation that enables deterministic generation when η = 0 and supports arbitrary sub-sampling schedules (10–50 steps) for faster inference.
  • The eta parameter in DDIM controls the trade-off between deterministic consistency and stochastic diversity, whereas DDPM is inherently random.
  • Both samplers inherit from DiffusionSampler and can be swapped without modifying the underlying LatentDiffusion model, as implemented in labml_nn/diffusion/stable_diffusion/sampler/.

Frequently Asked Questions

What is the main advantage of DDIM over DDPM?

DDIM allows for significantly faster inference by reducing the number of sampling steps from 1000 to as few as 10–50 while maintaining high image quality. Additionally, when configured with ddim_eta=0.0, it provides deterministic outputs that are reproducible across runs, unlike the inherently stochastic DDPM sampler.

Can I use DDIM with any pre-trained diffusion model?

Yes, DDIM sampling is compatible with any diffusion model trained with the standard DDPM objective because it uses the same noise prediction network ε_θ(x_t). The implementation in labml_nn/diffusion/stable_diffusion/sampler/ddim.py works directly with models like Stable Diffusion without requiring fine-tuning or architectural changes.

How does the ddim_eta parameter affect image generation?

The ddim_eta parameter (η) controls the amount of random noise added during the DDIM sampling process. When η = 0, sampling is deterministic and produces identical outputs for the same inputs. As η increases toward 1.0, the sampler injects more stochastic noise, increasing output diversity but potentially reducing consistency with the conditioning signal.

Why does DDPM require more steps than DDIM?

DDPM relies on a Markov chain where each step only considers the immediate previous timestep, requiring small step sizes to maintain the Gaussian approximation of the reverse process. DDIM computes the trajectory direction toward the predicted clean image x̂₀ directly, allowing it to skip intermediate timesteps without accumulating approximation error, as the implicit formula accounts for the entire remaining diffusion path.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →