How Stable Diffusion Uses Latent Space for Image Generation: Technical Architecture Explained

Stable Diffusion runs its entire diffusion process in a compressed latent space learned by a variational autoencoder, encoding images to latent vectors, applying a text-conditioned U-Net to predict noise, then decoding back to pixel space.

The labmlai/annotated_deep_learning_paper_implementations repository provides a clean, annotated implementation of this workflow. By operating on latent tensors rather than full-resolution pixels, the model achieves significant computational efficiency while maintaining high-quality image generation capabilities.

Encoding Images to Latent Representations

The latent space workflow begins in labml_nn/diffusion/stable_diffusion/model/autoencoder.py. The Autoencoder.encode method (lines 49‑57) processes input images through a convolutional encoder network, followed by a learned 1×1 convolution layer named quant_conv. This convolution produces mean and log‑variance parameters that parameterize a GaussianDistribution.

During the forward pass, the distribution samples a latent tensor, which is then multiplied by a latent scaling factor (latent_scaling_factor) before being fed to the diffusion model. This scaling normalizes the latent representation for stable training dynamics. The compression typically reduces spatial dimensions to 1/8 of the original image resolution, transforming a 512×512 RGB image into a 64×64×4 latent tensor.

The Diffusion Process in Compressed Space

Rather than applying Gaussian diffusion to raw pixels, Stable Diffusion performs all noising and denoising steps within the latent space defined by the autoencoder. The core architecture resides in labml_nn/diffusion/stable_diffusion/latent_diffusion.py.

The LatentDiffusion class wraps a U‑Net (UNetModel) inside a DiffusionWrapper (lines 34‑48). This wrapper simply forwards the tuple (x, t, context) to the U‑Net, preserving the exact module hierarchy from the original CompVis checkpoint. During training and inference, LatentDiffusion.forward (lines 36‑45) calls this wrapped model to predict noise on the latent tensor at each timestep.

This approach reduces memory consumption and computational cost dramatically compared to pixel‑space diffusion, as the U‑Net processes 64×64 tensors instead of 512×512 images.

Conditioning and Sampling Workflow

Text Conditioning with CLIP Embeddings

Text prompts are converted into conditioning vectors via CLIPTextEmbedder defined in labml_nn/diffusion/stable_diffusion/model/clip_embedder.py. The LatentDiffusion.get_text_conditioning method (lines 13‑17) returns CLIP embeddings that serve as the context c for the diffusion U‑Net.

These embeddings guide the denoising process, ensuring the generated latent tensor aligns with the semantic content of the prompt.

Sampling Latent Tensors with DDIM

The actual generation process uses samplers implemented in labml_nn/diffusion/stable_diffusion/sampler/ddim.py and ddpm.py. A sampler like DDIMSampler draws latent tensors from the learned diffusion trajectory, receiving both the text conditioning (cond) and an optional unconditional conditioning (un_cond) for classifier‑free guidance.

In Txt2Img.__call__ (lines 92‑96 in labml_nn/diffusion/stable_diffusion/scripts/text_to_image.py), the sampler receives the target shape—typically [batch_size, 4, 64, 64]—and iteratively denoises random Gaussian noise into a structured latent representation guided by the text embeddings.

Decoding Latents to Pixel Space

After sampling completes, the final step converts the latent tensor back into viewable pixels. The LatentDiffusion.autoencoder_decode method (lines 28‑34) first divides the latent tensor by the scaling factor applied during encoding, then passes it to Autoencoder.decode (lines 62‑71).

The decoder network, implemented in labml_nn/diffusion/stable_diffusion/model/autoencoder.py, runs a series of upsampling convolutions to reconstruct the pixel‑space image. This reconstruction restores the full resolution while preserving the semantic content and fine details encoded in the compressed latent representation.

End-to-End Implementation Example

The following example demonstrates the complete workflow using the repository’s utilities:

from pathlib import Path
from labml_nn.diffusion.stable_diffusion.latent_diffusion import LatentDiffusion
from labml_nn.diffusion.stable_diffusion.sampler.ddim import DDIMSampler
from labml_nn.diffusion.stable_diffusion.util import load_model, save_images, set_seed
import torch

# 1️⃣ Load a pretrained checkpoint

model: LatentDiffusion = load_model(
    Path('~/labml-data/stable-diffusion/sd-v1-4.ckpt')
).to('cuda' if torch.cuda.is_available() else 'cpu')

# 2️⃣ Prepare text conditioning

prompt = "a futuristic cityscape at sunset"
cond = model.get_text_conditioning([prompt])

# 3️⃣ Initialise the sampler (DDIM with 50 steps)

sampler = DDIMSampler(model, n_steps=50, ddim_eta=0.0)

# 4️⃣ Sample a latent tensor (batch size=1, 4 channels, 64×64 latent size)

latent = sampler.sample(
    cond=cond,
    shape=[1, 4, 64, 64],
    uncond_scale=7.5,
    uncond_cond=model.get_text_conditioning([""])
)

# 5️⃣ Decode the latent to an image

image = model.autoencoder_decode(latent)

# 6️⃣ Save the result

save_images(image, 'outputs', prefix='demo_')

For command‑line usage, the repository provides the Txt2Img wrapper in labml_nn/diffusion/stable_diffusion/scripts/text_to_image.py:

python -m labml_nn.diffusion.stable_diffusion.scripts.text_to_image \
    --prompt "an astronaut riding a horse on Mars" \
    --sampler ddim \
    --steps 50 \
    --scale 7.5 \
    --batch_size 4

Summary

  • Latent encoding in Autoencoder.encode compresses images via a variational autoencoder with a learned quant_conv layer and GaussianDistribution sampling.
  • Diffusion occurs in latent space through the DiffusionWrapper and UNetModel, processing 64×64×4 tensors rather than full‑resolution pixels.
  • Text guidance comes from CLIPTextEmbedder via LatentDiffusion.get_text_conditioning, feeding semantic context to the U‑Net.
  • Sampling algorithms like DDIMSampler iteratively denoise latents, utilizing classifier‑free guidance with conditional and unconditional prompts.
  • Latent decoding via LatentDiffusion.autoencoder_decode and Autoencoder.decode restores the scaling factor and reconstructs the final pixel image.

Frequently Asked Questions

What compression ratio does Stable Diffusion's latent space use?

According to the implementation in labml_nn/diffusion/stable_diffusion/model/autoencoder.py, Stable Diffusion compresses images to approximately 1/8 of their original spatial resolution. A 512×512 pixel image encodes to a 64×64 latent tensor with 4 channels, resulting in a 48× reduction in spatial dimensions.

Why does Stable Diffusion use latent space instead of pixel space?

Operating in latent space dramatically reduces computational cost and memory requirements. The U‑Net processes smaller tensors (64×64 instead of 512×512), allowing higher resolution generation on consumer hardware. As implemented in labml_nn/diffusion/stable_diffusion/latent_diffusion.py, this approach maintains generation quality while enabling real‑time inference.

How are text prompts converted for the diffusion model?

The CLIPTextEmbedder class in labml_nn/diffusion/stable_diffusion/model/clip_embedder.py tokenizes text prompts and processes them through OpenAI’s CLIP transformer. The LatentDiffusion.get_text_conditioning method (lines 13‑17) returns these embeddings, which the DiffusionWrapper passes to the U‑Net as the conditioning context c.

What is the purpose of the latent scaling factor?

The latent scaling factor (latent_scaling_factor) applied in Autoencoder.encode (lines 49‑57) and removed in LatentDiffusion.autoencoder_decode (lines 28‑34) normalizes the variance of latent tensors. This scaling ensures stable training dynamics for the diffusion model and prevents numerical instability during the noising and denoising process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →