How the Latent Preview System with TAESD Works in ComfyUI

TAESD (Tiny AutoEncoder for Stable Diffusion) enables real-time preview images during diffusion by decoding latent tensors through a lightweight auto-encoder approximation, bypassing the memory-heavy full VAE.

The latent preview system with TAESD allows ComfyUI users to visualize generation progress without waiting for the final VAE decode. This system leverages a compact neural network to rapidly convert latent representations into viewable images, providing immediate visual feedback while the diffusion sampler runs.

Architecture Overview

The preview system operates as a pluggable backend that intercepts latent tensors mid-generation. When activated via the --preview-method taesd CLI argument, ComfyUI instantiates a specialized decoder that trades fidelity for speed, consuming significantly less VRAM than the standard VAE decoder.

The architecture follows a factory pattern where get_previewer() selects the appropriate implementation based on the latent format and user preferences, returning either TAESDPreviewerImpl for standard image generation or TAEHVPreviewerImpl for video workflows.

Core Components

CLI Configuration and Method Selection

The preview method is defined in comfy/cli_args.py through the LatentPreviewMethod enum (lines 94-99). Users activate TAESD by launching ComfyUI with:

comfyui --preview-method taesd

This sets args.preview_method to LatentPreviewMethod.TAESD. Developers can also override this programmatically using set_preview_method() in latent_preview.py (lines 30-37), which updates the global configuration object to switch preview modes dynamically during execution.

Previewer Instantiation via get_previewer()

The function get_previewer(device, latent_format) in latent_preview.py (lines 78-105) serves as the factory for preview objects. It reads the current args.preview_method, locates the TAESD model weights in the vae_approx/ directory using folder_paths.py, and constructs the appropriate previewer instance based on the latent format's properties.

For video latents matching names in VIDEO_TAES, the system instantiates TAEHVPreviewerImpl; otherwise, it creates TAESDPreviewerImpl wrapping the standard TAESD model.

TAESD Decoder Loading and Model Selection

The decoder initialization logic (lines 96-103) distinguishes between two code paths:

  • Video Latents: Loads a lightweight comfy.sd.VAE instance and wraps it in TAEHVPreviewerImpl, skipping channel scaling since video decoders output normalized [0, 1] values directly.
  • Standard Latents: Instantiates comfy.taesd.taesd.TAESD and moves it to the execution device, preparing it for single-sample decoding.

This conditional loading ensures optimal memory usage across different generation modalities without loading unnecessary parameters.

Decoding and Image Generation

The TAESDPreviewerImpl class (lines 39-46) implements decode_latent_to_preview(), which processes the latent tensor x0 through three stages:

  1. Sample Extraction: Decodes only the first sample of the batch to minimize computation
  2. Channel Reordering: Calls movedim(0, 2) to transform the tensor from CHW to HWC format
  3. Image Conversion: Passes the result to preview_to_image() for final rasterization

For video previews, TAEHVPreviewerImpl (lines 47-50) extracts only the first channel (x0[:1, :, :1]) and bypasses the [-1, 1] normalization step.

End-to-End Execution Flow

When a sampler node initiates, prepare_callback() in latent_preview.py (lines 12-28) constructs the previewer once per run. The execution follows this pipeline:

  1. Initialization: get_previewer() resolves the device and latent format, loading the TAESD model from vae_approx/ into TAESDPreviewerImpl
  2. Step Callback: On every diffusion step, the sampler provides the current latent x0 to the previewer
  3. Fast Decoding: The previewer calls taesd.decode() on the latent sample, reordering channels via movedim(0, 2)
  4. Normalization: preview_to_image() (lines 16-30) scales float values from [-1, 1] to [0, 255], converts to uint8, and generates a PIL Image
  5. Streaming: The resulting JPEG bytes transmit to the frontend progress bar, updating the visual preview in real-time

This pipeline executes independently of the main generation thread, ensuring preview updates do not block the diffusion process.

Code Examples

Enabling TAESD via Command Line

Activate the latent preview system with TAESD at launch:


# Enable TAESD previews with 512px preview size

python main.py --preview-method taesd --preview-size 512

This configures args.preview_method globally for the session.

Programmatic Method Override

Override the preview method dynamically within custom nodes or scripts:

from comfy.latent_preview import set_preview_method

# Force TAESD for the next sampler execution

set_preview_method("taesd")

# ... execute sampler node ...

# Restore default behavior

set_preview_method("default")

The set_preview_method() function updates the global args.preview_method enum value in latent_preview.py.

Manual Previewer Creation

For custom sampling scripts outside the standard node graph:

import torch
from comfy.latent_preview import get_previewer
from comfy.model_management import get_torch_device

device = get_torch_device()

# latent_format is obtained from the loaded checkpoint/model

previewer = get_previewer(device, latent_format)

# latent_tensor shape: [batch, channels, height, width]

preview_image = previewer.decode_latent_to_preview(latent_tensor)
preview_image.show()

The get_previewer() call automatically selects TAESDPreviewerImpl when the latent format specifies taesd_decoder_name and the preview method is set to TAESD.

Key Source Files

File Purpose
comfy/cli_args.py Defines LatentPreviewMethod enum and parses --preview-method argument
latent_preview.py Core preview system containing get_previewer(), TAESDPreviewerImpl, and preview_to_image()
comfy/taesd/taesd.py TAESD model architecture implementing the lightweight auto-encoder
comfy/sd.py Provides fallback VAE class for video-TAE preview paths
folder_paths.py Locates model files in vae_approx/ directory

Summary

  • TAESD provides fast, low-memory previews by approximating the full VAE with a tiny auto-encoder
  • The system activates via --preview-method taesd or programmatic set_preview_method() calls
  • get_previewer() in latent_preview.py acts as a factory selecting between TAESDPreviewerImpl and TAEHVPreviewerImpl
  • Video latents bypass normalization and decode only the first channel, while standard latents decode the first sample with channel reordering
  • preview_to_image() converts decoded tensors from [-1, 1] float range to [0, 255] uint8 PIL Images for frontend display

Frequently Asked Questions

What is the difference between TAESD and the standard VAE decoder?

TAESD is a distilled, lightweight auto-encoder specifically trained to approximate VAE outputs with significantly fewer parameters. While the standard VAE provides maximum quality for final outputs, TAESD sacrifices minimal fidelity for substantial speed gains and lower VRAM usage, making it ideal for real-time previews during the diffusion process.

Why does the video previewer use a different implementation?

TAEHVPreviewerImpl handles video latents differently because video VAEs (TAE-HV) output values already normalized to [0, 1] range, unlike standard image VAEs that output [-1, 1]. Additionally, video previews extract only the first temporal channel (x0[:1, :, :1]) to generate a single representative frame rather than processing the entire temporal sequence.

Can I use TAESD previews with custom latent formats?

Yes, provided the latent_format object passed to get_previewer() includes a valid taesd_decoder_name property pointing to a compatible model in the vae_approx/ directory. The system automatically detects whether the format matches video TAE names (VIDEO_TAES) and instantiates the appropriate previewer implementation accordingly.

How does the preview system impact generation performance?

The TAESD decoder adds minimal overhead because it processes only the first sample of each batch and uses a highly optimized network architecture. The preview generation runs asynchronously to the diffusion sampler, ensuring that preview decoding does not block the main generation thread or significantly increase step time.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →