How to Implement VAE Encoding and Decoding Nodes in ComfyUI

ComfyUI implements VAE encoding and decoding through a declarative node architecture in nodes.py that wraps the core VAE class in comfy/sd.py, supporting both standard and tiled operations for high-resolution images and video.

Implementing VAE encoding and decoding nodes in ComfyUI requires understanding how the framework treats Variational Auto-Encoders as first-class data types that flow between nodes. The implementation spans declarative node definitions in nodes.py and the heavy-duty tensor processing logic in comfy/sd.py. This guide breaks down the exact implementation patterns used in the Comfy-Org/ComfyUI repository.

Understanding the VAE Node Architecture in nodes.py

All VAE nodes in ComfyUI follow a strict declarative protocol that separates UI metadata from execution logic.

The Declarative Node Protocol

Every VAE node class defines three critical class attributes that the ComfyUI engine uses to render sockets and route execution:

class VAEDecode:
    @classmethod
    def INPUT_TYPES(s):
        return {
            "required": {
                "samples": ("LATENT", {"tooltip": "The latent to be decoded."}),
                "vae": ("VAE", {"tooltip": "The VAE model used for decoding the latent."})
            }
        }
    RETURN_TYPES = ("IMAGE",)
    FUNCTION = "decode"
    CATEGORY = "latent"

INPUT_TYPES declares the input sockets, RETURN_TYPES specifies outputs, and FUNCTION maps the node to a Python method. The actual decode method remains a thin wrapper around the VAE object:

def decode(self, vae, samples):
    latent = samples["samples"]
    if latent.is_nested:  # handle batched latent tensors

        latent = latent.unbind()[0]
    images = vae.decode(latent)  # delegate to VAE class

    if len(images.shape) == 5:   # flatten video batches

        images = images.reshape(-1, images.shape[-3],
                               images.shape[-2], images.shape[-1])
    return (images,)

Core VAE Nodes and Their Implementation

The nodes.py file (lines 293-384) defines five primary VAE nodes that map directly to methods in the VAE class:

  • VAEDecode: Calls vae.decode(latent) to convert latent tensors to RGB images
  • VAEEncode: Calls vae.encode(pixels) to compress images into latent space
  • VAEDecodeTiled: Invokes vae.decode_tiled(...) for memory-efficient decoding of large images
  • VAEEncodeTiled: Invokes vae.encode_tiled(...) for tiled encoding operations
  • VAEEncodeForInpaint: Wraps vae.encode() with mask preprocessing for inpainting workflows (lines 94-124)

These nodes are registered in the global NODE_CLASS_MAPPINGS dictionary near line 2040, making them discoverable by the UI and API.

The VAE Class Implementation in comfy/sd.py

The VAE class in comfy/sd.py (lines 370-440) encapsulates all model-specific logic, checkpoint handling, and tensor transformations.

Checkpoint Handling and Model Detection

The constructor automatically detects VAE formats by probing for signature keys:

class VAE:
    def __init__(self, sd=None, device=None, config=None, dtype=None, metadata=None):
        # Auto-convert Diffusers format checkpoints

        if 'decoder.up_blocks.0.resnets.0.norm1.weight' in sd.keys():
            sd = diffusers_convert.convert_vae_state_dict(sd)
        
        self.downscale_ratio = 8
        self.upscale_ratio = 8
        self.latent_channels = 4
        self.output_channels = 3
        self.process_input = lambda img: img * 2.0 - 1.0  # Scale to [-1, 1]

        self.process_output = lambda img: torch.clamp((img + 1.0) / 2.0, 0.0, 1.0)

The class recognizes multiple VAE families including standard Stable Diffusion, TAESD (Tiny AutoEncoder), and Stable Cascade by checking for keys like "taesd_decoder.1.weight" or "decoder.conv_in.weight".

Encode and Decode Methods

The core methods handle the forward and inverse transformations between pixel and latent space:

def encode(self, pixels):
    latent = self.first_stage_model.encode(self.process_input(pixels))
    return latent

def decode(self, latent):
    images = self.first_stage_model.decode(latent)
    return self.process_output(images)

Where self.first_stage_model is an instance of AutoencodingEngine or a TAESD wrapper. The process_input lambda scales images from [0, 1] to [-1, 1] range expected by the encoder, while process_output clamps decoder outputs back to valid RGB ranges.

Tiled Operations for High-Resolution Processing

For memory-constrained environments, the VAE class provides tiled variants that process images in overlapping patches:

def encode_tiled(self, pixel_samples, tile_x=None, tile_y=None, overlap=None,
                 tile_t=None, overlap_t=None):
    # Compute optimal tile size based on compression ratio

    return self.first_stage_model.encode_tiled(...)

def decode_tiled(self, latent, tile_x, tile_y, overlap, tile_t=None, overlap_t=None):
    # Handle spatial and temporal tiling for video batches

    return self.first_stage_model.decode_tiled(...)

These methods enable processing of 4K images or long video sequences without out-of-memory errors by splitting tensors along spatial dimensions (and temporal dimension tile_t for video).

Loading VAE Models with VAELoader

The VAELoader node (lines 728-776 in nodes.py) handles checkpoint discovery and initialization:

class VAELoader:
    @staticmethod
    def vae_list(s):
        # Merge exact checkpoint files with approximate TAESD families

        return files + taesd_files
    
    @staticmethod
    def load_taesd(name):
        # Assemble encoder/decoder from separate files (taesd_encoder.pt, taesd_decoder.pt)

        # Inject scale/shift constants required by the engine

When executed, the loader calls comfy.sd.VAE(sd=sd, metadata=metadata), passing the merged state dictionary. The node outputs a VAE type that subsequent encoding/decoding nodes consume.

Practical Implementation Examples

API Workflow JSON

Below is a complete round-trip workflow demonstrating VAE encoding and decoding via the ComfyUI API:

{
  "1": {
    "class_type": "VAELoader",
    "inputs": { "vae_name": "taesd" }
  },
  "2": {
    "class_type": "LoadImage",
    "inputs": { "image_path": "inputs/example.png" }
  },
  "3": {
    "class_type": "VAEEncode",
    "inputs": {
      "pixels": ["2", 0],
      "vae": ["1", 0]
    }
  },
  "4": {
    "class_type": "VAEDecode",
    "inputs": {
      "samples": ["3", 0],
      "vae": ["1", 0]
    }
  },
  "5": {
    "class_type": "SaveImage",
    "inputs": {
      "images": ["4", 0],
      "filename_prefix": "vae_roundtrip"
    }
  }
}

This JSON structure corresponds to the example in script_examples/basic_api_example.py (lines 73-84).

Python Script Implementation

For programmatic use outside the node system:

import torch
from comfy.sd import VAE
from comfy.utils import load_image, save_image, load_torch_file

# 1. Load VAE checkpoint

sd = load_torch_file("models/vae/taesdxl.pt", safe_load=True)
vae = VAE(sd=sd)  # Initialize the engine

# 2. Load source image [1, H, W, 3] in [0,1] range

image = load_image("inputs/portrait.png")

# 3. Encode to latent space

latent = vae.encode(image)

# 4. Decode back to RGB

reconstructed = vae.decode(latent)

# 5. Save output

save_image(reconstructed, "outputs/reconstructed.png")

Tiled Video Processing

For video workflows with 5D tensors [B, T, C, H, W]:


# Decode with both spatial and temporal tiling

decoded = vae.decode_tiled(
    latent_video,
    tile_x=512, tile_y=512, overlap=64,
    tile_t=64, overlap_t=8
)

Summary

  • VAE nodes in ComfyUI follow a declarative pattern using INPUT_TYPES, RETURN_TYPES, and FUNCTION class attributes defined in nodes.py.
  • The VAE class in comfy/sd.py handles format detection, tensor scaling, and provides both standard and tiled encode/decode methods.
  • Tiled operations (encode_tiled, decode_tiled) prevent out-of-memory errors when processing high-resolution images or video sequences.
  • VAELoader manages checkpoint discovery including special handling for TAESD models that split encoder and decoder into separate files.
  • Nodes pass VAE objects as typed references, allowing different VAE models to be chained within the same workflow graph.

Frequently Asked Questions

What is the difference between VAEEncode and VAEEncodeTiled?

VAEEncode processes entire images in a single forward pass through vae.encode(), requiring sufficient GPU memory for the full resolution. VAEEncodeTiled splits large images into overlapping tiles (configurable via tile_x, tile_y, and overlap parameters) and processes them sequentially through vae.encode_tiled(), significantly reducing memory usage at the cost of slightly longer processing time.

How does ComfyUI handle different VAE formats like TAESD?

The VAE class constructor detects TAESD checkpoints by probing for keys like "taesd_decoder.1.weight". The VAELoader node assembles TAESD models from separate encoder and decoder files (e.g., taesd_encoder.pt and taesd_decoder.pt) and injects required scale/shift constants. The framework also auto-converts Diffusers-format checkpoints by detecting keys like "decoder.up_blocks.0.resnets.0.norm1.weight" and applying diffusers_convert.convert_vae_state_dict().

Can I use different VAEs for encoding and decoding in the same workflow?

Yes. Because ComfyUI treats VAE as a distinct data type that flows through graph edges, you can route the output of one VAELoader into a VAEEncode node and a different VAELoader into a subsequent VAEDecode node. This enables workflows such as encoding with a lightweight TAESD for speed, then decoding with a full-quality SD-VAE for final output quality.

What are the input and output tensor formats for VAE nodes?

Inputs: VAEEncode expects IMAGE tensors with shape [batch, height, width, 3] and values in [0, 1] range. VAEDecode expects LATENT dictionaries containing a "samples" key with tensors of shape [batch, channels, height//8, width//8]. Outputs: VAEEncode returns LATENT dictionaries, while VAEDecode returns IMAGE tensors of shape [batch, height, width, 3]. Video inputs use 5D tensors [batch, frames, channels, height, width] which the nodes flatten appropriately.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →