Deploying Sana with ComfyUI for Visual Workflow Editing

Sana integrates with ComfyUI through custom nodes that wrap the checkpoint loader, Gemma-2 text encoder, VAE, and Flow-Euler sampler, enabling node-based visual composition of high-resolution diffusion pipelines.

Deploying Sana with ComfyUI transforms NVIDIA's efficient diffusion model into a visual editing environment where you construct generation pipelines by connecting nodes. The NVlabs/Sana repository provides pre-configured JSON workflows and leverages the ComfyUI_ExtraModels custom node pack to expose Sana's native capabilities—Checkpoint loading, Gemma-2 text encoding, and DC-AE decoding—within ComfyUI's graph interface.

Installation and Prerequisites

To begin deploying Sana with ComfyUI, you must install the base ComfyUI framework followed by the extra models repository that contains Sana-specific node definitions.

Install ComfyUI and the Sana custom nodes in your custom_nodes directory:


# 1. Clone ComfyUI

git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI

# 2. Install Sana extra-models (contains custom node definitions)

git clone https://github.com/lawrence-cj/ComfyUI_ExtraModels.git custom_nodes/ComfyUI_ExtraModels

# 3. Launch the UI

python main.py

The custom nodes are automatically discovered when placed under ComfyUI/custom_nodes/ComfyUI_ExtraModels. No additional pip requirements are specified in the analysis, though standard ComfyUI dependencies apply.

The Sana Node Pipeline Architecture

The integration implements a seven-stage pipeline that mirrors Sana's native inference path. Understanding this data flow is critical for customizing workflows or troubleshooting generation issues.

Model Loading and Configuration

The workflow begins with SanaCheckpointLoader, which pulls the Sana model checkpoint (e.g., Efficient-Large-Model/Sana_1600M_1024px_BF16) and requires the matching architecture identifier SanaMS_1600M_P1_D20 to initialize the model structure correctly.

Simultaneously, the GemmaLoader node initializes the Gemma-2 2B text encoder. This is the specific encoder used by Sana models, loaded separately from the diffusion backbone to allow shared memory management and efficient scheduling.

Resolution Selection and Latent Creation

The SanaResolutionSelect node determines the target output dimensions (e.g., 1024px or 4096px) and passes standardized width/height integers to the EmptySanaLatentImage node. This node creates an empty latent tensor of the requested size, which serves as the initial noise input for the diffusion process.

Conditioning and Sampling

Prompts flow through GemmaTextEncode (or the alternative SanaTextEncode for Sana-specific conditioning) to produce embedding tensors compatible with the model's attention layers. These conditioning tensors feed into the standard KSampler node, configured to use the Flow-Euler sampler as implemented in Sana's inference pipeline.

Decoding and Preview

The final latent representation passes to VAEDecode, which utilizes the DC-AE VAE (mit-han-lab/dc-ae-f32c32-sana-1.1-diffusers) to convert latent tensors into RGB images. The PreviewImage node then renders the result within the ComfyUI canvas.

Loading Sample Workflows

The NVlabs/Sana repository ships with ready-to-use workflow files located in docs/ComfyUI/. These JSON files pre-wire the node connections described above.

To load a workflow:

  1. Click "Load workflow" in the ComfyUI interface.
  2. Select either Sana_FlowEuler.json (for 1024px generation) or Sana_FlowEuler_4K.json (for 4096px generation) from the repository.
  3. Edit the prompt inside the SanaTextEncode or GemmaTextEncode node (default: "a dog and a cat").
  4. Press "Queue Prompt" to trigger automatic checkpoint downloads and begin generation.

The workflow files serve as the definitive reference for node wiring and parameter defaults, as detailed in docs/ComfyUI/comfyui.md.

Hardware Requirements for 4K Generation

While the standard 1024px workflow runs on consumer GPUs, deploying the 4K workflow requires substantial VRAM. The Sana_FlowEuler_4K.json workflow sets SanaResolutionSelect to 4096px and requires at least 18GB of GPU memory to execute without offloading. Ensure your hardware configuration meets this threshold before attempting high-resolution generation.

Model Conversion and Pipeline Integration

Under the hood, the ComfyUI nodes interface with the core Sana implementation found in app/sana_pipeline.py. This file contains the Python classes that handle the actual diffusion scheduling and noise prediction.

For checkpoints not distributed in Diffusers format, use the conversion utility at tools/convert_scripts/convert_sana_to_diffusers.py. This script transforms native Sana checkpoints into the structure expected by SanaCheckpointLoader, ensuring compatibility with the ComfyUI node system.

Summary

  • Deploying Sana with ComfyUI requires installing the ComfyUI_ExtraModels repository into ComfyUI/custom_nodes/ to access Sana-specific nodes like SanaCheckpointLoader and GemmaLoader.
  • The node pipeline follows seven stages: checkpoint loading, text encoding, resolution selection, latent creation, sampling via Flow-Euler, DC-AE decoding, and preview.
  • Pre-configured workflows are available in docs/ComfyUI/Sana_FlowEuler.json (1024px) and Sana_FlowEuler_4K.json (4096px).
  • The 4K workflow demands 18GB GPU VRAM and utilizes the same node architecture scaled to higher resolutions.
  • Core pipeline logic is implemented in app/sana_pipeline.py, with conversion utilities available in tools/convert_scripts/convert_sana_to_diffusers.py for checkpoint format standardization.

Frequently Asked Questions

What GPU memory is required for Sana ComfyUI workflows?

The 1024px generation workflow runs on standard consumer GPUs with 8–12GB VRAM, while the 4K workflow configured in Sana_FlowEuler_4K.json requires a minimum of 18GB GPU memory to process 4096px latents without model offloading.

Where are the custom node definitions for Sana located?

The Python node implementations reside in the ComfyUI_ExtraModels repository, which must be cloned into ComfyUI/custom_nodes/ComfyUI_ExtraModels. The JSON workflow files that wire these nodes together are shipped in the main NVlabs/Sana repository under docs/ComfyUI/.

How do I convert Sana checkpoints for ComfyUI compatibility?

Use the convert_sana_to_diffusers.py script located in tools/convert_scripts/ within the Sana repository. This utility converts native Sana checkpoints into the Diffusers format required by the SanaCheckpointLoader node, ensuring proper architecture mapping and weight loading.

Can I modify the text encoder configuration in the workflow?

Yes, the default workflows use GemmaTextEncode linked to the GemmaLoader node for the Gemma-2 2B encoder. You can substitute this with SanaTextEncode for direct Sana-specific conditioning, or modify the loader node parameters to point to alternative checkpoint variants while maintaining the same graph structure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →