Deploying Sana with ComfyUI for Visual Workflow Editing
Sana integrates with ComfyUI through custom nodes that wrap the checkpoint loader, Gemma-2 text encoder, VAE, and Flow-Euler sampler, enabling node-based visual composition of high-resolution diffusion pipelines.
Deploying Sana with ComfyUI transforms NVIDIA's efficient diffusion model into a visual editing environment where you construct generation pipelines by connecting nodes. The NVlabs/Sana repository provides pre-configured JSON workflows and leverages the ComfyUI_ExtraModels custom node pack to expose Sana's native capabilities—Checkpoint loading, Gemma-2 text encoding, and DC-AE decoding—within ComfyUI's graph interface.
Installation and Prerequisites
To begin deploying Sana with ComfyUI, you must install the base ComfyUI framework followed by the extra models repository that contains Sana-specific node definitions.
Install ComfyUI and the Sana custom nodes in your custom_nodes directory:
# 1. Clone ComfyUI
git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI
# 2. Install Sana extra-models (contains custom node definitions)
git clone https://github.com/lawrence-cj/ComfyUI_ExtraModels.git custom_nodes/ComfyUI_ExtraModels
# 3. Launch the UI
python main.py
The custom nodes are automatically discovered when placed under ComfyUI/custom_nodes/ComfyUI_ExtraModels. No additional pip requirements are specified in the analysis, though standard ComfyUI dependencies apply.
The Sana Node Pipeline Architecture
The integration implements a seven-stage pipeline that mirrors Sana's native inference path. Understanding this data flow is critical for customizing workflows or troubleshooting generation issues.
Model Loading and Configuration
The workflow begins with SanaCheckpointLoader, which pulls the Sana model checkpoint (e.g., Efficient-Large-Model/Sana_1600M_1024px_BF16) and requires the matching architecture identifier SanaMS_1600M_P1_D20 to initialize the model structure correctly.
Simultaneously, the GemmaLoader node initializes the Gemma-2 2B text encoder. This is the specific encoder used by Sana models, loaded separately from the diffusion backbone to allow shared memory management and efficient scheduling.
Resolution Selection and Latent Creation
The SanaResolutionSelect node determines the target output dimensions (e.g., 1024px or 4096px) and passes standardized width/height integers to the EmptySanaLatentImage node. This node creates an empty latent tensor of the requested size, which serves as the initial noise input for the diffusion process.
Conditioning and Sampling
Prompts flow through GemmaTextEncode (or the alternative SanaTextEncode for Sana-specific conditioning) to produce embedding tensors compatible with the model's attention layers. These conditioning tensors feed into the standard KSampler node, configured to use the Flow-Euler sampler as implemented in Sana's inference pipeline.
Decoding and Preview
The final latent representation passes to VAEDecode, which utilizes the DC-AE VAE (mit-han-lab/dc-ae-f32c32-sana-1.1-diffusers) to convert latent tensors into RGB images. The PreviewImage node then renders the result within the ComfyUI canvas.
Loading Sample Workflows
The NVlabs/Sana repository ships with ready-to-use workflow files located in docs/ComfyUI/. These JSON files pre-wire the node connections described above.
To load a workflow:
- Click "Load workflow" in the ComfyUI interface.
- Select either
Sana_FlowEuler.json(for 1024px generation) orSana_FlowEuler_4K.json(for 4096px generation) from the repository. - Edit the prompt inside the
SanaTextEncodeorGemmaTextEncodenode (default: "a dog and a cat"). - Press "Queue Prompt" to trigger automatic checkpoint downloads and begin generation.
The workflow files serve as the definitive reference for node wiring and parameter defaults, as detailed in docs/ComfyUI/comfyui.md.
Hardware Requirements for 4K Generation
While the standard 1024px workflow runs on consumer GPUs, deploying the 4K workflow requires substantial VRAM. The Sana_FlowEuler_4K.json workflow sets SanaResolutionSelect to 4096px and requires at least 18GB of GPU memory to execute without offloading. Ensure your hardware configuration meets this threshold before attempting high-resolution generation.
Model Conversion and Pipeline Integration
Under the hood, the ComfyUI nodes interface with the core Sana implementation found in app/sana_pipeline.py. This file contains the Python classes that handle the actual diffusion scheduling and noise prediction.
For checkpoints not distributed in Diffusers format, use the conversion utility at tools/convert_scripts/convert_sana_to_diffusers.py. This script transforms native Sana checkpoints into the structure expected by SanaCheckpointLoader, ensuring compatibility with the ComfyUI node system.
Summary
- Deploying Sana with ComfyUI requires installing the
ComfyUI_ExtraModelsrepository intoComfyUI/custom_nodes/to access Sana-specific nodes likeSanaCheckpointLoaderandGemmaLoader. - The node pipeline follows seven stages: checkpoint loading, text encoding, resolution selection, latent creation, sampling via Flow-Euler, DC-AE decoding, and preview.
- Pre-configured workflows are available in
docs/ComfyUI/Sana_FlowEuler.json(1024px) andSana_FlowEuler_4K.json(4096px). - The 4K workflow demands 18GB GPU VRAM and utilizes the same node architecture scaled to higher resolutions.
- Core pipeline logic is implemented in
app/sana_pipeline.py, with conversion utilities available intools/convert_scripts/convert_sana_to_diffusers.pyfor checkpoint format standardization.
Frequently Asked Questions
What GPU memory is required for Sana ComfyUI workflows?
The 1024px generation workflow runs on standard consumer GPUs with 8–12GB VRAM, while the 4K workflow configured in Sana_FlowEuler_4K.json requires a minimum of 18GB GPU memory to process 4096px latents without model offloading.
Where are the custom node definitions for Sana located?
The Python node implementations reside in the ComfyUI_ExtraModels repository, which must be cloned into ComfyUI/custom_nodes/ComfyUI_ExtraModels. The JSON workflow files that wire these nodes together are shipped in the main NVlabs/Sana repository under docs/ComfyUI/.
How do I convert Sana checkpoints for ComfyUI compatibility?
Use the convert_sana_to_diffusers.py script located in tools/convert_scripts/ within the Sana repository. This utility converts native Sana checkpoints into the Diffusers format required by the SanaCheckpointLoader node, ensuring proper architecture mapping and weight loading.
Can I modify the text encoder configuration in the workflow?
Yes, the default workflows use GemmaTextEncode linked to the GemmaLoader node for the Gemma-2 2B encoder. You can substitute this with SanaTextEncode for direct Sana-specific conditioning, or modify the loader node parameters to point to alternative checkpoint variants while maintaining the same graph structure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →