How to Use the Diffusers Integration with SanaPipeline
The SanaPipeline from NVlabs/Sana is a native diffusers-compatible pipeline that enables standard .from_pretrained() loading and inference via the familiar pipe() API, requiring only a conversion step for native checkpoints.
The NVlabs/Sana repository provides a high-performance text-to-image generation model designed to work seamlessly within the 🤗 diffusers ecosystem. By leveraging the diffusers integration with SanaPipeline, you can instantiate models using standard diffusers patterns, convert native PyTorch checkpoints, and customize inference with scheduler substitutions. This integration follows the standard diffusers architecture of VAE, text encoder, transformer, and scheduler components.
Architecture of the SanaPipeline Diffusers Integration
The Sana pipeline implements the standard diffusers contract, making it interchangeable with other text-to-image pipelines. Understanding this architecture is essential for customizing components or debugging inference.
Core Component Mapping
The pipeline composes four primary components that map directly to diffusers classes:
- VAE: Implements latent encoding/decoding using
AutoencoderDCfrom diffusers, responsible for converting between pixel space and latent space via thevae_decodemethod. - Text Encoder: Wraps
AutoModelForCausalLMfrom HuggingFace, tokenized viaself.tokenizerto produce text embeddings. - Transformer: The core Sana model (
SanaTransformer2DModelfrom diffusers) generates latent updates conditioned on text embeddings and timestep information. - Scheduler: Supports
DPMSolverMultistepScheduler,FlowMatchEulerDiscreteScheduler, orSCMSchedulerdepending on the sampling strategy (Flow-Euler, DPMSolver, or SCM for Sana-Sprint variants).
Initialization and Forward Pass Implementation
In app/sana_pipeline.py, the SanaPipeline.__init__ method orchestrates component initialization using pyrallis configuration management:
- Loads hyperparameters into a
SanaInferencedataclass. - Constructs components via
build_vae(),build_text_encoder(), andbuild_sana_model(). - Pre-computes null caption embeddings for classifier-free guidance.
The SanaPipeline.forward method (lines 66‑78 in app/sana_pipeline.py) implements the standard diffusion loop:
- Tokenizes prompts and generates text embeddings via
self.text_encoder. - Initializes latents either from noise or provided inputs.
- Instantiates the solver (either
FlowEulerorDPMSolverMultistepScheduler) based onself.vis_sampler. - Runs the denoising scheduler for
num_inference_stepsiterations. - Decodes final latents through the VAE, optionally resizing via
resize_and_crop_tensor.
Converting Native Checkpoints to Diffusers Format
Native Sana checkpoints (.pth files) require conversion to the diffusers directory structure before using the standard API.
Using the Conversion Script
The repository provides tools/convert_scripts/convert_sana_to_diffusers.py to handle parameter remapping and pipeline packaging:
python tools/convert_scripts/convert_sana_to_diffusers.py \
--orig_ckpt_path=/path/to/Sana_1600M_1024px.pth \
--image_size=1024 \
--model_type=SanaMS_1600M_P1_D20 \
--scheduler_type=flow-dpm_solver \
--dump_path=./my_sana_pipeline \
--save_full_pipeline \
--dtype=fp16
The script performs the following operations:
- Loads the native checkpoint using
torch.load()and remaps parameter names to the diffusers naming convention. - Instantiates
SanaTransformer2DModelwith appropriatemodel_kwargs, optionally usingaccelerateempty-weight initialization for memory efficiency. - Constructs the VAE via
AutoencoderDC.from_pretrained()and text encoder viaAutoModelForCausalLM. - Selects the appropriate scheduler based on
model_type(SCM for Sana-Sprint, otherwise DPMSolver or Flow-Euler). - Packages components into a
SanaPipeline(orSanaSprintPipeline) and exports viasave_pretrained().
After execution, ./my_sana_pipeline contains a valid diffusers pipeline with config.json, transformer/, vae/, and scheduler/ subdirectories.
Loading and Running Inference
Once converted, the pipeline follows standard diffusers usage patterns.
Loading a Pre-converted Pipeline
Import the pipeline class and load from the converted directory:
from diffusers import SanaPipeline
import torch
pipe = SanaPipeline.from_pretrained(
"./my_sana_pipeline", # or HuggingFace hub path
torch_dtype=torch.float16,
)
pipe.to("cuda")
Standard Text-to-Image Generation
The __call__ method accepts standard diffusers arguments that map 1-to-1 to the SanaPipeline.forward signature:
image = pipe(
prompt="A futuristic cityscape at sunset with flying vehicles",
height=1024,
width=1024,
num_inference_steps=20,
guidance_scale=4.5,
pag_guidance_scale=1.0,
).images[0]
image.save("output.png")
Key parameters include:
guidance_scale: Controls classifier-free guidance strength.pag_guidance_scale: Implements perturbed attention guidance specific to Sana.num_inference_steps: Denoising steps (20-30 recommended for Flow-Euler).
Accessing Log Probabilities for Research
For uncertainty quantification or research applications requiring intermediate states, use the log-probability-aware variant in diffusion/post_training/diffusers_patch/pipeline_with_logprob.py:
from diffusion.post_training.diffusers_patch import pipeline_with_logprob
image, latents, log_probs = pipeline_with_logprob.pipeline_with_logprob_sana(
transformer=pipe.transformer,
vae=pipe.vae,
prompt_embeds=pipe.text_encoder("A futuristic cityscape"),
num_inference_steps=30,
guidance_scale=5.0,
solver="flow",
deterministic=False,
)
This returns the final PIL image, a list of intermediate latent tensors, and per-step log-probabilities compatible with reinforcement learning or direct preference optimization workflows.
Key Integration Files
When working with the diffusers integration, these files define the critical touchpoints:
| File | Purpose |
|---|---|
app/sana_pipeline.py |
Native pipeline implementation with build_vae, build_text_encoder, and the forward pass logic. |
tools/convert_scripts/convert_sana_to_diffusers.py |
Checkpoint conversion utility that remaps weights to SanaTransformer2DModel. |
diffusion/post_training/diffusers_patch/pipeline_with_logprob.py |
Research variant exposing pipeline_with_logprob_sana for log-probability extraction. |
train_scripts/train_dreambooth_lora_sana.py |
Fine-tuning example demonstrating LoRA adaptation within the diffusers framework. |
app/app_sana.py |
CLI wrapper demonstrating end-to-end pipeline loading and inference execution. |
Summary
- SanaPipeline implements the standard diffusers contract, enabling drop-in replacement with other text-to-image pipelines.
- Use
tools/convert_scripts/convert_sana_to_diffusers.pyto convert native.pthcheckpoints into loadable diffusers directories. - The pipeline supports multiple schedulers including DPMSolverMultistepScheduler, FlowMatchEulerDiscreteScheduler, and SCMScheduler.
- For research requiring intermediate computations, import
pipeline_with_logprob_sanafrom the diffusers patch module. - All standard diffusers parameters (
guidance_scale,num_inference_steps, etc.) map directly to the underlyingSanaPipeline.forward implementationinapp/sana_pipeline.py.
Frequently Asked Questions
How do I convert a native Sana checkpoint to diffusers format?
Use the convert_sana_to_diffusers.py script located in tools/convert_scripts/. Provide the original checkpoint path, specify the model type (e.g., SanaMS_1600M_P1_D20), and set the output directory via --dump_path. The script handles parameter remapping, instantiates the SanaTransformer2DModel, and packages the VAE, text encoder, and scheduler into a standard diffusers pipeline structure ready for from_pretrained() loading.
What scheduler options are supported by SanaPipeline?
According to the source code in app/sana_pipeline.py, the pipeline supports three scheduler configurations: DPMSolverMultistepScheduler for standard DPM solving, FlowMatchEulerDiscreteScheduler for Flow-Euler sampling, and SCMScheduler specifically for Sana-Sprint models. The conversion script automatically selects the appropriate scheduler based on the --model_type argument, or you can manually override using the --scheduler_type flag.
How can I access intermediate latents and log probabilities during inference?
Import the pipeline_with_logprob_sana function from diffusion/post_training/diffusers_patch/pipeline_with_logprob.py. This function accepts the same core components (transformer, VAE, prompt embeddings) as the standard pipeline but returns a tuple containing the final image, a list of intermediate latent tensors, and per-step log-probabilities. This is particularly useful for research applications involving direct preference optimization or uncertainty quantification.
What is the difference between the native Sana pipeline and the diffusers implementation?
The native implementation in app/sana_pipeline.py uses pyrallis for configuration management and includes Sana-specific utilities like build_sana_model(). When converted via convert_sana_to_diffusers.py, the model becomes a standard diffusers pipeline that can be imported with from diffusers import SanaPipeline. The converted version uses the exact same weights and architecture but follows the diffusers API contract, enabling interoperability with the broader ecosystem of diffusers tools, schedulers, and training scripts like train_dreambooth_lora_sana.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →