How to Use the Diffusers Integration with SanaPipeline

The SanaPipeline from NVlabs/Sana is a native diffusers-compatible pipeline that enables standard .from_pretrained() loading and inference via the familiar pipe() API, requiring only a conversion step for native checkpoints.

The NVlabs/Sana repository provides a high-performance text-to-image generation model designed to work seamlessly within the 🤗 diffusers ecosystem. By leveraging the diffusers integration with SanaPipeline, you can instantiate models using standard diffusers patterns, convert native PyTorch checkpoints, and customize inference with scheduler substitutions. This integration follows the standard diffusers architecture of VAE, text encoder, transformer, and scheduler components.

Architecture of the SanaPipeline Diffusers Integration

The Sana pipeline implements the standard diffusers contract, making it interchangeable with other text-to-image pipelines. Understanding this architecture is essential for customizing components or debugging inference.

Core Component Mapping

The pipeline composes four primary components that map directly to diffusers classes:

  • VAE: Implements latent encoding/decoding using AutoencoderDC from diffusers, responsible for converting between pixel space and latent space via the vae_decode method.
  • Text Encoder: Wraps AutoModelForCausalLM from HuggingFace, tokenized via self.tokenizer to produce text embeddings.
  • Transformer: The core Sana model (SanaTransformer2DModel from diffusers) generates latent updates conditioned on text embeddings and timestep information.
  • Scheduler: Supports DPMSolverMultistepScheduler, FlowMatchEulerDiscreteScheduler, or SCMScheduler depending on the sampling strategy (Flow-Euler, DPMSolver, or SCM for Sana-Sprint variants).

Initialization and Forward Pass Implementation

In app/sana_pipeline.py, the SanaPipeline.__init__ method orchestrates component initialization using pyrallis configuration management:

  1. Loads hyperparameters into a SanaInference dataclass.
  2. Constructs components via build_vae(), build_text_encoder(), and build_sana_model().
  3. Pre-computes null caption embeddings for classifier-free guidance.

The SanaPipeline.forward method (lines 66‑78 in app/sana_pipeline.py) implements the standard diffusion loop:

  • Tokenizes prompts and generates text embeddings via self.text_encoder.
  • Initializes latents either from noise or provided inputs.
  • Instantiates the solver (either FlowEuler or DPMSolverMultistepScheduler) based on self.vis_sampler.
  • Runs the denoising scheduler for num_inference_steps iterations.
  • Decodes final latents through the VAE, optionally resizing via resize_and_crop_tensor.

Converting Native Checkpoints to Diffusers Format

Native Sana checkpoints (.pth files) require conversion to the diffusers directory structure before using the standard API.

Using the Conversion Script

The repository provides tools/convert_scripts/convert_sana_to_diffusers.py to handle parameter remapping and pipeline packaging:

python tools/convert_scripts/convert_sana_to_diffusers.py \
    --orig_ckpt_path=/path/to/Sana_1600M_1024px.pth \
    --image_size=1024 \
    --model_type=SanaMS_1600M_P1_D20 \
    --scheduler_type=flow-dpm_solver \
    --dump_path=./my_sana_pipeline \
    --save_full_pipeline \
    --dtype=fp16

The script performs the following operations:

  • Loads the native checkpoint using torch.load() and remaps parameter names to the diffusers naming convention.
  • Instantiates SanaTransformer2DModel with appropriate model_kwargs, optionally using accelerate empty-weight initialization for memory efficiency.
  • Constructs the VAE via AutoencoderDC.from_pretrained() and text encoder via AutoModelForCausalLM.
  • Selects the appropriate scheduler based on model_type (SCM for Sana-Sprint, otherwise DPMSolver or Flow-Euler).
  • Packages components into a SanaPipeline (or SanaSprintPipeline) and exports via save_pretrained().

After execution, ./my_sana_pipeline contains a valid diffusers pipeline with config.json, transformer/, vae/, and scheduler/ subdirectories.

Loading and Running Inference

Once converted, the pipeline follows standard diffusers usage patterns.

Loading a Pre-converted Pipeline

Import the pipeline class and load from the converted directory:

from diffusers import SanaPipeline
import torch

pipe = SanaPipeline.from_pretrained(
    "./my_sana_pipeline",  # or HuggingFace hub path

    torch_dtype=torch.float16,
)
pipe.to("cuda")

Standard Text-to-Image Generation

The __call__ method accepts standard diffusers arguments that map 1-to-1 to the SanaPipeline.forward signature:

image = pipe(
    prompt="A futuristic cityscape at sunset with flying vehicles",
    height=1024,
    width=1024,
    num_inference_steps=20,
    guidance_scale=4.5,
    pag_guidance_scale=1.0,
).images[0]

image.save("output.png")

Key parameters include:

  • guidance_scale: Controls classifier-free guidance strength.
  • pag_guidance_scale: Implements perturbed attention guidance specific to Sana.
  • num_inference_steps: Denoising steps (20-30 recommended for Flow-Euler).

Accessing Log Probabilities for Research

For uncertainty quantification or research applications requiring intermediate states, use the log-probability-aware variant in diffusion/post_training/diffusers_patch/pipeline_with_logprob.py:

from diffusion.post_training.diffusers_patch import pipeline_with_logprob

image, latents, log_probs = pipeline_with_logprob.pipeline_with_logprob_sana(
    transformer=pipe.transformer,
    vae=pipe.vae,
    prompt_embeds=pipe.text_encoder("A futuristic cityscape"),
    num_inference_steps=30,
    guidance_scale=5.0,
    solver="flow",
    deterministic=False,
)

This returns the final PIL image, a list of intermediate latent tensors, and per-step log-probabilities compatible with reinforcement learning or direct preference optimization workflows.

Key Integration Files

When working with the diffusers integration, these files define the critical touchpoints:

File Purpose
app/sana_pipeline.py Native pipeline implementation with build_vae, build_text_encoder, and the forward pass logic.
tools/convert_scripts/convert_sana_to_diffusers.py Checkpoint conversion utility that remaps weights to SanaTransformer2DModel.
diffusion/post_training/diffusers_patch/pipeline_with_logprob.py Research variant exposing pipeline_with_logprob_sana for log-probability extraction.
train_scripts/train_dreambooth_lora_sana.py Fine-tuning example demonstrating LoRA adaptation within the diffusers framework.
app/app_sana.py CLI wrapper demonstrating end-to-end pipeline loading and inference execution.

Summary

  • SanaPipeline implements the standard diffusers contract, enabling drop-in replacement with other text-to-image pipelines.
  • Use tools/convert_scripts/convert_sana_to_diffusers.py to convert native .pth checkpoints into loadable diffusers directories.
  • The pipeline supports multiple schedulers including DPMSolverMultistepScheduler, FlowMatchEulerDiscreteScheduler, and SCMScheduler.
  • For research requiring intermediate computations, import pipeline_with_logprob_sana from the diffusers patch module.
  • All standard diffusers parameters (guidance_scale, num_inference_steps, etc.) map directly to the underlying SanaPipeline.forward implementation in app/sana_pipeline.py.

Frequently Asked Questions

How do I convert a native Sana checkpoint to diffusers format?

Use the convert_sana_to_diffusers.py script located in tools/convert_scripts/. Provide the original checkpoint path, specify the model type (e.g., SanaMS_1600M_P1_D20), and set the output directory via --dump_path. The script handles parameter remapping, instantiates the SanaTransformer2DModel, and packages the VAE, text encoder, and scheduler into a standard diffusers pipeline structure ready for from_pretrained() loading.

What scheduler options are supported by SanaPipeline?

According to the source code in app/sana_pipeline.py, the pipeline supports three scheduler configurations: DPMSolverMultistepScheduler for standard DPM solving, FlowMatchEulerDiscreteScheduler for Flow-Euler sampling, and SCMScheduler specifically for Sana-Sprint models. The conversion script automatically selects the appropriate scheduler based on the --model_type argument, or you can manually override using the --scheduler_type flag.

How can I access intermediate latents and log probabilities during inference?

Import the pipeline_with_logprob_sana function from diffusion/post_training/diffusers_patch/pipeline_with_logprob.py. This function accepts the same core components (transformer, VAE, prompt embeddings) as the standard pipeline but returns a tuple containing the final image, a list of intermediate latent tensors, and per-step log-probabilities. This is particularly useful for research applications involving direct preference optimization or uncertainty quantification.

What is the difference between the native Sana pipeline and the diffusers implementation?

The native implementation in app/sana_pipeline.py uses pyrallis for configuration management and includes Sana-specific utilities like build_sana_model(). When converted via convert_sana_to_diffusers.py, the model becomes a standard diffusers pipeline that can be imported with from diffusers import SanaPipeline. The converted version uses the exact same weights and architecture but follows the diffusers API contract, enabling interoperability with the broader ecosystem of diffusers tools, schedulers, and training scripts like train_dreambooth_lora_sana.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →