# How to Use the Diffusers Integration with SanaPipeline

> Learn how to use the diffusers integration with SanaPipeline for seamless loading and inference of native checkpoints. Effortlessly use the familiar pipe API.

- Repository: [NVIDIA Research Projects/Sana](https://github.com/NVlabs/Sana)
- Tags: how-to-guide
- Published: 2026-05-19

---

**The SanaPipeline from NVlabs/Sana is a native diffusers-compatible pipeline that enables standard `.from_pretrained()` loading and inference via the familiar `pipe()` API, requiring only a conversion step for native checkpoints.**

The NVlabs/Sana repository provides a high-performance text-to-image generation model designed to work seamlessly within the 🤗 diffusers ecosystem. By leveraging the diffusers integration with SanaPipeline, you can instantiate models using standard diffusers patterns, convert native PyTorch checkpoints, and customize inference with scheduler substitutions. This integration follows the standard diffusers architecture of VAE, text encoder, transformer, and scheduler components.

## Architecture of the SanaPipeline Diffusers Integration

The Sana pipeline implements the standard diffusers contract, making it interchangeable with other text-to-image pipelines. Understanding this architecture is essential for customizing components or debugging inference.

### Core Component Mapping

The pipeline composes four primary components that map directly to diffusers classes:

- **VAE**: Implements latent encoding/decoding using `AutoencoderDC` from diffusers, responsible for converting between pixel space and latent space via the `vae_decode` method.
- **Text Encoder**: Wraps `AutoModelForCausalLM` from HuggingFace, tokenized via `self.tokenizer` to produce text embeddings.
- **Transformer**: The core Sana model (`SanaTransformer2DModel` from diffusers) generates latent updates conditioned on text embeddings and timestep information.
- **Scheduler**: Supports `DPMSolverMultistepScheduler`, `FlowMatchEulerDiscreteScheduler`, or `SCMScheduler` depending on the sampling strategy (Flow-Euler, DPMSolver, or SCM for Sana-Sprint variants).

### Initialization and Forward Pass Implementation

In [`app/sana_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_pipeline.py), the `SanaPipeline.__init__` method orchestrates component initialization using `pyrallis` configuration management:

1. Loads hyperparameters into a `SanaInference` dataclass.
2. Constructs components via `build_vae()`, `build_text_encoder()`, and `build_sana_model()`.
3. Pre-computes null caption embeddings for classifier-free guidance.

The `SanaPipeline.forward` method (lines 66‑78 in [`app/sana_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_pipeline.py)) implements the standard diffusion loop:

- Tokenizes prompts and generates text embeddings via `self.text_encoder`.
- Initializes latents either from noise or provided inputs.
- Instantiates the solver (either `FlowEuler` or `DPMSolverMultistepScheduler`) based on `self.vis_sampler`.
- Runs the denoising scheduler for `num_inference_steps` iterations.
- Decodes final latents through the VAE, optionally resizing via `resize_and_crop_tensor`.

## Converting Native Checkpoints to Diffusers Format

Native Sana checkpoints (`.pth` files) require conversion to the diffusers directory structure before using the standard API.

### Using the Conversion Script

The repository provides [`tools/convert_scripts/convert_sana_to_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/tools/convert_scripts/convert_sana_to_diffusers.py) to handle parameter remapping and pipeline packaging:

```bash
python tools/convert_scripts/convert_sana_to_diffusers.py \
    --orig_ckpt_path=/path/to/Sana_1600M_1024px.pth \
    --image_size=1024 \
    --model_type=SanaMS_1600M_P1_D20 \
    --scheduler_type=flow-dpm_solver \
    --dump_path=./my_sana_pipeline \
    --save_full_pipeline \
    --dtype=fp16

```

The script performs the following operations:
- Loads the native checkpoint using `torch.load()` and remaps parameter names to the diffusers naming convention.
- Instantiates `SanaTransformer2DModel` with appropriate `model_kwargs`, optionally using `accelerate` empty-weight initialization for memory efficiency.
- Constructs the VAE via `AutoencoderDC.from_pretrained()` and text encoder via `AutoModelForCausalLM`.
- Selects the appropriate scheduler based on `model_type` (SCM for Sana-Sprint, otherwise DPMSolver or Flow-Euler).
- Packages components into a `SanaPipeline` (or `SanaSprintPipeline`) and exports via `save_pretrained()`.

After execution, `./my_sana_pipeline` contains a valid diffusers pipeline with [`config.json`](https://github.com/NVlabs/Sana/blob/main/config.json), `transformer/`, `vae/`, and `scheduler/` subdirectories.

## Loading and Running Inference

Once converted, the pipeline follows standard diffusers usage patterns.

### Loading a Pre-converted Pipeline

Import the pipeline class and load from the converted directory:

```python
from diffusers import SanaPipeline
import torch

pipe = SanaPipeline.from_pretrained(
    "./my_sana_pipeline",  # or HuggingFace hub path

    torch_dtype=torch.float16,
)
pipe.to("cuda")

```

### Standard Text-to-Image Generation

The `__call__` method accepts standard diffusers arguments that map 1-to-1 to the `SanaPipeline.forward` signature:

```python
image = pipe(
    prompt="A futuristic cityscape at sunset with flying vehicles",
    height=1024,
    width=1024,
    num_inference_steps=20,
    guidance_scale=4.5,
    pag_guidance_scale=1.0,
).images[0]

image.save("output.png")

```

**Key parameters** include:
- `guidance_scale`: Controls classifier-free guidance strength.
- `pag_guidance_scale`: Implements perturbed attention guidance specific to Sana.
- `num_inference_steps`: Denoising steps (20-30 recommended for Flow-Euler).

### Accessing Log Probabilities for Research

For uncertainty quantification or research applications requiring intermediate states, use the log-probability-aware variant in [`diffusion/post_training/diffusers_patch/pipeline_with_logprob.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/post_training/diffusers_patch/pipeline_with_logprob.py):

```python
from diffusion.post_training.diffusers_patch import pipeline_with_logprob

image, latents, log_probs = pipeline_with_logprob.pipeline_with_logprob_sana(
    transformer=pipe.transformer,
    vae=pipe.vae,
    prompt_embeds=pipe.text_encoder("A futuristic cityscape"),
    num_inference_steps=30,
    guidance_scale=5.0,
    solver="flow",
    deterministic=False,
)

```

This returns the final PIL image, a list of intermediate latent tensors, and per-step log-probabilities compatible with reinforcement learning or direct preference optimization workflows.

## Key Integration Files

When working with the diffusers integration, these files define the critical touchpoints:

| File | Purpose |
|------|---------|
| [`app/sana_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_pipeline.py) | Native pipeline implementation with `build_vae`, `build_text_encoder`, and the forward pass logic. |
| [`tools/convert_scripts/convert_sana_to_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/tools/convert_scripts/convert_sana_to_diffusers.py) | Checkpoint conversion utility that remaps weights to `SanaTransformer2DModel`. |
| [`diffusion/post_training/diffusers_patch/pipeline_with_logprob.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/post_training/diffusers_patch/pipeline_with_logprob.py) | Research variant exposing `pipeline_with_logprob_sana` for log-probability extraction. |
| [`train_scripts/train_dreambooth_lora_sana.py`](https://github.com/NVlabs/Sana/blob/main/train_scripts/train_dreambooth_lora_sana.py) | Fine-tuning example demonstrating LoRA adaptation within the diffusers framework. |
| [`app/app_sana.py`](https://github.com/NVlabs/Sana/blob/main/app/app_sana.py) | CLI wrapper demonstrating end-to-end pipeline loading and inference execution. |

## Summary

- **SanaPipeline** implements the standard diffusers contract, enabling drop-in replacement with other text-to-image pipelines.
- Use **[`tools/convert_scripts/convert_sana_to_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/tools/convert_scripts/convert_sana_to_diffusers.py)** to convert native `.pth` checkpoints into loadable diffusers directories.
- The pipeline supports multiple schedulers including **DPMSolverMultistepScheduler**, **FlowMatchEulerDiscreteScheduler**, and **SCMScheduler**.
- For research requiring intermediate computations, import **`pipeline_with_logprob_sana`** from the diffusers patch module.
- All standard diffusers parameters (`guidance_scale`, `num_inference_steps`, etc.) map directly to the underlying `SanaPipeline.forward implementation` in [`app/sana_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_pipeline.py).

## Frequently Asked Questions

### How do I convert a native Sana checkpoint to diffusers format?

Use the [`convert_sana_to_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/convert_sana_to_diffusers.py) script located in `tools/convert_scripts/`. Provide the original checkpoint path, specify the model type (e.g., `SanaMS_1600M_P1_D20`), and set the output directory via `--dump_path`. The script handles parameter remapping, instantiates the `SanaTransformer2DModel`, and packages the VAE, text encoder, and scheduler into a standard diffusers pipeline structure ready for `from_pretrained()` loading.

### What scheduler options are supported by SanaPipeline?

According to the source code in [`app/sana_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_pipeline.py), the pipeline supports three scheduler configurations: **DPMSolverMultistepScheduler** for standard DPM solving, **FlowMatchEulerDiscreteScheduler** for Flow-Euler sampling, and **SCMScheduler** specifically for Sana-Sprint models. The conversion script automatically selects the appropriate scheduler based on the `--model_type` argument, or you can manually override using the `--scheduler_type` flag.

### How can I access intermediate latents and log probabilities during inference?

Import the `pipeline_with_logprob_sana` function from [`diffusion/post_training/diffusers_patch/pipeline_with_logprob.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/post_training/diffusers_patch/pipeline_with_logprob.py). This function accepts the same core components (transformer, VAE, prompt embeddings) as the standard pipeline but returns a tuple containing the final image, a list of intermediate latent tensors, and per-step log-probabilities. This is particularly useful for research applications involving direct preference optimization or uncertainty quantification.

### What is the difference between the native Sana pipeline and the diffusers implementation?

The native implementation in [`app/sana_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_pipeline.py) uses `pyrallis` for configuration management and includes Sana-specific utilities like `build_sana_model()`. When converted via [`convert_sana_to_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/convert_sana_to_diffusers.py), the model becomes a standard diffusers pipeline that can be imported with `from diffusers import SanaPipeline`. The converted version uses the exact same weights and architecture but follows the diffusers API contract, enabling interoperability with the broader ecosystem of diffusers tools, schedulers, and training scripts like [`train_dreambooth_lora_sana.py`](https://github.com/NVlabs/Sana/blob/main/train_dreambooth_lora_sana.py).