# Running Sana with 4-bit Quantization using SVDQuant and Nunchaku

> Run Sana with 4-bit quantization using SVDQuant and Nunchaku. Slash GPU memory to 8GB while preserving quality. Easily integrate NunchakuSanaTransformer2DModel for efficient inference.

- Repository: [NVIDIA Research Projects/Sana](https://github.com/NVlabs/Sana)
- Tags: how-to-guide
- Published: 2026-05-19

---

**You can run SANA inference with 4-bit quantization by replacing the standard `SanaTransformer2DModel` with `NunchakuSanaTransformer2DModel` from the Nunchaku library, reducing GPU memory usage to approximately 8 GB while maintaining generation quality through SVDQuant compression.**

SANA is a Linear Diffusion Transformer (DiT) developed by NVIDIA Labs that generates high-resolution images efficiently. Running Sana with 4-bit quantization using SVDQuant and Nunchaku enables high-resolution generation on consumer GPUs with minimal latency overhead. This workflow replaces only the transformer backbone with an int4-quantized version while keeping the text encoder and VAE in BF16 precision.

## Architecture of the 4-bit Pipeline

The SANA pipeline consists of three core components: a decoder-only text encoder for prompt embeddings, the `SanaTransformer2DModel` as the DiT backbone, and an `AutoencoderDC` VAE for latent decoding. During standard inference, the DiT runs in FP16 or BF16 precision.

4-bit quantization replaces the FP16 transformer with an **int4-quantized version** supplied by the Nunchaku library. According to the NVlabs/Sana source code, this quantized model is a drop-in replacement for the original `SanaTransformer2DModel`, meaning the text encoder, VAE, and scheduler remain unchanged.

**SVDQuant** compresses the model by factorizing weight matrices using singular-value decomposition and storing the factors in 4-bit format. The **Nunchaku** runtime efficiently dequantizes these factors on-the-fly during inference, executing the transformer with minimal overhead.

## Converting Checkpoints to SVDQuant Format

Before running inference, you must convert full-precision checkpoints to the schema expected by Nunchaku. The repository provides [`tools/convert_scripts/convert_sana_to_svdquant.py`](https://github.com/NVlabs/Sana/blob/main/tools/convert_scripts/convert_sana_to_svdquant.py) to handle this transformation.

The conversion script rewrites weight names, interpolates positional embeddings, and reshapes linear-attention matrices. It patches patch-embeddings, caption projection layers, and time-embedding modules to match the Nunchaku implementation.

```bash
python -m tools.convert_scripts.convert_sana_to_svdquant \
    --orig_ckpt_path Efficient-Large-Model/Sana_1600M_1024px_BF16_diffusers \
    --model_type SanaMS_1600M_P1_D20 \
    --image_size 1024 \
    --output_path ./svdq_int4_checkpoint

```

This produces a checkpoint compatible with `NunchakuSanaTransformer2DModel`, preserving model quality while preparing weights for 4-bit storage.

## Loading the Quantized Model for Inference

To execute 4-bit inference, instantiate the quantized transformer and pass it to the standard `SanaPipeline`. The implementation in [`app/app_sana_4bit.py`](https://github.com/NVlabs/Sana/blob/main/app/app_sana_4bit.py) demonstrates this pattern.

```python
import torch
from diffusers import SanaPipeline
from nunchaku.models.transformer_sana import NunchakuSanaTransformer2DModel

# Load the int4-quantized transformer (SVDQuant checkpoint)

transformer = NunchakuSanaTransformer2DModel.from_pretrained(
    "mit-han-lab/svdq-int4-sana-1600m"
)

# Build the pipeline with the quantized transformer

pipe = SanaPipeline.from_pretrained(
    "Efficient-Large-Model/Sana_1600M_1024px_BF16_diffusers",
    transformer=transformer,
    variant="bf16",
    torch_dtype=torch.bfloat16,
).to("cuda")

# Optional: ensure text encoder and VAE remain in BF16

pipe.text_encoder.to(torch.bfloat16)
pipe.vae.to(torch.bfloat16)

prompt = "a cyberpunk cat with a neon sign that says \"Sana\""
image = pipe(
    prompt=prompt,
    height=1024,
    width=1024,
    guidance_scale=4.5,
    num_inference_steps=20,
    generator=torch.Generator(device="cuda").manual_seed(42),
)[0]

image.save("sana_4bit.png")

```

The only change from standard inference is the **int4 transformer** import. All other pipeline components load in BF16 mode as configured in [`diffusion/utils/config.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/utils/config.py), and the scheduler from [`diffusion/scheduler/flow_euler_sampler.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/scheduler/flow_euler_sampler.py) operates unchanged.

## Running the Gradio Demo

For interactive testing, use the specialized demo application that ships with the repository.

```bash
python -m app.app_sana_4bit --share

```

This launches a web interface identical to the full-precision demo but loads the int4 transformer automatically. The application handles prompt processing, seed management, and image generation through the quantized pipeline, providing a direct comparison of the user experience between precisions.

## Memory and Performance Benefits

Running Sana with 4-bit quantization using SVDQuant and Nunchaku reduces the **GPU memory footprint to approximately 8 GB**, enabling 4K image generation on single consumer GPUs. The linear attention mechanism in SANA already reduces token-to-token quadratic complexity, making the model memory-efficient before quantization even begins.

The SVDQuant method preserves most original model quality by decomposing matrices rather than applying uniform quantization. Nunchaku's runtime minimizes dequantization overhead, keeping inference latency comparable to the FP16 baseline while dramatically reducing VRAM requirements.

## Summary

- **SVDQuant** enables 4-bit compression through singular-value decomposition of weight matrices.
- **Nunchaku** provides the `NunchakuSanaTransformer2DModel` runtime for efficient int4 inference.
- Convert existing checkpoints using [`tools/convert_scripts/convert_sana_to_svdquant.py`](https://github.com/NVlabs/Sana/blob/main/tools/convert_scripts/convert_sana_to_svdquant.py) before deployment.
- The quantized transformer drops into the standard `SanaPipeline` without modifying text encoders, VAEs, or schedulers.
- 4-bit inference requires approximately **8 GB of GPU memory** for high-resolution generation.
- Reference implementation available in [`app/app_sana_4bit.py`](https://github.com/NVlabs/Sana/blob/main/app/app_sana_4bit.py).

## Frequently Asked Questions

### What hardware requirements are needed for 4-bit Sana inference?

You need a CUDA-capable GPU with approximately **8 GB of VRAM** to run the int4-quantized model at resolutions up to 4K. The Nunchaku runtime optimizes memory access patterns specifically for NVIDIA hardware, making consumer-grade cards like the RTX 3070 or 4060 viable for production inference.

### Does 4-bit quantization affect image quality compared to BF16?

SVDQuant preserves most of the original model quality by factorizing weight matrices rather than applying aggressive rounding. According to the NVlabs/Sana implementation, the perceptual difference between 4-bit and BF16 outputs is minimal because SVDQuant stores the singular value decomposition factors rather than naive quantized weights, maintaining the model's representational capacity.

### Can I use custom trained checkpoints with the 4-bit pipeline?

Yes, but you must first convert your checkpoint using [`tools/convert_scripts/convert_sana_to_svdquant.py`](https://github.com/NVlabs/Sana/blob/main/tools/convert_scripts/convert_sana_to_svdquant.py). The script handles weight name remapping, positional embedding interpolation, and linear-attention matrix reshaping required by the Nunchaku format. After conversion, load the resulting checkpoint with `NunchakuSanaTransformer2DModel.from_pretrained()` instead of the standard transformer.

### How does the Nunchaku transformer differ from the standard SanaTransformer2DModel?

`NunchakuSanaTransformer2DModel` implements the same forward interface as the standard `SanaTransformer2DModel` but executes operations using 4-bit quantized weights internally. According to the source code in [`app/app_sana_4bit.py`](https://github.com/NVlabs/Sana/blob/main/app/app_sana_4bit.py), it accepts identical inputs and configuration parameters from [`diffusion/utils/config.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/utils/config.py), making it a true drop-in replacement that requires no changes to pipeline logic or scheduler code from [`diffusion/scheduler/flow_euler_sampler.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/scheduler/flow_euler_sampler.py).