Running Sana with 4-bit Quantization using SVDQuant and Nunchaku

You can run SANA inference with 4-bit quantization by replacing the standard SanaTransformer2DModel with NunchakuSanaTransformer2DModel from the Nunchaku library, reducing GPU memory usage to approximately 8 GB while maintaining generation quality through SVDQuant compression.

SANA is a Linear Diffusion Transformer (DiT) developed by NVIDIA Labs that generates high-resolution images efficiently. Running Sana with 4-bit quantization using SVDQuant and Nunchaku enables high-resolution generation on consumer GPUs with minimal latency overhead. This workflow replaces only the transformer backbone with an int4-quantized version while keeping the text encoder and VAE in BF16 precision.

Architecture of the 4-bit Pipeline

The SANA pipeline consists of three core components: a decoder-only text encoder for prompt embeddings, the SanaTransformer2DModel as the DiT backbone, and an AutoencoderDC VAE for latent decoding. During standard inference, the DiT runs in FP16 or BF16 precision.

4-bit quantization replaces the FP16 transformer with an int4-quantized version supplied by the Nunchaku library. According to the NVlabs/Sana source code, this quantized model is a drop-in replacement for the original SanaTransformer2DModel, meaning the text encoder, VAE, and scheduler remain unchanged.

SVDQuant compresses the model by factorizing weight matrices using singular-value decomposition and storing the factors in 4-bit format. The Nunchaku runtime efficiently dequantizes these factors on-the-fly during inference, executing the transformer with minimal overhead.

Converting Checkpoints to SVDQuant Format

Before running inference, you must convert full-precision checkpoints to the schema expected by Nunchaku. The repository provides tools/convert_scripts/convert_sana_to_svdquant.py to handle this transformation.

The conversion script rewrites weight names, interpolates positional embeddings, and reshapes linear-attention matrices. It patches patch-embeddings, caption projection layers, and time-embedding modules to match the Nunchaku implementation.

python -m tools.convert_scripts.convert_sana_to_svdquant \
    --orig_ckpt_path Efficient-Large-Model/Sana_1600M_1024px_BF16_diffusers \
    --model_type SanaMS_1600M_P1_D20 \
    --image_size 1024 \
    --output_path ./svdq_int4_checkpoint

This produces a checkpoint compatible with NunchakuSanaTransformer2DModel, preserving model quality while preparing weights for 4-bit storage.

Loading the Quantized Model for Inference

To execute 4-bit inference, instantiate the quantized transformer and pass it to the standard SanaPipeline. The implementation in app/app_sana_4bit.py demonstrates this pattern.

import torch
from diffusers import SanaPipeline
from nunchaku.models.transformer_sana import NunchakuSanaTransformer2DModel

# Load the int4-quantized transformer (SVDQuant checkpoint)

transformer = NunchakuSanaTransformer2DModel.from_pretrained(
    "mit-han-lab/svdq-int4-sana-1600m"
)

# Build the pipeline with the quantized transformer

pipe = SanaPipeline.from_pretrained(
    "Efficient-Large-Model/Sana_1600M_1024px_BF16_diffusers",
    transformer=transformer,
    variant="bf16",
    torch_dtype=torch.bfloat16,
).to("cuda")

# Optional: ensure text encoder and VAE remain in BF16

pipe.text_encoder.to(torch.bfloat16)
pipe.vae.to(torch.bfloat16)

prompt = "a cyberpunk cat with a neon sign that says \"Sana\""
image = pipe(
    prompt=prompt,
    height=1024,
    width=1024,
    guidance_scale=4.5,
    num_inference_steps=20,
    generator=torch.Generator(device="cuda").manual_seed(42),
)[0]

image.save("sana_4bit.png")

The only change from standard inference is the int4 transformer import. All other pipeline components load in BF16 mode as configured in diffusion/utils/config.py, and the scheduler from diffusion/scheduler/flow_euler_sampler.py operates unchanged.

Running the Gradio Demo

For interactive testing, use the specialized demo application that ships with the repository.

python -m app.app_sana_4bit --share

This launches a web interface identical to the full-precision demo but loads the int4 transformer automatically. The application handles prompt processing, seed management, and image generation through the quantized pipeline, providing a direct comparison of the user experience between precisions.

Memory and Performance Benefits

Running Sana with 4-bit quantization using SVDQuant and Nunchaku reduces the GPU memory footprint to approximately 8 GB, enabling 4K image generation on single consumer GPUs. The linear attention mechanism in SANA already reduces token-to-token quadratic complexity, making the model memory-efficient before quantization even begins.

The SVDQuant method preserves most original model quality by decomposing matrices rather than applying uniform quantization. Nunchaku's runtime minimizes dequantization overhead, keeping inference latency comparable to the FP16 baseline while dramatically reducing VRAM requirements.

Summary

  • SVDQuant enables 4-bit compression through singular-value decomposition of weight matrices.
  • Nunchaku provides the NunchakuSanaTransformer2DModel runtime for efficient int4 inference.
  • Convert existing checkpoints using tools/convert_scripts/convert_sana_to_svdquant.py before deployment.
  • The quantized transformer drops into the standard SanaPipeline without modifying text encoders, VAEs, or schedulers.
  • 4-bit inference requires approximately 8 GB of GPU memory for high-resolution generation.
  • Reference implementation available in app/app_sana_4bit.py.

Frequently Asked Questions

What hardware requirements are needed for 4-bit Sana inference?

You need a CUDA-capable GPU with approximately 8 GB of VRAM to run the int4-quantized model at resolutions up to 4K. The Nunchaku runtime optimizes memory access patterns specifically for NVIDIA hardware, making consumer-grade cards like the RTX 3070 or 4060 viable for production inference.

Does 4-bit quantization affect image quality compared to BF16?

SVDQuant preserves most of the original model quality by factorizing weight matrices rather than applying aggressive rounding. According to the NVlabs/Sana implementation, the perceptual difference between 4-bit and BF16 outputs is minimal because SVDQuant stores the singular value decomposition factors rather than naive quantized weights, maintaining the model's representational capacity.

Can I use custom trained checkpoints with the 4-bit pipeline?

Yes, but you must first convert your checkpoint using tools/convert_scripts/convert_sana_to_svdquant.py. The script handles weight name remapping, positional embedding interpolation, and linear-attention matrix reshaping required by the Nunchaku format. After conversion, load the resulting checkpoint with NunchakuSanaTransformer2DModel.from_pretrained() instead of the standard transformer.

How does the Nunchaku transformer differ from the standard SanaTransformer2DModel?

NunchakuSanaTransformer2DModel implements the same forward interface as the standard SanaTransformer2DModel but executes operations using 4-bit quantized weights internally. According to the source code in app/app_sana_4bit.py, it accepts identical inputs and configuration parameters from diffusion/utils/config.py, making it a true drop-in replacement that requires no changes to pipeline logic or scheduler code from diffusion/scheduler/flow_euler_sampler.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →