# Configuring Inference-Time Scaling in SANA-1.5 for Improved Image Quality

> Enhance SANA-1.5 image quality using inference-time scaling. Generate numerous candidates and select top-K via tournament for superior GenEval scores without retraining.

- Repository: [NVIDIA Research Projects/Sana](https://github.com/NVlabs/Sana)
- Tags: performance
- Published: 2026-05-19

---

**To improve image quality in SANA-1.5, generate thousands of candidate images per prompt and use the NVILA-Verifier in a tournament-style selection process to pick the top-K highest quality outputs, significantly boosting GenEval scores without retraining.**

The **NVlabs/Sana** repository implements a powerful **inference-time scaling** strategy that transforms the SANA-1.5 diffusion model from a single-shot generator into a quality-optimized pipeline. By leveraging massive candidate generation followed by intelligent verification, you can extract substantially better perceptual quality from existing checkpoints without modifying the underlying model weights.

## How Inference-Time Scaling Works in SANA-1.5

SANA-1.5 treats image generation as a search problem over a large candidate space. The architecture follows a three-stage pipeline that decouples generation from selection, allowing the lightweight NVILA-Verifier to surface the best results from thousands of candidates.

### Stage 1: Massive Candidate Generation

The pipeline begins by generating **N candidate images** for each input prompt using the standard SANA inference entry point. In [`scripts/inference_sana_sprint.py`](https://github.com/NVlabs/Sana/blob/main/scripts/inference_sana_sprint.py), the generation process respects the `--img_nums_per_sample` flag, which controls the number of candidates produced per prompt (commonly set to 2,048 for maximum quality gains).

The script writes outputs to a structured directory format (`<output_dir>/<idx>/samples/*.png`) alongside a `metadata.jsonl` file that preserves the original prompt text. This metadata is crucial for the subsequent verification stage, as the NVILA-Verifier requires the original text to perform accurate pairwise comparisons.

### Stage 2: NVILA-Verifier Selection

Once candidates are generated, [`tools/inference_scaling/nvila_sana_pick.py`](https://github.com/NVlabs/Sana/blob/main/tools/inference_scaling/nvila_sana_pick.py) loads the **NVILA-Lite-2B-Verifier** model using `AutoModel.from_pretrained` with `device_map="auto"` for automatic GPU allocation. The verifier implements a **tournament-style pairwise comparison** that repeatedly halves the candidate set until only the desired **K images** remain (defaulting to 4).

The core logic resides in the `nvila_compare` function, which presents image pairs to the verifier and receives "yes/no" decisions along with token scores (`yes_id`, `no_id`) used to break ties. This process reads prompts via the `get_prompt` helper function that parses the `metadata.jsonl` files generated in Stage 1. Selected images are copied to `output/nvila_pick/best_K_of_N/` for final evaluation.

### Stage 3: GenEval Quality Assessment

The final stage computes objective quality metrics using [`tools/metrics/compute_geneval.sh`](https://github.com/NVlabs/Sana/blob/main/tools/metrics/compute_geneval.sh). This script aggregates the selected top-K images and runs the official GenEval metric to quantify the improvement. The documented scaling curve demonstrates that **more candidates directly correlate with higher GenEval scores**, while the verifier efficiently compresses the large pool to a manageable number of high-quality outputs.

## Setting Up the Inference Pipeline

Before executing the scaling workflow, ensure your environment meets the specific dependencies documented in [`docs/inference_scaling/inference_scaling.md`](https://github.com/NVlabs/Sana/blob/main/docs/inference_scaling/inference_scaling.md).

### Environment Requirements

The inference-time scaling feature requires exact package versions to ensure reproducible results:

- **transformers==4.46** for the NVILA model compatibility
- **scaling_on_scales** library for tournament logic utilities
- Standard SANA dependencies (PyTorch, diffusers, etc.)

Install these within your existing SANA conda environment or create a dedicated environment following the instructions in the repository's inference scaling documentation.

### Directory Structure and Output Paths

The pipeline expects a specific folder hierarchy:

- **Generation output**: `$output_dir/<sample_idx>/samples/*.png` containing the raw candidates
- **Metadata**: `$output_dir/metadata.jsonl` storing prompt strings indexed by sample ID
- **Selection output**: `output/nvila_pick/best_${K}_of_${N}/` containing the winning images
- **Final metrics**: Reports generated via the GenEval evaluation scripts

## Step-by-Step Implementation

Execute the full inference-time scaling pipeline using the provided shell wrappers and Python utilities.

### Generating Candidate Images

First, download a SANA-1.5 checkpoint and generate your candidate pool. The `--img_nums_per_sample` parameter controls the inference-time scaling factor:

```bash

# Download the 600M checkpoint (example)

huggingface-cli download Efficient-Large-Model/Sana_600M_512px \
    --repo-type model --local-dir output/Sana_600M_512px --local-dir-use-symlinks False

# Configure scaling parameters

n_samples=32  # Use 2048 for production quality

output_dir=output/geneval_generated_path

# Generate candidates using the SANA inference script

bash scripts/infer_run_inference_geneval.sh \
    configs/sana_config/512ms/Sana_600M_img512.yaml \
    output/Sana_600M_512px/checkpoints/Sana_600M_512px_MultiLing.pth \
    --img_nums_per_sample=$n_samples \
    --output_dir=$output_dir

```

This invokes the standard inference entry point and populates the output directory with candidate images and metadata.

### Running the Tournament Selection

Execute the NVILA verifier to select the top-K images from your generated candidates:

```bash
pick_number=4  # Final number of images to keep per prompt

bash tools/inference_scaling/nvila_sana_pick.sh \
    $output_dir \
    $n_samples \
    $pick_number

```

The shell wrapper calls `python -m tools.inference_scaling.nvila_sana_pick` with the specified arguments. The verifier performs pairwise comparisons tournament-style, reducing the N candidates to the best K images based on visual alignment with the original prompt.

### Computing Final Quality Metrics

Evaluate the selected images using the GenEval metric to quantify the quality improvement:

```bash

# Activate the Geneval environment (see tools/metrics/geneval/geneval_env.md)

conda activate geneval

DIR_AFTER_PICK="output/nvila_pick/best_${pick_number}_of_${n_samples}/${output_dir}"
bash tools/metrics/compute_geneval.sh $(dirname "$DIR_AFTER_PICK") $(basename "$DIR_AFTER_PICK")

```

This produces the final GenEval score demonstrating the effectiveness of your inference-time scaling configuration.

## Technical Implementation Details

Understanding the verifier architecture helps optimize the pipeline for your specific hardware constraints.

### The NVILA Verifier Component

The [`tools/inference_scaling/nvila_sana_pick.py`](https://github.com/NVlabs/Sana/blob/main/tools/inference_scaling/nvila_sana_pick.py) utility is deliberately decoupled from the main SANA codebase, allowing you to swap in alternative verifiers if needed. It initializes the NVILA-Lite-2B-Verifier once and reuses the model instance across all pairwise comparisons in the tournament, minimizing GPU memory overhead.

The verifier runs on the same device as the generation step (via `device_map="auto"`), enabling the entire pipeline to execute on a single GPU node without additional orchestration infrastructure.

### Scaling Trade-offs and Performance

The relationship between candidate count and quality follows a **scaling curve** documented in [`docs/inference_scaling/inference_scaling.md`](https://github.com/NVlabs/Sana/blob/main/docs/inference_scaling/inference_scaling.md). While generating 2,048 candidates yields the highest GenEval scores, the tournament selection efficiently reduces this to 4 images, keeping the total inference latency manageable compared to brute-force enumeration.

Because the verifier performs log(N) comparisons rather than evaluating all candidates individually, the selection overhead grows logarithmically with candidate count, making large-N scaling practical for production deployments.

## Summary

- **Inference-time scaling** in SANA-1.5 generates multiple candidate images per prompt and selects the best using the NVILA-Verifier.
- The pipeline uses three stages: candidate generation via [`scripts/inference_sana_sprint.py`](https://github.com/NVlabs/Sana/blob/main/scripts/inference_sana_sprint.py), tournament selection via [`tools/inference_scaling/nvila_sana_pick.py`](https://github.com/NVlabs/Sana/blob/main/tools/inference_scaling/nvila_sana_pick.py), and quality scoring via [`tools/metrics/compute_geneval.sh`](https://github.com/NVlabs/Sana/blob/main/tools/metrics/compute_geneval.sh).
- The NVILA-Lite-2B-Verifier performs pairwise comparisons in a tournament structure, efficiently reducing thousands of candidates to a handful of high-quality outputs.
- Configuration requires specific dependencies including `transformers==4.46` and the `scaling_on_scales` library.
- The entire workflow runs on a single GPU node, with the verifier initialized using `device_map="auto"` for automatic resource allocation.
- Increasing the candidate count (controlled by `--img_nums_per_sample`) directly improves GenEval scores without model retraining.

## Frequently Asked Questions

### How many candidate images should I generate for optimal results?

According to the SANA-1.5 source code documentation, generating **2,048 candidates** per prompt delivers the highest GenEval scores, though you can use as few as 32 for quick testing. The NVILA-Verifier efficiently handles large candidate pools through its logarithmic tournament selection algorithm, so the primary constraint becomes generation time rather than verification overhead.

### Can I use a different verifier model instead of NVILA-Lite-2B?

Yes. The [`tools/inference_scaling/nvila_sana_pick.py`](https://github.com/NVlabs/Sana/blob/main/tools/inference_scaling/nvila_sana_pick.py) script is intentionally decoupled from the main SANA codebase, allowing you to substitute any verification model that implements a pairwise comparison interface. Replace the `AutoModel.from_pretrained` call and the `nvila_compare` function logic with your preferred verifier's inference code while maintaining the same input/output contract.

### What specific dependencies are required for the inference scaling pipeline?

You must install **transformers==4.46** and the **scaling_on_scales** library alongside your standard SANA environment. These requirements are documented in [`docs/inference_scaling/inference_scaling.md`](https://github.com/NVlabs/Sana/blob/main/docs/inference_scaling/inference_scaling.md). The specific version constraint on transformers ensures compatibility with the NVILA model architecture and token scoring mechanisms used in the tournament selection process.

### Does inference-time scaling work with all SANA-1.5 model sizes?

Yes. The scaling strategy operates at the inference script level ([`scripts/infer_run_inference_geneval.sh`](https://github.com/NVlabs/Sana/blob/main/scripts/infer_run_inference_geneval.sh)) and works with any SANA checkpoint compatible with the standard configuration files in `configs/sana_config/`. Whether using the 600M or larger variants, the `--img_nums_per_sample` flag and NVILA-Verifier selection process remain identical.