Configuring Inference-Time Scaling in SANA-1.5 for Improved Image Quality

To improve image quality in SANA-1.5, generate thousands of candidate images per prompt and use the NVILA-Verifier in a tournament-style selection process to pick the top-K highest quality outputs, significantly boosting GenEval scores without retraining.

The NVlabs/Sana repository implements a powerful inference-time scaling strategy that transforms the SANA-1.5 diffusion model from a single-shot generator into a quality-optimized pipeline. By leveraging massive candidate generation followed by intelligent verification, you can extract substantially better perceptual quality from existing checkpoints without modifying the underlying model weights.

How Inference-Time Scaling Works in SANA-1.5

SANA-1.5 treats image generation as a search problem over a large candidate space. The architecture follows a three-stage pipeline that decouples generation from selection, allowing the lightweight NVILA-Verifier to surface the best results from thousands of candidates.

Stage 1: Massive Candidate Generation

The pipeline begins by generating N candidate images for each input prompt using the standard SANA inference entry point. In scripts/inference_sana_sprint.py, the generation process respects the --img_nums_per_sample flag, which controls the number of candidates produced per prompt (commonly set to 2,048 for maximum quality gains).

The script writes outputs to a structured directory format (<output_dir>/<idx>/samples/*.png) alongside a metadata.jsonl file that preserves the original prompt text. This metadata is crucial for the subsequent verification stage, as the NVILA-Verifier requires the original text to perform accurate pairwise comparisons.

Stage 2: NVILA-Verifier Selection

Once candidates are generated, tools/inference_scaling/nvila_sana_pick.py loads the NVILA-Lite-2B-Verifier model using AutoModel.from_pretrained with device_map="auto" for automatic GPU allocation. The verifier implements a tournament-style pairwise comparison that repeatedly halves the candidate set until only the desired K images remain (defaulting to 4).

The core logic resides in the nvila_compare function, which presents image pairs to the verifier and receives "yes/no" decisions along with token scores (yes_id, no_id) used to break ties. This process reads prompts via the get_prompt helper function that parses the metadata.jsonl files generated in Stage 1. Selected images are copied to output/nvila_pick/best_K_of_N/ for final evaluation.

Stage 3: GenEval Quality Assessment

The final stage computes objective quality metrics using tools/metrics/compute_geneval.sh. This script aggregates the selected top-K images and runs the official GenEval metric to quantify the improvement. The documented scaling curve demonstrates that more candidates directly correlate with higher GenEval scores, while the verifier efficiently compresses the large pool to a manageable number of high-quality outputs.

Setting Up the Inference Pipeline

Before executing the scaling workflow, ensure your environment meets the specific dependencies documented in docs/inference_scaling/inference_scaling.md.

Environment Requirements

The inference-time scaling feature requires exact package versions to ensure reproducible results:

  • transformers==4.46 for the NVILA model compatibility
  • scaling_on_scales library for tournament logic utilities
  • Standard SANA dependencies (PyTorch, diffusers, etc.)

Install these within your existing SANA conda environment or create a dedicated environment following the instructions in the repository's inference scaling documentation.

Directory Structure and Output Paths

The pipeline expects a specific folder hierarchy:

  • Generation output: $output_dir/<sample_idx>/samples/*.png containing the raw candidates
  • Metadata: $output_dir/metadata.jsonl storing prompt strings indexed by sample ID
  • Selection output: output/nvila_pick/best_${K}_of_${N}/ containing the winning images
  • Final metrics: Reports generated via the GenEval evaluation scripts

Step-by-Step Implementation

Execute the full inference-time scaling pipeline using the provided shell wrappers and Python utilities.

Generating Candidate Images

First, download a SANA-1.5 checkpoint and generate your candidate pool. The --img_nums_per_sample parameter controls the inference-time scaling factor:


# Download the 600M checkpoint (example)

huggingface-cli download Efficient-Large-Model/Sana_600M_512px \
    --repo-type model --local-dir output/Sana_600M_512px --local-dir-use-symlinks False

# Configure scaling parameters

n_samples=32  # Use 2048 for production quality

output_dir=output/geneval_generated_path

# Generate candidates using the SANA inference script

bash scripts/infer_run_inference_geneval.sh \
    configs/sana_config/512ms/Sana_600M_img512.yaml \
    output/Sana_600M_512px/checkpoints/Sana_600M_512px_MultiLing.pth \
    --img_nums_per_sample=$n_samples \
    --output_dir=$output_dir

This invokes the standard inference entry point and populates the output directory with candidate images and metadata.

Running the Tournament Selection

Execute the NVILA verifier to select the top-K images from your generated candidates:

pick_number=4  # Final number of images to keep per prompt

bash tools/inference_scaling/nvila_sana_pick.sh \
    $output_dir \
    $n_samples \
    $pick_number

The shell wrapper calls python -m tools.inference_scaling.nvila_sana_pick with the specified arguments. The verifier performs pairwise comparisons tournament-style, reducing the N candidates to the best K images based on visual alignment with the original prompt.

Computing Final Quality Metrics

Evaluate the selected images using the GenEval metric to quantify the quality improvement:


# Activate the Geneval environment (see tools/metrics/geneval/geneval_env.md)

conda activate geneval

DIR_AFTER_PICK="output/nvila_pick/best_${pick_number}_of_${n_samples}/${output_dir}"
bash tools/metrics/compute_geneval.sh $(dirname "$DIR_AFTER_PICK") $(basename "$DIR_AFTER_PICK")

This produces the final GenEval score demonstrating the effectiveness of your inference-time scaling configuration.

Technical Implementation Details

Understanding the verifier architecture helps optimize the pipeline for your specific hardware constraints.

The NVILA Verifier Component

The tools/inference_scaling/nvila_sana_pick.py utility is deliberately decoupled from the main SANA codebase, allowing you to swap in alternative verifiers if needed. It initializes the NVILA-Lite-2B-Verifier once and reuses the model instance across all pairwise comparisons in the tournament, minimizing GPU memory overhead.

The verifier runs on the same device as the generation step (via device_map="auto"), enabling the entire pipeline to execute on a single GPU node without additional orchestration infrastructure.

Scaling Trade-offs and Performance

The relationship between candidate count and quality follows a scaling curve documented in docs/inference_scaling/inference_scaling.md. While generating 2,048 candidates yields the highest GenEval scores, the tournament selection efficiently reduces this to 4 images, keeping the total inference latency manageable compared to brute-force enumeration.

Because the verifier performs log(N) comparisons rather than evaluating all candidates individually, the selection overhead grows logarithmically with candidate count, making large-N scaling practical for production deployments.

Summary

  • Inference-time scaling in SANA-1.5 generates multiple candidate images per prompt and selects the best using the NVILA-Verifier.
  • The pipeline uses three stages: candidate generation via scripts/inference_sana_sprint.py, tournament selection via tools/inference_scaling/nvila_sana_pick.py, and quality scoring via tools/metrics/compute_geneval.sh.
  • The NVILA-Lite-2B-Verifier performs pairwise comparisons in a tournament structure, efficiently reducing thousands of candidates to a handful of high-quality outputs.
  • Configuration requires specific dependencies including transformers==4.46 and the scaling_on_scales library.
  • The entire workflow runs on a single GPU node, with the verifier initialized using device_map="auto" for automatic resource allocation.
  • Increasing the candidate count (controlled by --img_nums_per_sample) directly improves GenEval scores without model retraining.

Frequently Asked Questions

How many candidate images should I generate for optimal results?

According to the SANA-1.5 source code documentation, generating 2,048 candidates per prompt delivers the highest GenEval scores, though you can use as few as 32 for quick testing. The NVILA-Verifier efficiently handles large candidate pools through its logarithmic tournament selection algorithm, so the primary constraint becomes generation time rather than verification overhead.

Can I use a different verifier model instead of NVILA-Lite-2B?

Yes. The tools/inference_scaling/nvila_sana_pick.py script is intentionally decoupled from the main SANA codebase, allowing you to substitute any verification model that implements a pairwise comparison interface. Replace the AutoModel.from_pretrained call and the nvila_compare function logic with your preferred verifier's inference code while maintaining the same input/output contract.

What specific dependencies are required for the inference scaling pipeline?

You must install transformers==4.46 and the scaling_on_scales library alongside your standard SANA environment. These requirements are documented in docs/inference_scaling/inference_scaling.md. The specific version constraint on transformers ensures compatibility with the NVILA model architecture and token scoring mechanisms used in the tournament selection process.

Does inference-time scaling work with all SANA-1.5 model sizes?

Yes. The scaling strategy operates at the inference script level (scripts/infer_run_inference_geneval.sh) and works with any SANA checkpoint compatible with the standard configuration files in configs/sana_config/. Whether using the 600M or larger variants, the --img_nums_per_sample flag and NVILA-Verifier selection process remain identical.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →