# Implementing Video-to-Video Editing with Sana Video Refiner: A Complete Guide

> Master video-to-video editing with Sana Video Refiner. Learn how to achieve high-fidelity results using diffusion and upsampling with this comprehensive guide.

- Repository: [NVIDIA Research Projects/Sana](https://github.com/NVlabs/Sana)
- Tags: how-to-guide
- Published: 2026-05-19

---

**The Sana Video Refiner enables high-fidelity video-to-video editing by combining the Sana video diffusion model with Lightricks LTX-2 upsampling through a six-stage latent refinement pipeline.**

The **Sana Video Refiner** in the NVlabs/Sana repository provides a modular approach to video editing that leverages latent diffusion techniques to transform input prompts into polished, high-resolution video outputs. According to the source code in [`app/sana_video_refiner_pipeline_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_video_refiner_pipeline_diffusers.py), the pipeline orchestrates two distinct models: Sana Video for initial latent generation and LTX-2 for Stage-2 refinement with specialized LoRA weights.

## What Is the Sana Video Refiner?

The refiner is a hybrid pipeline that bridges **Sana Video** (a 720p-capable video diffusion model) with **LTX-2** (a high-fidelity upsampling and refinement system). Unlike standard video generation scripts, this implementation keeps everything in the latent space until the final encoding step, minimizing GPU memory usage while maximizing output quality.

The architecture is intentionally modular, allowing you to swap LTX-2 checkpoints, inject custom LoRA weights for style adaptation, and bypass default Diffusers normalization to match the original LTX-2 implementation.

## Prerequisites and Setup

Before running the refiner, ensure your environment meets these requirements:

- **Python 3.8+** with PyTorch 2.0 or higher
- **Diffusers 0.32+** for pipeline compatibility
- CUDA-capable GPU with sufficient VRAM for 720p video generation

Install dependencies from the repository root:

```bash
pip install -r requirements.txt

```

The refiner depends on specific model checkpoints: `Efficient-Large-Model/SANA-Video_2B_720p` for the base generation and `Lightricks/LTX-2` (or compatible checkpoints) for Stage-2 refinement.

## The Six-Stage Refinement Workflow

The implementation in [`sana_video_refiner_pipeline_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/sana_video_refiner_pipeline_diffusers.py) follows a precise six-step workflow, with key operations occurring at specific line ranges.

### Stage 1: Latent Generation with SanaVideoPipeline

The process begins by generating a latent representation rather than a decoded video. In lines 53-71, the `SanaVideoPipeline` creates a compressed spatio-temporal tensor that respects the prompt's motion characteristics.

```python
sana_pipe = SanaVideoPipeline.from_pretrained(
    sana_model_id, torch_dtype=dtype
)
sana_pipe.text_encoder.to(dtype)
sana_pipe.enable_model_cpu_offload(gpu_id=0)

video_latent = sana_pipe(
    prompt=full_prompt,
    negative_prompt=negative_prompt,
    height=sana_height,
    width=sana_width,
    frames=sana_frames,
    guidance_scale=sana_guidance_scale,
    num_inference_steps=sana_num_steps,
    generator=generator,
    output_type="latent",  # Critical: returns latent tensor

    return_dict=True,
).latents  # Shape: (B, C, T, H, W)

```

Setting `output_type="latent"` prevents premature VAE decoding, preserving the tensor for refinement.

### Stage 2: Optional VAE Decode and Latent Upsampling

Lines 94-154 handle optional intermediate steps. If `save_sana_output` is enabled, the latent is denormalized using VAE statistics and decoded to a low-resolution video for visual verification.

When `skip_upsampler=False`, the `LTX2LatentUpsamplerModel` increases spatial resolution. The upsampler consumes the original latent and returns a higher-resolution version that is re-normalized to the LTX-2 VAE's latent space before proceeding.

### Stage 3: Manual Latent Packing

Stage 3 (lines 168-174) manually packs latents using `LTX2Pipeline._pack_latents` to bypass Diffusers' default normalization. This step matches the original LTX-2 implementation's expected input format, ensuring compatibility with the Stage-2 model.

### Stage 4: Audio Latent Synthesis

Lines 180-197 create a **zero-audio latent** with the same shape expected by the LTX-2 audio VAE. After LTX-2's normalization passes, this results in a silent audio stream unless you substitute your own audio latent.

```python

# Audio latent shape matches LTX-2 expectations

audio_latent = torch.zeros(
    batch_size, audio_channels, audio_length, 
    dtype=torch.float32, device=device
)

```

### Stage 5: Stage-2 Refinement with LTX-2

The critical refinement step occurs in lines 202-274. The LTX-2 pipeline loads distilled LoRA weights (`stage_2_distilled`) and performs only **three diffusion steps** using the `FlowMatchEulerDiscreteScheduler` with fixed sigma values (`STAGE_2_DISTILLED_SIGMA_VALUES`).

```python
ltx2_pipe.load_lora_weights(
    ltx2_model_id, 
    adapter_name="stage_2_distilled",
    weight_name="ltx-2-19b-distilled-lora-384.safetensors"
)
ltx2_pipe.set_adapters("stage_2_distilled", 1.0)

video, audio = ltx2_pipe(
    latents=packed_video_latent,
    audio_latents=audio_latent,
    prompt=prompt,
    negative_prompt=negative_prompt,
    height=pixel_height,
    width=pixel_width,
    num_frames=pixel_num_frames,
    num_inference_steps=3,  # Distilled model requires only 3 steps

    noise_scale=STAGE_2_DISTILLED_SIGMA_VALUES[0],
    sigmas=STAGE_2_DISTILLED_SIGMA_VALUES,
    guidance_scale=1.0,
    frame_rate=frame_rate,
    generator=generator,
    output_type="np",
    return_dict=False,
)

```

### Stage 6: Video Encoding and Output

Finally, lines 330-444 encode the refined tensor to MP4 format using `encode_video` from `diffusers.pipelines.ltx2.utils`. This utility accepts the `(T, H, W, C)` numpy array and writes the final video file, optionally embedding the audio stream.

## Implementation Examples

### Command-Line Interface

The refiner exposes a CLI through Python Fire (`@cli` decorator), making every function argument available as a command-line flag:

```bash
python -m app.sana_video_refiner_pipeline_diffusers \
    --prompt "A cat and a dog baking a cake together in a kitchen." \
    --sana_model_id /path/to/sana_checkpoint \
    --ltx2_model_id Lightricks/LTX-2 \
    --output_path refined.mp4 \
    --sana_height 720 \
    --sana_width 1280 \
    --motion_score 30

```

### Python API Integration

For programmatic use, import the `sana_video_ltx2_refine` function directly:

```python
from app.sana_video_refiner_pipeline_diffusers import sana_video_ltx2_refine

sana_video_ltx2_refine(
    prompt="A futuristic cityscape at sunset, flying cars in the sky.",
    negative_prompt="low-res, blurry, artifacted",
    sana_model_id="Efficient-Large-Model/SANA-Video_2B_720p",
    ltx2_model_id="Lightricks/LTX-2",
    sana_height=720,
    sana_width=1280,
    sana_frames=81,
    motion_score=30,
    sana_guidance_scale=6.0,
    sana_num_steps=50,
    frame_rate=16.0,
    seed=123,
    output_path="city_refined.mp4",
    save_sana_output="city_raw.mp4",
    skip_upsampler=False,
)

```

The function writes files to `output_path` (refined) and optionally `save_sana_output` (raw Sana output) without returning values.

### Integrating with Existing Sana Video Outputs

If you have existing latents from [`inference_sana_video.py`](https://github.com/NVlabs/Sana/blob/main/inference_sana_video.py), modify the refiner to accept external tensors. You would bypass Stage 1 (lines 53-71) and inject your latent directly before the packing stage (lines 168-174), though this requires minor source modifications to skip the initial generation steps.

## Advanced Configuration Options

**Custom LoRA Weights**: Replace the default `stage_2_distilled` adapter by modifying the `adapter_name` and `weight_name` parameters in the `load_lora_weights` call (lines 202-274).

**Audio Editing**: Substitute the zero-audio latent (lines 180-197) with real audio latents generated by an external audio VAE to retain or modify soundtracks.

**Resolution Scaling**: Increase `sana_height` and `sana_width` beyond 720p, ensuring `skip_upsampler=False` to activate the `LTX2LatentUpsamplerModel` for detail enhancement.

**Checkpoint Flexibility**: Point `ltx2_model_id` to any HuggingFace-hosted LTX-2 checkpoint; the pipeline automatically adapts VAE configurations accordingly.

## Summary

- The **Sana Video Refiner** combines Sana Video generation with LTX-2 Stage-2 refinement in [`app/sana_video_refiner_pipeline_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_video_refiner_pipeline_diffusers.py).
- The workflow maintains latents through six stages: generation, optional decode/upsample, packing, audio synthesis, LTX-2 refinement, and encoding.
- **Stage-2 refinement** uses only 3 diffusion steps with distilled LoRA weights for computational efficiency.
- The pipeline supports both **CLI execution** via Fire and **Python API** integration.
- **Zero-audio latents** ensure silent output unless custom audio is injected.

## Frequently Asked Questions

### How does the Sana Video Refiner differ from standard Sana Video inference?

Standard inference decodes the VAE immediately after generation, producing the final video. The refiner keeps the representation in **latent space** and passes it through LTX-2's Stage-2 model with specialized LoRA weights, resulting in higher fidelity and detail preservation without retraining the full model.

### What hardware requirements are needed for 720p video refinement?

You need a CUDA-capable GPU with sufficient VRAM to hold both the Sana Video model (2B parameters) and the LTX-2 model simultaneously, or sequential loading via `enable_model_cpu_offload`. The latent upsampler (when enabled) requires additional memory proportional to the target resolution.

### Can I use custom audio instead of the silent default?

Yes. Replace the zero-audio latent creation in lines 180-197 with a tensor generated by an audio VAE compatible with LTX-2's expectations. The shape must match `(batch_size, audio_channels, audio_length)` as defined in the LTX-2 configuration.

### Why does Stage-2 use only 3 inference steps?

The Stage-2 model utilizes **distilled LoRA weights** (`ltx-2-19b-distilled-lora-384.safetensors`) trained to converge in fewer steps. According to the source code, these distilled weights work with the `FlowMatchEulerDiscreteScheduler` using fixed `STAGE_2_DISTILLED_SIGMA_VALUES`, making 3 steps sufficient for high-quality refinement while maintaining generation speed.