Implementing Video-to-Video Editing with Sana Video Refiner: A Complete Guide
The Sana Video Refiner enables high-fidelity video-to-video editing by combining the Sana video diffusion model with Lightricks LTX-2 upsampling through a six-stage latent refinement pipeline.
The Sana Video Refiner in the NVlabs/Sana repository provides a modular approach to video editing that leverages latent diffusion techniques to transform input prompts into polished, high-resolution video outputs. According to the source code in app/sana_video_refiner_pipeline_diffusers.py, the pipeline orchestrates two distinct models: Sana Video for initial latent generation and LTX-2 for Stage-2 refinement with specialized LoRA weights.
What Is the Sana Video Refiner?
The refiner is a hybrid pipeline that bridges Sana Video (a 720p-capable video diffusion model) with LTX-2 (a high-fidelity upsampling and refinement system). Unlike standard video generation scripts, this implementation keeps everything in the latent space until the final encoding step, minimizing GPU memory usage while maximizing output quality.
The architecture is intentionally modular, allowing you to swap LTX-2 checkpoints, inject custom LoRA weights for style adaptation, and bypass default Diffusers normalization to match the original LTX-2 implementation.
Prerequisites and Setup
Before running the refiner, ensure your environment meets these requirements:
- Python 3.8+ with PyTorch 2.0 or higher
- Diffusers 0.32+ for pipeline compatibility
- CUDA-capable GPU with sufficient VRAM for 720p video generation
Install dependencies from the repository root:
pip install -r requirements.txt
The refiner depends on specific model checkpoints: Efficient-Large-Model/SANA-Video_2B_720p for the base generation and Lightricks/LTX-2 (or compatible checkpoints) for Stage-2 refinement.
The Six-Stage Refinement Workflow
The implementation in sana_video_refiner_pipeline_diffusers.py follows a precise six-step workflow, with key operations occurring at specific line ranges.
Stage 1: Latent Generation with SanaVideoPipeline
The process begins by generating a latent representation rather than a decoded video. In lines 53-71, the SanaVideoPipeline creates a compressed spatio-temporal tensor that respects the prompt's motion characteristics.
sana_pipe = SanaVideoPipeline.from_pretrained(
sana_model_id, torch_dtype=dtype
)
sana_pipe.text_encoder.to(dtype)
sana_pipe.enable_model_cpu_offload(gpu_id=0)
video_latent = sana_pipe(
prompt=full_prompt,
negative_prompt=negative_prompt,
height=sana_height,
width=sana_width,
frames=sana_frames,
guidance_scale=sana_guidance_scale,
num_inference_steps=sana_num_steps,
generator=generator,
output_type="latent", # Critical: returns latent tensor
return_dict=True,
).latents # Shape: (B, C, T, H, W)
Setting output_type="latent" prevents premature VAE decoding, preserving the tensor for refinement.
Stage 2: Optional VAE Decode and Latent Upsampling
Lines 94-154 handle optional intermediate steps. If save_sana_output is enabled, the latent is denormalized using VAE statistics and decoded to a low-resolution video for visual verification.
When skip_upsampler=False, the LTX2LatentUpsamplerModel increases spatial resolution. The upsampler consumes the original latent and returns a higher-resolution version that is re-normalized to the LTX-2 VAE's latent space before proceeding.
Stage 3: Manual Latent Packing
Stage 3 (lines 168-174) manually packs latents using LTX2Pipeline._pack_latents to bypass Diffusers' default normalization. This step matches the original LTX-2 implementation's expected input format, ensuring compatibility with the Stage-2 model.
Stage 4: Audio Latent Synthesis
Lines 180-197 create a zero-audio latent with the same shape expected by the LTX-2 audio VAE. After LTX-2's normalization passes, this results in a silent audio stream unless you substitute your own audio latent.
# Audio latent shape matches LTX-2 expectations
audio_latent = torch.zeros(
batch_size, audio_channels, audio_length,
dtype=torch.float32, device=device
)
Stage 5: Stage-2 Refinement with LTX-2
The critical refinement step occurs in lines 202-274. The LTX-2 pipeline loads distilled LoRA weights (stage_2_distilled) and performs only three diffusion steps using the FlowMatchEulerDiscreteScheduler with fixed sigma values (STAGE_2_DISTILLED_SIGMA_VALUES).
ltx2_pipe.load_lora_weights(
ltx2_model_id,
adapter_name="stage_2_distilled",
weight_name="ltx-2-19b-distilled-lora-384.safetensors"
)
ltx2_pipe.set_adapters("stage_2_distilled", 1.0)
video, audio = ltx2_pipe(
latents=packed_video_latent,
audio_latents=audio_latent,
prompt=prompt,
negative_prompt=negative_prompt,
height=pixel_height,
width=pixel_width,
num_frames=pixel_num_frames,
num_inference_steps=3, # Distilled model requires only 3 steps
noise_scale=STAGE_2_DISTILLED_SIGMA_VALUES[0],
sigmas=STAGE_2_DISTILLED_SIGMA_VALUES,
guidance_scale=1.0,
frame_rate=frame_rate,
generator=generator,
output_type="np",
return_dict=False,
)
Stage 6: Video Encoding and Output
Finally, lines 330-444 encode the refined tensor to MP4 format using encode_video from diffusers.pipelines.ltx2.utils. This utility accepts the (T, H, W, C) numpy array and writes the final video file, optionally embedding the audio stream.
Implementation Examples
Command-Line Interface
The refiner exposes a CLI through Python Fire (@cli decorator), making every function argument available as a command-line flag:
python -m app.sana_video_refiner_pipeline_diffusers \
--prompt "A cat and a dog baking a cake together in a kitchen." \
--sana_model_id /path/to/sana_checkpoint \
--ltx2_model_id Lightricks/LTX-2 \
--output_path refined.mp4 \
--sana_height 720 \
--sana_width 1280 \
--motion_score 30
Python API Integration
For programmatic use, import the sana_video_ltx2_refine function directly:
from app.sana_video_refiner_pipeline_diffusers import sana_video_ltx2_refine
sana_video_ltx2_refine(
prompt="A futuristic cityscape at sunset, flying cars in the sky.",
negative_prompt="low-res, blurry, artifacted",
sana_model_id="Efficient-Large-Model/SANA-Video_2B_720p",
ltx2_model_id="Lightricks/LTX-2",
sana_height=720,
sana_width=1280,
sana_frames=81,
motion_score=30,
sana_guidance_scale=6.0,
sana_num_steps=50,
frame_rate=16.0,
seed=123,
output_path="city_refined.mp4",
save_sana_output="city_raw.mp4",
skip_upsampler=False,
)
The function writes files to output_path (refined) and optionally save_sana_output (raw Sana output) without returning values.
Integrating with Existing Sana Video Outputs
If you have existing latents from inference_sana_video.py, modify the refiner to accept external tensors. You would bypass Stage 1 (lines 53-71) and inject your latent directly before the packing stage (lines 168-174), though this requires minor source modifications to skip the initial generation steps.
Advanced Configuration Options
Custom LoRA Weights: Replace the default stage_2_distilled adapter by modifying the adapter_name and weight_name parameters in the load_lora_weights call (lines 202-274).
Audio Editing: Substitute the zero-audio latent (lines 180-197) with real audio latents generated by an external audio VAE to retain or modify soundtracks.
Resolution Scaling: Increase sana_height and sana_width beyond 720p, ensuring skip_upsampler=False to activate the LTX2LatentUpsamplerModel for detail enhancement.
Checkpoint Flexibility: Point ltx2_model_id to any HuggingFace-hosted LTX-2 checkpoint; the pipeline automatically adapts VAE configurations accordingly.
Summary
- The Sana Video Refiner combines Sana Video generation with LTX-2 Stage-2 refinement in
app/sana_video_refiner_pipeline_diffusers.py. - The workflow maintains latents through six stages: generation, optional decode/upsample, packing, audio synthesis, LTX-2 refinement, and encoding.
- Stage-2 refinement uses only 3 diffusion steps with distilled LoRA weights for computational efficiency.
- The pipeline supports both CLI execution via Fire and Python API integration.
- Zero-audio latents ensure silent output unless custom audio is injected.
Frequently Asked Questions
How does the Sana Video Refiner differ from standard Sana Video inference?
Standard inference decodes the VAE immediately after generation, producing the final video. The refiner keeps the representation in latent space and passes it through LTX-2's Stage-2 model with specialized LoRA weights, resulting in higher fidelity and detail preservation without retraining the full model.
What hardware requirements are needed for 720p video refinement?
You need a CUDA-capable GPU with sufficient VRAM to hold both the Sana Video model (2B parameters) and the LTX-2 model simultaneously, or sequential loading via enable_model_cpu_offload. The latent upsampler (when enabled) requires additional memory proportional to the target resolution.
Can I use custom audio instead of the silent default?
Yes. Replace the zero-audio latent creation in lines 180-197 with a tensor generated by an audio VAE compatible with LTX-2's expectations. The shape must match (batch_size, audio_channels, audio_length) as defined in the LTX-2 configuration.
Why does Stage-2 use only 3 inference steps?
The Stage-2 model utilizes distilled LoRA weights (ltx-2-19b-distilled-lora-384.safetensors) trained to converge in fewer steps. According to the source code, these distilled weights work with the FlowMatchEulerDiscreteScheduler using fixed STAGE_2_DISTILLED_SIGMA_VALUES, making 3 steps sufficient for high-quality refinement while maintaining generation speed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →