How to Use the LTX-2 Retake Pipeline to Regenerate Specific Time Regions

The LTX-2 Retake pipeline allows you to replace a selected time interval of an existing video while keeping the rest untouched by using TemporalRegionMask and ModalitySpec configurations in the RetakePipeline class.

The RetakePipeline in the Lightricks/LTX-2 repository provides a dedicated video-editing interface for localized regeneration. Unlike full video generation, this pipeline preserves your original content outside a specified time window, making it ideal for correcting errors or updating specific scenes without re-rendering entire clips.

How the Retake Pipeline Works

The pipeline operates through a six-stage process defined in packages/ltx-pipelines/src/ltx_pipelines/retake.py. Understanding these stages helps you optimize your regeneration parameters and debug issues when they arise.

Step 1 — Loading and Encoding Source Media

First, the pipeline encodes your source video and audio into latent tensors. In RetakePipeline.__call__ (lines 90-104), the image_conditioner and audio_conditioner process the full input streams. These latents serve as the frozen background for regions you do not want to modify.

Step 2 — Building Temporal Masks and Modality Specs

The core mechanism for targeting specific time regions occurs in lines 180-203. Here, the pipeline constructs ModalitySpec objects that wrap:

  • A TemporalRegionMask defining the [start_time, end_time] slice to regenerate
  • A frozen flag indicating that latent values outside the mask should remain unchanged
  • The initial latent tensors from Step 1

The TemporalRegionMask class is implemented in packages/ltx-core/src/ltx_core/conditioning/types/noise_mask_cond.py.

Step 3 — Denoising and Regeneration

The pipeline supports two operational modes selected in lines 236-260:

  • Distilled mode (default for CLI): Uses a fixed 8-step sigma schedule (DISTILLED_SIGMAS) and a SimpleDenoiser that skips classifier-free guidance (CFG) for speed.
  • Full-model mode: Loads an LTX2Scheduler and applies a GuidedDenoiser with multi-modal CFG and STG support.

The DiffusionStage (lines 264-274) runs the denoising loop only on latent regions covered by your mask, leaving frozen regions untouched.

Step 4 — Decoding and Output

Finally, the VideoDecoder and AudioDecoder convert the modified latents back to pixel space and waveform data. The encode_video function from packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py writes the result to disk (lines 278-284).

Input Requirements and Constraints

The RetakePipeline.main function (lines 284-340) enforces strict validation rules that your input must satisfy:

  • Frame count: The source video must contain exactly 8k + 1 frames (e.g., 97, 193, 257). The CLI raises a ValueError if this constraint is violated.
  • Resolution: Width and height must be multiples of 32.
  • Time ordering: start_time must be strictly less than end_time.

If your source video lacks an audio track, the pipeline automatically sets initial_audio_latent to None and processes only the video modality.

Running the Retake Pipeline

CLI Quick Start

The fastest way to regenerate a time region is through the command-line interface provided by video_editing_arg_parser in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py. The main() function handles validation, pipeline construction, and execution.

python -m ltx_pipelines.retake \
    --distilled-checkpoint-path /path/to/ltx-2-distilled.safetensors \
    --gemma-root /path/to/gemma \
    --video-path /data/input.mp4 \
    --start-time 3.5 \
    --end-time 7.2 \
    --prompt "A bright sunrise over a calm lake" \
    --output-path /data/output.mp4 \
    --seed 1234
  • Time values are in seconds.
  • The CLI defaults to distilled=True for fast inference.
  • Use --regenerate-audio or --no-regenerate-video to target specific modalities.

Python API for Full Control

For advanced use cases requiring CFG, LoRA adapters, or custom guidance, instantiate RetakePipeline directly with distilled=False.

from pathlib import Path
import torch

from ltx_pipelines.retake import RetakePipeline
from ltx_core.loader import LoraPathStrengthAndSDOps, LTXV_LORA_COMFY_RENAMING_MAP
from ltx_core.components.guiders import MultiModalGuiderParams
from ltx_core.quantization import build_fp8_cast_policy
from ltx_core.model.video_vae import TilingConfig, get_video_chunks_number
from ltx_pipelines.utils.media_io import encode_video, get_videostream_metadata

# Configuration

ckpt = Path("/models/ltx-2-full.safetensors")
gemma_root = Path("/models/gemma")
video_path = Path("/data/source.mp4")
output_path = Path("/data/retaken.mp4")

# Optional LoRA adapters

loras = [
    LoraPathStrengthAndSDOps(
        "/models/custom_lora.safetensors",
        0.8,
        LTXV_LORA_COMFY_RENAMING_MAP,
    )
]

# Optional quantization for memory reduction

quant = build_fp8_cast_policy(str(ckpt))

# Initialize pipeline with full model capabilities

pipeline = RetakePipeline(
    checkpoint_path=str(ckpt),
    gemma_root=str(gemma_root),
    loras=loras,
    quantization=quant,
    distilled=False,          # Enable CFG and advanced guidance

    offload_mode="cpu",       # Stream weights to CPU for 16GB GPUs

)

# Configure guidance parameters

video_guider = MultiModalGuiderParams(cfg_scale=3.0, stg_scale=1.0)
audio_guider = MultiModalGuiderParams(cfg_scale=5.0, stg_scale=0.8)

# Execute regeneration

video_iter, audio_wave = pipeline(
    video_path=str(video_path),
    prompt="A flock of birds flying over a mountain lake at sunrise",
    start_time=3.5,
    end_time=7.2,
    seed=42,
    video_guider_params=video_guider,
    audio_guider_params=audio_guider,
    regenerate_video=True,
    regenerate_audio=False,      # Preserve original audio

    enhance_prompt=False,
    tiling_config=TilingConfig.default(),
)

# Encode output

meta = get_videostream_metadata(str(video_path))
chunks = get_video_chunks_number(meta.frames, TilingConfig.default())

encode_video(
    video=video_iter,
    fps=int(meta.fps),
    audio=audio_wave,
    output_path=str(output_path),
    video_chunks_number=chunks,
)

Selective Modality Regeneration

You can regenerate only video, only audio, or both by toggling boolean flags:

video_iter, audio_wave = pipeline(
    video_path="/data/source.mp4",
    prompt="Gentle ocean waves with distant seagull calls",
    start_time=0.0,
    end_time=5.0,
    seed=777,
    regenerate_video=False,      # Keep original video

    regenerate_audio=True,       # Regenerate only audio segment

    tiling_config=TilingConfig.default(),
)

Optimizing Performance and Memory

The RetakePipeline accepts offload_mode and quantization arguments to manage GPU memory constraints:

  • Offloading: Set offload_mode="cpu" to stream model weights to system memory between inference steps, reducing VRAM usage at the cost of slight latency.
  • Quantization: Pass a quantization policy from build_fp8_cast_policy() to run the diffusion model in 8-bit precision, significantly reducing memory footprint for long videos.

These options are particularly important when processing high-resolution content (1920×1080 or higher) where latent tensors consume substantial VRAM.

Summary

  • The LTX-2 Retake pipeline preserves existing video content outside a specified time window by using TemporalRegionMask and frozen ModalitySpec configurations.
  • Input videos must satisfy frames = 8k + 1 and resolution multiples of 32, validated in packages/ltx-pipelines/src/ltx_pipelines/retake.py.
  • Distilled mode (default CLI) offers 8-step fast inference without CFG, while full-model mode enables classifier-free guidance and STG through MultiModalGuiderParams.
  • You can regenerate video, audio, or both independently using the regenerate_video and regenerate_audio flags.
  • Memory optimization is available via offload_mode and build_fp8_cast_policy() quantization.

Frequently Asked Questions

What video formats does the Retake pipeline support?

The pipeline accepts standard video formats readable by the underlying FFmpeg-based loader, though it outputs MP4 files via encode_video. The source video must meet the specific frame count (8k + 1) and resolution constraints (multiples of 32) enforced by the validation logic in the main() function.

Why does my video need exactly 8k+1 frames?

This constraint aligns with the temporal compression ratio of the LTX-2 Video VAE architecture. The autoencoder downsamples the time dimension by a factor of 8 during encoding, requiring the frame count to follow the 8k + 1 pattern to ensure symmetrical encoding and decoding without padding artifacts.

Can I use LoRA adapters with the Retake pipeline?

Yes. The RetakePipeline constructor accepts a loras parameter expecting a list of LoraPathStrengthAndSDOps objects. You can load multiple adapters simultaneously, and the pipeline will apply them during the diffusion stage to influence the style of the regenerated region.

How do I fix "Resolution must be multiple of 32" errors?

The VAE architecture requires spatial dimensions divisible by 32 due to 4× spatial downsampling in both height and width. Resize your input video to valid dimensions (e.g., 1920×1080, 1280×720) before processing, or use a preprocessing script to pad frames to the nearest multiple of 32.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →