How to Use RetakePipeline for Selective Video Region Regeneration in LTX‑2

Use ltx_pipelines.retake.RetakePipeline to regenerate a specific time segment of a video while freezing the rest, by passing start_time, end_time, and optional regenerate_video/regenerate_audio flags.

The Retake Pipeline in LTX‑2 enables precise, non‑destructive editing of video content. Instead of regenerating an entire video, you can target a specific temporal window—measured in seconds—and replace only that region with AI‑generated content driven by a text prompt. This guide walks through the architecture, key parameters, and practical implementations for both Python and CLI workflows.

What the Retake Pipeline Does

RetakePipeline is purpose‑built for selective video region regeneration. It accepts a source video (file path or EXR‑frame folder), encodes it into latent space, applies a temporal mask to isolate the target segment, runs diffusion only on that masked region, and decodes the result back to pixels and audio.

The pipeline leaves unmasked portions of the video and audio intact. You can independently control whether video frames, audio waveform, or both get regenerated within the specified window.

Core Architecture and File Locations

Understanding the internal flow helps debug issues and customize behavior. The pipeline is implemented in packages/ltx-pipelines/src/ltx_pipelines/retake.py.

Pipeline Initialization and Sub‑modules

The RetakePipeline.__init__ method (lines 53‑90) constructs all required components:

  • Prompt encoder for text conditioning
  • Image/audio conditioners for cross‑modal guidance
  • Diffusion stage for latent denoising
  • Video and audio decoders for final output

It also sets defaults for device, dtype, distilled mode, and handles LoRA loading.

Input Validation Requirements

The CLI entry point uses video_editing_arg_parser from packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 75‑84). This enforces a critical rule:

  • EXR‑frame folders require --frame-rate (no container FPS exists)
  • Standard video files must not receive --frame-rate (FPS extracted from container)

This validation occurs in _verify_media_path_args and prevents silent frame‑rate mismatches in HDR workflows.

Temporal Masking Mechanism

The TemporalRegionMask class (defined in packages/ltx-core/src/ltx_core/conditioning/types/noise_mask_cond.py) converts start_time and end_time (seconds) into frame indices using the video FPS. These masks attach to video and audio ModalitySpec instances in retake.py (lines 68‑75).

Each ModalitySpec carries a frozen flag. When regenerate_video=False or regenerate_audio=False, the mask list is empty and the corresponding latent remains unchanged through the diffusion process.

Denoising Modes: Distilled vs. Full

RetakePipeline supports two inference modes controlled by the distilled parameter:

Mode Denoiser Steps Guidance
Distilled (distilled=True, default for CLI) SimpleDenoiser Fixed 8‑step sigma schedule (DISTILLED_SIGMAS) None (fast)
Full model (distilled=False) GuidedDenoiser + MultiModalGuider User‑defined via num_inference_steps Full video/audio guidance

Selection logic appears in retake.py lines 90‑106.

Python API: Complete Working Example

Here is a full workflow using the full model with custom inference steps:

from ltx_pipelines.retake import RetakePipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.media_io import (
    encode_video,
    get_videostream_metadata,
    get_video_chunks_number,
)

# 1. Initialize model paths from checkpoint directory

model_paths = ModelPaths.from_checkpoint_dir("/path/to/checkpoint_dir")

# 2. Build pipeline in full-model mode

pipeline = RetakePipeline(
    model_paths=model_paths,
    loras=[],                       # Optional LoRA weights

    distilled=False,                # Enable full guidance

    device="cuda",
)

# 3. Execute selective regeneration

video_iter, audio_wave, tiling_cfg = pipeline(
    video_path="source_video.mp4",
    prompt="A futuristic cityscape with neon lights",
    start_time=3.0,                 # Start regeneration at 3 seconds

    end_time=6.5,                   # End at 6.5 seconds

    seed=42,
    fps=None,                       # Auto-detect from MP4 container

    regenerate_video=True,
    regenerate_audio=False,         # Preserve original audio

    num_inference_steps=60,         # Higher quality, slower inference

)

# 4. Encode and save output

metadata = get_videostream_metadata("source_video.mp4")
chunks = get_video_chunks_number(metadata.frames, tiling_cfg)

encode_video(
    video=video_iter,
    fps=int(metadata.fps),
    audio=audio_wave,
    output_path="video_segment_regenerated.mp4",
    video_chunks_number=chunks,
)

Key parameters explained:

  • regenerate_audio=False — Sets frozen=True on the audio ModalitySpec, preserving original sound
  • distilled=False — Activates GuidedDenoiser with full MultiModalGuider support
  • num_inference_steps — Only respected when distilled=False

CLI Usage: Quick Reference

Standard Video with Distilled Model

python -m ltx_pipelines.retake \
    --distilled-checkpoint-path /models/ltx2-distilled.safetensors \
    --video-path input.mp4 \
    --prompt "A dramatic sunset over the ocean" \
    --start-time 5.0 \
    --end-time 8.0 \
    --seed 123 \
    --output-path sunset_retake.mp4

No --frame-rate needed—FPS is read from the MP4 container.

EXR Sequence (HDR Workflow)

python -m ltx_pipelines.retake \
    --distilled-checkpoint-path /models/ltx2-distilled.safetensors \
    --video-path /data/scene_exr/ \
    --frame-rate 24 \
    --prompt "Morning fog in a forest" \
    --start-time 1.0 \
    --end-time 3.5 \
    --seed 456 \
    --output-path forest_fog.mp4

Critical: --frame-rate 24 is mandatory for EXR folders. The validation in args.py will reject the command otherwise.

Audio‑Only Regeneration

python -m ltx_pipelines.retake \
    --distilled-checkpoint-path /models/ltx2-distilled.safetensors \
    --video-path input.mp4 \
    --prompt "Thunderstorm ambience" \
    --start-time 10.0 \
    --end-time 14.0 \
    --seed 789 \
    --regenerate-video false \
    --regenerate-audio true \
    --output-path storm_audio.mp4

--regenerate-video false freezes all video latents; only the 4‑second audio window is regenerated.

Advanced Configuration Options

Controlling Latent Freezing

Pass regenerate_video and regenerate_audio as booleans to RetakePipeline.__call__. These map directly to the frozen attribute in ModalitySpec (defined in packages/ltx-pipelines/src/ltx_pipelines/utils/types.py):

  • True — Latent is masked and denoised
  • False — Latent is frozen, original content preserved

Sigma Schedules and Step Counts

The distilled parameter determines schedule behavior:

  • Distilled mode — Fixed schedule DISTILLED_SIGMAS (8 steps), num_inference_steps ignored
  • Full mode — User‑defined num_inference_steps; schedule computed per‑call

Full mode enables quality‑vs‑speed tradeoffs for production workflows.

LoRA Integration

Pass trained LoRA paths to the loras list during pipeline construction:

pipeline = RetakePipeline(
    model_paths=model_paths,
    loras=["/path/to/style_lora.safetensors"],
    distilled=True,
)

LoRA weights are loaded once at initialization and applied during prompt encoding.

Key Source Files Reference

File Purpose
packages/ltx-pipelines/src/ltx_pipelines/retake.py RetakePipeline class, mask handling, diffusion execution, CLI entry point
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py video_editing_arg_parser, EXR/frame‑rate validation
packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py video_latent_from_file, audio_latent_from_file, encode_video
packages/ltx-pipelines/src/ltx_pipelines/utils/types.py ModalitySpec, OffloadMode definitions
packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py ModelPaths checkpoint resolution
packages/ltx-core/src/ltx_core/conditioning/types/noise_mask_cond.py TemporalRegionMask implementation

Summary

  • RetakePipeline enables selective video region regeneration by time window, preserving content outside the mask
  • TemporalRegionMask converts start_time/end_time to frame indices using source FPS
  • EXR sequences require --frame-rate; standard videos must omit it
  • regenerate_video/regenerate_audio flags independently control which modalities are modified
  • distilled=True (default) uses fast 8‑step inference; distilled=False enables full guidance with configurable steps
  • All core logic resides in retake.py with validation in args.py and media I/O in media_io.py

Frequently Asked Questions

How do I regenerate only audio without touching video?

Set regenerate_video=False and regenerate_audio=True. In CLI: --regenerate-video false --regenerate-audio true. This freezes the video latent and applies diffusion only to the audio latent within the specified time window. The original video frames remain pixel‑identical.

Why does my EXR sequence fail with a frame‑rate error?

EXR folders lack embedded FPS metadata. The validation in _verify_media_path_args (lines 75‑84 of args.py) requires --frame-rate for EXR inputs and rejects it for containerized video files. Add --frame-rate 24 (or your target FPS) to resolve.

Can I use different step counts with the distilled model?

No. The distilled checkpoint is optimized for a fixed 8‑step sigma schedule (DISTILLED_SIGMAS). To control num_inference_steps, set distilled=False when constructing RetakePipeline. This activates GuidedDenoiser with MultiModalGuider for full conditional guidance.

What happens if start_time and end_time span the entire video?

The temporal mask covers all frames, effectively performing full video regeneration. However, regenerate_video=False or regenerate_audio=False will still freeze the respective modality. For true full regeneration, use ltx_pipelines.video_generation instead—RetakePipeline carries overhead for masking logic you don't need.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →