How to Use the LTX-2 RetakePipeline for Regenerating Specific Video Segments

The LTX-2 RetakePipeline regenerates a precise time window of an existing video while preserving surrounding frames intact, combining latent conditioning, temporal region masking, and selective diffusion denoising.

Local video editing with AI generation demands surgical precision—regenerating only the problematic seconds without re-rendering entire sequences. The RetakePipeline in Lightricks/LTX-2 solves this through a sophisticated multi-stage architecture that masks diffusion to user-specified temporal regions. This article examines the pipeline's implementation, from CLI invocation to Python API integration, based on the actual source code in packages/ltx-pipelines/src/ltx_pipelines/retake.py.


Core Architecture of LTX-2 RetakePipeline

The RetakePipeline orchestrates eight specialized components defined across the ltx-pipelines package. Each stage handles a distinct transformation from source media to regenerated output.

Conditioning and Encoding Stack

The pipeline relies on three primary conditioners defined in [ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py):

  • PromptEncoder – Encodes text prompts into video and audio conditioning embeddings. In distilled mode, only positive prompts are processed.
  • ImageConditioner – Converts source video frames (or HDR EXR sequences) into VAE latent representations.
  • AudioConditioner – Transforms the source audio track into audio VAE latents for synchronized regeneration.

These conditioners feed into the DiffusionStage, which wraps the LTX-2 transformer checkpoint with optional LoRA weights, quantization, and compilation optimizations.

Temporal Region Control with ModalitySpec

The critical innovation for selective regeneration appears in [ltx_core/conditioning/types/noise_mask_cond.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/conditioning/types/noise_mask_cond.py). The TemporalRegionMask class creates a boolean mask over the latent timeline, marking frames within [start_time, end_time] as mutable while freezing all others.

This mask integrates with ModalitySpec objects (defined in [ltx_pipelines/utils/types.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py)) that declare:

  • conditionings=[TemporalRegionMask(...)] – Specifies which latent regions receive diffusion updates.
  • frozen=True – Explicitly locks unmasked regions against modification.

Denoising Implementations

Two denoiser classes in [ltx_pipelines/utils/denoisers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py) handle the diffusion process:


Complete LTX-2 RetakePipeline Execution Flow

Understanding the thirteen-stage pipeline execution clarifies how frames remain untouched outside the regeneration window:

  1. Metadata extraction – get_videostream_metadata from [media_io.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) reads resolution, frame count, and FPS.

  2. Latent encoding – ImageConditioner and AudioConditioner compress source media into VAE-compatible latent spaces.

  3. Prompt encoding – PromptEncoder generates conditioning embeddings; distilled mode skips negative prompts.

  4. Modality specification – ModalitySpec instantiations define mutable and frozen latent regions.

  5. Sigma schedule selection – Distilled mode uses DISTILLED_SIGMAS; standard mode generates dynamic schedules via LTX2Scheduler.

  6. Regional diffusion – DiffusionStage executes the configured denoiser, applying noise updates exclusively within masked temporal regions.

  7. Latent decoding – VideoDecoder and AudioDecoder reconstruct pixel-space frames and waveforms.

  8. Tiling validation – ensure_tiling_config from [helpers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) verifies VAE compatibility.

  9. Chunk calculation – get_video_chunks_number determines tiling parameters for high-resolution inputs.

  10. Temporal stitching – Masks ensure decoded latents blend seamlessly with preserved surrounding content.

  11. Color-space handling – HDR EXR sequences receive appropriate color transformation via media_io utilities.

  12. Video encoding – encode_video writes the final output with synchronized audio.

  13. Constraint validation – Final output enforces LTX-2 requirements: frame counts of 8k+1, dimensions multiples of 32, and matching VAE tiling configuration.


CLI Usage for LTX-2 RetakePipeline

The command-line interface provides immediate access without Python scripting. The entry point resides at [ltx_pipelines/retake.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/retake.py), invoked via module execution:

python -m ltx_pipelines.retake \
    --video_path /data/input.mp4 \
    --prompt "A sunrise over a misty forest" \
    --start_time 5.0 \
    --end_time 7.5 \
    --seed 42 \
    --output_path /data/output_retaken.mp4

Key CLI Arguments from args.py

The [ltx_pipelines/utils/args.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) module defines video_editing_arg_parser with these critical parameters:

Argument Type Description
--video_path string Source video file or EXR frame directory
--prompt string Generation prompt for the retaken region
--negative_prompt string Undesired content specification (non-distilled only)
--start_time float Retake interval start in seconds
--end_time float Retake interval end in seconds
--seed int Deterministic generation seed
--output_path string Destination for regenerated video
--distilled flag Enable 8-step distilled mode (default: true)
--enhance_prompt flag Activate prompt enhancement preprocessing

Python API for LTX-2 RetakePipeline Integration

Custom workflows require direct class instantiation. The RetakePipeline constructor accepts extensive configuration through ModelPaths and runtime parameters.

Basic RetakePipeline Instantiation

from ltx_pipelines.retake import RetakePipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Initialize model path resolution

model_paths = ModelPaths(root="/models/LTX-2")

# Construct pipeline with distilled inference

pipeline = RetakePipeline(
    model_paths=model_paths,
    loras=[],                     # Optional LoRA configuration list

    device=None,                  # Auto-detect CUDA; specify "cpu" or device index

    distilled=True,               # Use 8-step schedule vs. full CFG

)

Executing the Retake Operation

The __call__ method accepts temporal boundaries and returns a three-tuple of generators and configuration:

video_iter, audio_waveform, tiling_cfg = pipeline(
    video_path="/data/input.mp4",
    prompt="A sunrise over a misty forest",
    start_time=5.0,
    end_time=7.5,
    seed=42,
    fps=None,                     # Inferred from source metadata when None

    enhance_prompt=False,
)

Return values:

  • video_iter – Generator yielding decoded frame tensors
  • audio_waveform – Synchronized audio samples for the retaken region
  • tiling_cfg – Tiling parameters for high-resolution handling

Encoding and Saving Results

The [media_io.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) utilities handle final output generation:

from ltx_pipelines.utils.media_io import encode_video, get_videostream_metadata

# Extract source properties for accurate encoding

meta = get_videostream_metadata("/data/input.mp4")

# Write final video with proper frame count and audio synchronization

encode_video(
    video=video_iter,
    fps=int(meta.fps),
    audio=audio_waveform,
    output_path="/data/output_retaken.mp4",
    video_chunks_number=meta.frames,   # Adjust for tiling if resolution exceeds VAE limits

    color_space=None,                  # Auto-detect or specify HDR color space

)

High-Resolution Handling and Tiling Configuration

Videos exceeding the VAE's native resolution require automatic tiling. The ensure_tiling_config function in [helpers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) calculates compatible tile dimensions:

from ltx_pipelines.utils.helpers import ensure_tiling_config, get_video_chunks_number

tiling_config = ensure_tiling_config(
    width=3840,
    height=2160,
    vae_tile_size=256,        # Checkpoint-specific VAE tile dimension

    overlap=32,               # Blending overlap between adjacent tiles

)

chunks = get_video_chunks_number(
    frames=meta.frames,
    tiling_config=tiling_config,
)

The pipeline automatically applies this configuration during VideoDecoder operation, stitching tiled outputs with gradient blending to eliminate seam artifacts.


Summary

  • RetakePipeline regenerates precise video segments via TemporalRegionMask while preserving surrounding content through ModalitySpec frozen region declarations.
  • The architecture separates concerns across conditioning (PromptEncoder, ImageConditioner, AudioConditioner), diffusion (DiffusionStage with GuidedDenoiser or SimpleDenoiser), and I/O (media_io encoding/decoding).
  • Distilled mode (default) uses 8 fixed sigmas for fast inference; full CFG mode enables negative prompting at higher computational cost.
  • Implementation imposes LTX-2 constraints: frame counts of 8k+1, dimensions divisible by 32, and VAE-compatible tiling configurations.
  • Source files cluster in packages/ltx-pipelines/src/ltx_pipelines/ with retake.py as the primary entry point.

Frequently Asked Questions

What video formats does LTX-2 RetakePipeline accept?

The pipeline accepts standard video files (MP4, MOV via get_videostream_metadata) and HDR EXR frame sequences. The media_io.py module handles format detection and color-space conversion automatically.

How does the temporal mask prevent changes to protected frames?

TemporalRegionMask in noise_mask_cond.py creates a boolean latent mask where True values permit diffusion updates. The DiffusionStage multiplies noise predictions by this mask, ensuring zero gradient flow to frozen regions regardless of denoiser iterations.

Can LoRA weights be applied during video retakes?

Yes. Pass LoRA configuration dictionaries to the loras parameter in RetakePipeline.__init__. The DiffusionStage loads and merges these weights before the first inference call, applying trained modifications to the base transformer checkpoint.

Why must frame counts equal 8k+1?

This constraint derives from the LTX-2 transformer architecture's temporal attention pattern. The model processes frames in non-overlapping groups of 8, requiring a single additional frame to complete the final group. The pipeline validates this condition before latent encoding to prevent runtime failures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →