How to Use the LTX-2 RetakePipeline for Regenerating Specific Video Segments
The LTX-2 RetakePipeline regenerates a precise time window of an existing video while preserving surrounding frames intact, combining latent conditioning, temporal region masking, and selective diffusion denoising.
Local video editing with AI generation demands surgical precision—regenerating only the problematic seconds without re-rendering entire sequences. The RetakePipeline in Lightricks/LTX-2 solves this through a sophisticated multi-stage architecture that masks diffusion to user-specified temporal regions. This article examines the pipeline's implementation, from CLI invocation to Python API integration, based on the actual source code in packages/ltx-pipelines/src/ltx_pipelines/retake.py.
Core Architecture of LTX-2 RetakePipeline
The RetakePipeline orchestrates eight specialized components defined across the ltx-pipelines package. Each stage handles a distinct transformation from source media to regenerated output.
Conditioning and Encoding Stack
The pipeline relies on three primary conditioners defined in [ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py):
- PromptEncoder – Encodes text prompts into video and audio conditioning embeddings. In distilled mode, only positive prompts are processed.
- ImageConditioner – Converts source video frames (or HDR EXR sequences) into VAE latent representations.
- AudioConditioner – Transforms the source audio track into audio VAE latents for synchronized regeneration.
These conditioners feed into the DiffusionStage, which wraps the LTX-2 transformer checkpoint with optional LoRA weights, quantization, and compilation optimizations.
Temporal Region Control with ModalitySpec
The critical innovation for selective regeneration appears in [ltx_core/conditioning/types/noise_mask_cond.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/conditioning/types/noise_mask_cond.py). The TemporalRegionMask class creates a boolean mask over the latent timeline, marking frames within [start_time, end_time] as mutable while freezing all others.
This mask integrates with ModalitySpec objects (defined in [ltx_pipelines/utils/types.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py)) that declare:
conditionings=[TemporalRegionMask(...)]– Specifies which latent regions receive diffusion updates.frozen=True– Explicitly locks unmasked regions against modification.
Denoising Implementations
Two denoiser classes in [ltx_pipelines/utils/denoisers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py) handle the diffusion process:
- GuidedDenoiser – Implements classifier-free guidance (CFG) for full model inference with separate positive and negative prompt conditioning.
- SimpleDenoiser – Lightweight variant for distilled mode using a fixed 8-step sigma schedule (
DISTILLED_SIGMASin [constants.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py)).
Complete LTX-2 RetakePipeline Execution Flow
Understanding the thirteen-stage pipeline execution clarifies how frames remain untouched outside the regeneration window:
-
Metadata extraction –
get_videostream_metadatafrom [media_io.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) reads resolution, frame count, and FPS. -
Latent encoding –
ImageConditionerandAudioConditionercompress source media into VAE-compatible latent spaces. -
Prompt encoding –
PromptEncodergenerates conditioning embeddings; distilled mode skips negative prompts. -
Modality specification –
ModalitySpecinstantiations define mutable and frozen latent regions. -
Sigma schedule selection – Distilled mode uses
DISTILLED_SIGMAS; standard mode generates dynamic schedules viaLTX2Scheduler. -
Regional diffusion –
DiffusionStageexecutes the configured denoiser, applying noise updates exclusively within masked temporal regions. -
Latent decoding –
VideoDecoderandAudioDecoderreconstruct pixel-space frames and waveforms. -
Tiling validation –
ensure_tiling_configfrom [helpers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) verifies VAE compatibility. -
Chunk calculation –
get_video_chunks_numberdetermines tiling parameters for high-resolution inputs. -
Temporal stitching – Masks ensure decoded latents blend seamlessly with preserved surrounding content.
-
Color-space handling – HDR EXR sequences receive appropriate color transformation via
media_ioutilities. -
Video encoding –
encode_videowrites the final output with synchronized audio. -
Constraint validation – Final output enforces LTX-2 requirements: frame counts of
8k+1, dimensions multiples of 32, and matching VAE tiling configuration.
CLI Usage for LTX-2 RetakePipeline
The command-line interface provides immediate access without Python scripting. The entry point resides at [ltx_pipelines/retake.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/retake.py), invoked via module execution:
python -m ltx_pipelines.retake \
--video_path /data/input.mp4 \
--prompt "A sunrise over a misty forest" \
--start_time 5.0 \
--end_time 7.5 \
--seed 42 \
--output_path /data/output_retaken.mp4
Key CLI Arguments from args.py
The [ltx_pipelines/utils/args.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) module defines video_editing_arg_parser with these critical parameters:
| Argument | Type | Description |
|---|---|---|
--video_path |
string | Source video file or EXR frame directory |
--prompt |
string | Generation prompt for the retaken region |
--negative_prompt |
string | Undesired content specification (non-distilled only) |
--start_time |
float | Retake interval start in seconds |
--end_time |
float | Retake interval end in seconds |
--seed |
int | Deterministic generation seed |
--output_path |
string | Destination for regenerated video |
--distilled |
flag | Enable 8-step distilled mode (default: true) |
--enhance_prompt |
flag | Activate prompt enhancement preprocessing |
Python API for LTX-2 RetakePipeline Integration
Custom workflows require direct class instantiation. The RetakePipeline constructor accepts extensive configuration through ModelPaths and runtime parameters.
Basic RetakePipeline Instantiation
from ltx_pipelines.retake import RetakePipeline
from ltx_pipelines.utils.model_paths import ModelPaths
# Initialize model path resolution
model_paths = ModelPaths(root="/models/LTX-2")
# Construct pipeline with distilled inference
pipeline = RetakePipeline(
model_paths=model_paths,
loras=[], # Optional LoRA configuration list
device=None, # Auto-detect CUDA; specify "cpu" or device index
distilled=True, # Use 8-step schedule vs. full CFG
)
Executing the Retake Operation
The __call__ method accepts temporal boundaries and returns a three-tuple of generators and configuration:
video_iter, audio_waveform, tiling_cfg = pipeline(
video_path="/data/input.mp4",
prompt="A sunrise over a misty forest",
start_time=5.0,
end_time=7.5,
seed=42,
fps=None, # Inferred from source metadata when None
enhance_prompt=False,
)
Return values:
video_iter– Generator yielding decoded frame tensorsaudio_waveform– Synchronized audio samples for the retaken regiontiling_cfg– Tiling parameters for high-resolution handling
Encoding and Saving Results
The [media_io.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) utilities handle final output generation:
from ltx_pipelines.utils.media_io import encode_video, get_videostream_metadata
# Extract source properties for accurate encoding
meta = get_videostream_metadata("/data/input.mp4")
# Write final video with proper frame count and audio synchronization
encode_video(
video=video_iter,
fps=int(meta.fps),
audio=audio_waveform,
output_path="/data/output_retaken.mp4",
video_chunks_number=meta.frames, # Adjust for tiling if resolution exceeds VAE limits
color_space=None, # Auto-detect or specify HDR color space
)
High-Resolution Handling and Tiling Configuration
Videos exceeding the VAE's native resolution require automatic tiling. The ensure_tiling_config function in [helpers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) calculates compatible tile dimensions:
from ltx_pipelines.utils.helpers import ensure_tiling_config, get_video_chunks_number
tiling_config = ensure_tiling_config(
width=3840,
height=2160,
vae_tile_size=256, # Checkpoint-specific VAE tile dimension
overlap=32, # Blending overlap between adjacent tiles
)
chunks = get_video_chunks_number(
frames=meta.frames,
tiling_config=tiling_config,
)
The pipeline automatically applies this configuration during VideoDecoder operation, stitching tiled outputs with gradient blending to eliminate seam artifacts.
Summary
- RetakePipeline regenerates precise video segments via
TemporalRegionMaskwhile preserving surrounding content throughModalitySpecfrozen region declarations. - The architecture separates concerns across conditioning (
PromptEncoder,ImageConditioner,AudioConditioner), diffusion (DiffusionStagewithGuidedDenoiserorSimpleDenoiser), and I/O (media_ioencoding/decoding). - Distilled mode (default) uses 8 fixed sigmas for fast inference; full CFG mode enables negative prompting at higher computational cost.
- Implementation imposes LTX-2 constraints: frame counts of
8k+1, dimensions divisible by 32, and VAE-compatible tiling configurations. - Source files cluster in
packages/ltx-pipelines/src/ltx_pipelines/withretake.pyas the primary entry point.
Frequently Asked Questions
What video formats does LTX-2 RetakePipeline accept?
The pipeline accepts standard video files (MP4, MOV via get_videostream_metadata) and HDR EXR frame sequences. The media_io.py module handles format detection and color-space conversion automatically.
How does the temporal mask prevent changes to protected frames?
TemporalRegionMask in noise_mask_cond.py creates a boolean latent mask where True values permit diffusion updates. The DiffusionStage multiplies noise predictions by this mask, ensuring zero gradient flow to frozen regions regardless of denoiser iterations.
Can LoRA weights be applied during video retakes?
Yes. Pass LoRA configuration dictionaries to the loras parameter in RetakePipeline.__init__. The DiffusionStage loads and merges these weights before the first inference call, applying trained modifications to the base transformer checkpoint.
Why must frame counts equal 8k+1?
This constraint derives from the LTX-2 transformer architecture's temporal attention pattern. The model processes frames in non-overlapping groups of 8, requiring a single additional frame to complete the final group. The pipeline validates this condition before latent encoding to prevent runtime failures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →