How to Use RetakePipeline for Selective Video Region Regeneration in LTX‑2
Use ltx_pipelines.retake.RetakePipeline to regenerate a specific time segment of a video while freezing the rest, by passing start_time, end_time, and optional regenerate_video/regenerate_audio flags.
The Retake Pipeline in LTX‑2 enables precise, non‑destructive editing of video content. Instead of regenerating an entire video, you can target a specific temporal window—measured in seconds—and replace only that region with AI‑generated content driven by a text prompt. This guide walks through the architecture, key parameters, and practical implementations for both Python and CLI workflows.
What the Retake Pipeline Does
RetakePipeline is purpose‑built for selective video region regeneration. It accepts a source video (file path or EXR‑frame folder), encodes it into latent space, applies a temporal mask to isolate the target segment, runs diffusion only on that masked region, and decodes the result back to pixels and audio.
The pipeline leaves unmasked portions of the video and audio intact. You can independently control whether video frames, audio waveform, or both get regenerated within the specified window.
Core Architecture and File Locations
Understanding the internal flow helps debug issues and customize behavior. The pipeline is implemented in packages/ltx-pipelines/src/ltx_pipelines/retake.py.
Pipeline Initialization and Sub‑modules
The RetakePipeline.__init__ method (lines 53‑90) constructs all required components:
- Prompt encoder for text conditioning
- Image/audio conditioners for cross‑modal guidance
- Diffusion stage for latent denoising
- Video and audio decoders for final output
It also sets defaults for device, dtype, distilled mode, and handles LoRA loading.
Input Validation Requirements
The CLI entry point uses video_editing_arg_parser from packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 75‑84). This enforces a critical rule:
- EXR‑frame folders require
--frame-rate(no container FPS exists) - Standard video files must not receive
--frame-rate(FPS extracted from container)
This validation occurs in _verify_media_path_args and prevents silent frame‑rate mismatches in HDR workflows.
Temporal Masking Mechanism
The TemporalRegionMask class (defined in packages/ltx-core/src/ltx_core/conditioning/types/noise_mask_cond.py) converts start_time and end_time (seconds) into frame indices using the video FPS. These masks attach to video and audio ModalitySpec instances in retake.py (lines 68‑75).
Each ModalitySpec carries a frozen flag. When regenerate_video=False or regenerate_audio=False, the mask list is empty and the corresponding latent remains unchanged through the diffusion process.
Denoising Modes: Distilled vs. Full
RetakePipeline supports two inference modes controlled by the distilled parameter:
| Mode | Denoiser | Steps | Guidance |
|---|---|---|---|
Distilled (distilled=True, default for CLI) |
SimpleDenoiser |
Fixed 8‑step sigma schedule (DISTILLED_SIGMAS) |
None (fast) |
Full model (distilled=False) |
GuidedDenoiser + MultiModalGuider |
User‑defined via num_inference_steps |
Full video/audio guidance |
Selection logic appears in retake.py lines 90‑106.
Python API: Complete Working Example
Here is a full workflow using the full model with custom inference steps:
from ltx_pipelines.retake import RetakePipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.media_io import (
encode_video,
get_videostream_metadata,
get_video_chunks_number,
)
# 1. Initialize model paths from checkpoint directory
model_paths = ModelPaths.from_checkpoint_dir("/path/to/checkpoint_dir")
# 2. Build pipeline in full-model mode
pipeline = RetakePipeline(
model_paths=model_paths,
loras=[], # Optional LoRA weights
distilled=False, # Enable full guidance
device="cuda",
)
# 3. Execute selective regeneration
video_iter, audio_wave, tiling_cfg = pipeline(
video_path="source_video.mp4",
prompt="A futuristic cityscape with neon lights",
start_time=3.0, # Start regeneration at 3 seconds
end_time=6.5, # End at 6.5 seconds
seed=42,
fps=None, # Auto-detect from MP4 container
regenerate_video=True,
regenerate_audio=False, # Preserve original audio
num_inference_steps=60, # Higher quality, slower inference
)
# 4. Encode and save output
metadata = get_videostream_metadata("source_video.mp4")
chunks = get_video_chunks_number(metadata.frames, tiling_cfg)
encode_video(
video=video_iter,
fps=int(metadata.fps),
audio=audio_wave,
output_path="video_segment_regenerated.mp4",
video_chunks_number=chunks,
)
Key parameters explained:
regenerate_audio=False— Setsfrozen=Trueon the audioModalitySpec, preserving original sounddistilled=False— ActivatesGuidedDenoiserwith fullMultiModalGuidersupportnum_inference_steps— Only respected whendistilled=False
CLI Usage: Quick Reference
Standard Video with Distilled Model
python -m ltx_pipelines.retake \
--distilled-checkpoint-path /models/ltx2-distilled.safetensors \
--video-path input.mp4 \
--prompt "A dramatic sunset over the ocean" \
--start-time 5.0 \
--end-time 8.0 \
--seed 123 \
--output-path sunset_retake.mp4
No --frame-rate needed—FPS is read from the MP4 container.
EXR Sequence (HDR Workflow)
python -m ltx_pipelines.retake \
--distilled-checkpoint-path /models/ltx2-distilled.safetensors \
--video-path /data/scene_exr/ \
--frame-rate 24 \
--prompt "Morning fog in a forest" \
--start-time 1.0 \
--end-time 3.5 \
--seed 456 \
--output-path forest_fog.mp4
Critical: --frame-rate 24 is mandatory for EXR folders. The validation in args.py will reject the command otherwise.
Audio‑Only Regeneration
python -m ltx_pipelines.retake \
--distilled-checkpoint-path /models/ltx2-distilled.safetensors \
--video-path input.mp4 \
--prompt "Thunderstorm ambience" \
--start-time 10.0 \
--end-time 14.0 \
--seed 789 \
--regenerate-video false \
--regenerate-audio true \
--output-path storm_audio.mp4
--regenerate-video false freezes all video latents; only the 4‑second audio window is regenerated.
Advanced Configuration Options
Controlling Latent Freezing
Pass regenerate_video and regenerate_audio as booleans to RetakePipeline.__call__. These map directly to the frozen attribute in ModalitySpec (defined in packages/ltx-pipelines/src/ltx_pipelines/utils/types.py):
True— Latent is masked and denoisedFalse— Latent is frozen, original content preserved
Sigma Schedules and Step Counts
The distilled parameter determines schedule behavior:
- Distilled mode — Fixed schedule
DISTILLED_SIGMAS(8 steps),num_inference_stepsignored - Full mode — User‑defined
num_inference_steps; schedule computed per‑call
Full mode enables quality‑vs‑speed tradeoffs for production workflows.
LoRA Integration
Pass trained LoRA paths to the loras list during pipeline construction:
pipeline = RetakePipeline(
model_paths=model_paths,
loras=["/path/to/style_lora.safetensors"],
distilled=True,
)
LoRA weights are loaded once at initialization and applied during prompt encoding.
Key Source Files Reference
| File | Purpose |
|---|---|
packages/ltx-pipelines/src/ltx_pipelines/retake.py |
RetakePipeline class, mask handling, diffusion execution, CLI entry point |
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py |
video_editing_arg_parser, EXR/frame‑rate validation |
packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py |
video_latent_from_file, audio_latent_from_file, encode_video |
packages/ltx-pipelines/src/ltx_pipelines/utils/types.py |
ModalitySpec, OffloadMode definitions |
packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py |
ModelPaths checkpoint resolution |
packages/ltx-core/src/ltx_core/conditioning/types/noise_mask_cond.py |
TemporalRegionMask implementation |
Summary
- RetakePipeline enables selective video region regeneration by time window, preserving content outside the mask
- TemporalRegionMask converts
start_time/end_timeto frame indices using source FPS - EXR sequences require
--frame-rate; standard videos must omit it regenerate_video/regenerate_audioflags independently control which modalities are modifieddistilled=True(default) uses fast 8‑step inference;distilled=Falseenables full guidance with configurable steps- All core logic resides in
retake.pywith validation inargs.pyand media I/O inmedia_io.py
Frequently Asked Questions
How do I regenerate only audio without touching video?
Set regenerate_video=False and regenerate_audio=True. In CLI: --regenerate-video false --regenerate-audio true. This freezes the video latent and applies diffusion only to the audio latent within the specified time window. The original video frames remain pixel‑identical.
Why does my EXR sequence fail with a frame‑rate error?
EXR folders lack embedded FPS metadata. The validation in _verify_media_path_args (lines 75‑84 of args.py) requires --frame-rate for EXR inputs and rejects it for containerized video files. Add --frame-rate 24 (or your target FPS) to resolve.
Can I use different step counts with the distilled model?
No. The distilled checkpoint is optimized for a fixed 8‑step sigma schedule (DISTILLED_SIGMAS). To control num_inference_steps, set distilled=False when constructing RetakePipeline. This activates GuidedDenoiser with MultiModalGuider for full conditional guidance.
What happens if start_time and end_time span the entire video?
The temporal mask covers all frames, effectively performing full video regeneration. However, regenerate_video=False or regenerate_audio=False will still freeze the respective modality. For true full regeneration, use ltx_pipelines.video_generation instead—RetakePipeline carries overhead for masking logic you don't need.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →