How the RetakePipeline Facilitates Selective Regeneration of Specific Time Regions in LTX-2
The RetakePipeline regenerates only a user-defined time interval of an existing video by applying temporal region masks that freeze frames outside the target window, enabling efficient, localized video editing without reprocessing the entire clip.
The RetakePipeline in Lightricks' LTX-2 is purpose-built for targeted video editing. Unlike full-video regeneration pipelines, it isolates a specific time window—defined by start_time and end_time—and runs diffusion only on that segment while preserving everything else. This architecture dramatically reduces computational cost and maintains visual continuity between edited and untouched frames.
Temporal Region Masking with TemporalRegionMask
At the core of selective regeneration is the TemporalRegionMask class, defined in [ltx_core/conditioning/types/noise_mask_cond.py](ltx_core/conditioning/types/noise_mask_cond.py). The pipeline wraps user-provided time boundaries into this mask and attaches it to modality specifications:
video_modality_spec = ModalitySpec(
context=v_context_p,
conditionings=[
TemporalRegionMask(
start_time=start_time,
end_time=end_time,
fps=output_shape.fps
)
] if regenerate_video else [],
initial_latent=initial_video_latent,
frozen=not regenerate_video,
)
This mask serves two critical functions in the retake pipeline:
- Frames inside the mask — Passed to the diffusion denoiser for regeneration
- Frames outside the mask — Marked with
frozen=True, causing the pipeline to skip them entirely
The mask is instantiated at lines 12–13 of [packages/ltx-pipelines/src/ltx_pipelines/retake.py](packages/ltx-pipelines/src/ltx_pipelines/retake.py) and propagated through the conditioning system.
Modality-Aware Diffusion Processing
The DiffusionStage receives both video_modality_spec and audio_modality_spec objects. During the denoising loop, it evaluates each modality's conditionings list:
- Latents carrying a
TemporalRegionMaskare routed through the full denoising process - Latents without masking—or with
frozen=True—are copied unchanged from the input latent
This design ensures computational cost scales with the length of the edited segment, not the entire video duration. A 5-second edit in a 60-second video requires only ~8% of the diffusion steps that full regeneration would demand.
Independent Video and Audio Control
The public __call__ signature at lines 64–66 of [retake.py](packages/ltx-pipelines/src/ltx_pipelines/retake.py) exposes two independent flags:
| Parameter | Purpose |
|---|---|
regenerate_video |
Enable/disable diffusion on the video track |
regenerate_audio |
Enable/disable diffusion on the audio track |
Both modalities can share the same TemporalRegionMask instance, or each can carry its own. This decoupling allows scenarios such as:
- Video-only retake — Regenerate visual content while preserving original audio
- Audio-only retake — Replace sound design without altering frames
- Synchronized retake — Update both modalities across the same time window
CLI and Python Usage Examples
Command-Line Retake
Regenerate the 10-second to 15-second segment using the retake CLI:
python -m ltx_pipelines.retake \
--video_path input.mp4 \
--prompt "A sunrise over mountains" \
--start_time 10 \
--end_time 15 \
--seed 42 \
--output_path output.mp4
The --start_time and --end_time arguments are parsed in [packages/ltx-pipelines/utils/args.py](packages/ltx-pipelines/utils/args.py) and validated before pipeline initialization.
Programmatic Retake in Python
For integration into larger workflows, instantiate RetakePipeline directly:
from ltx_pipelines.retake import RetakePipeline
from ltx_pipelines.utils.model_paths import ModelPaths
# Initialize with model paths and optional LoRAs
pipeline = RetakePipeline(
model_paths=ModelPaths(),
loras=[],
distilled=True, # Use distilled sigma schedule for faster inference
)
# Execute selective regeneration
video_iter, audio, tiling_cfg = pipeline(
video_path="input.mp4",
prompt="A sunrise over mountains",
start_time=10.0,
end_time=15.0,
seed=42,
regenerate_video=True,
regenerate_audio=False, # Preserve original audio track
)
# Encode final output using media utilities
from ltx_pipelines.utils.media_io import encode_video, get_videostream_metadata
source_meta = get_videostream_metadata("input.mp4")
encode_video(
video=video_iter,
fps=int(source_meta.fps),
audio=audio,
output_path="output.mp4",
video_chunks_number=1,
color_space="SRGB_LINEAR",
)
The ModalitySpec and OffloadMode types—defined in [packages/ltx-pipelines/utils/types.py](packages/ltx-pipelines/utils/types.py)—orchestrate how masking and freezing semantics are communicated between pipeline stages.
Key Source Files in the Retake Pipeline
| File | Role |
|---|---|
[packages/ltx-pipelines/src/ltx_pipelines/retake.py](packages/ltx-pipelines/src/ltx_pipelines/retake.py) |
Core RetakePipeline class, __call__ implementation, and CLI entry point |
[ltx_core/conditioning/types/noise_mask_cond.py](ltx_core/conditioning/types/noise_mask_cond.py) |
TemporalRegionMask definition for time-window isolation |
[packages/ltx-pipelines/utils/types.py](packages/ltx-pipelines/utils/types.py) |
ModalitySpec and OffloadMode type definitions |
[packages/ltx-pipelines/utils/args.py](packages/ltx-pipelines/utils/args.py) |
CLI argument parsing for --start_time and --end_time |
[packages/ltx-pipelines/utils/media_io.py](packages/ltx-pipelines/utils/media_io.py) |
Video metadata extraction and output encoding helpers |
Summary
- TemporalRegionMask isolates the target time window and marks surrounding frames as frozen
- ModalitySpec objects carry masking instructions separately for video and audio tracks
- Independent
regenerate_video/regenerate_audioflags enable selective editing of either modality - Frozen latents bypass diffusion entirely, reducing computation to proportional-to-segment cost
- Seamless blending is achieved by preserving original frames outside the masked region
Frequently Asked Questions
How does the RetakePipeline prevent changes outside the specified time region?
The pipeline constructs a TemporalRegionMask with start_time and end_time, then sets frozen=not regenerate_video on the modality spec. Frozen latents are never passed to the denoiser—instead, they are copied unchanged from the input. This ensures pixels outside the mask remain identical to the source video.
Can I regenerate audio without touching the video, or vice versa?
Yes. The __call__ method accepts independent boolean flags for each modality. Set regenerate_video=True, regenerate_audio=False to update only frames, or invert them for audio-only retake. Both modalities reference the same TemporalRegionMask by default but can be configured separately.
What file formats and color spaces does the retake pipeline support?
According to [media_io.py](packages/ltx-pipelines/utils/media_io.py), the pipeline reads standard video formats through get_videostream_metadata() and outputs to formats supported by encode_video(), including SRGB_LINEAR color space. EXR folder detection is also available for high-bit-depth workflows.
Where is the time window validation performed in the codebase?
Time boundaries are parsed as CLI arguments in [args.py](packages/ltx-pipelines/utils/args.py), then passed directly to TemporalRegionMask. The mask class validates that start_time < end_time and converts temporal coordinates to frame indices using the fps parameter, ensuring precise alignment with the latent timeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →