How the RetakePipeline Facilitates Selective Regeneration of Specific Time Regions in LTX-2

The RetakePipeline regenerates only a user-defined time interval of an existing video by applying temporal region masks that freeze frames outside the target window, enabling efficient, localized video editing without reprocessing the entire clip.

The RetakePipeline in Lightricks' LTX-2 is purpose-built for targeted video editing. Unlike full-video regeneration pipelines, it isolates a specific time window—defined by start_time and end_time—and runs diffusion only on that segment while preserving everything else. This architecture dramatically reduces computational cost and maintains visual continuity between edited and untouched frames.

Temporal Region Masking with TemporalRegionMask

At the core of selective regeneration is the TemporalRegionMask class, defined in [ltx_core/conditioning/types/noise_mask_cond.py](ltx_core/conditioning/types/noise_mask_cond.py). The pipeline wraps user-provided time boundaries into this mask and attaches it to modality specifications:

video_modality_spec = ModalitySpec(
    context=v_context_p,
    conditionings=[
        TemporalRegionMask(
            start_time=start_time,
            end_time=end_time,
            fps=output_shape.fps
        )
    ] if regenerate_video else [],
    initial_latent=initial_video_latent,
    frozen=not regenerate_video,
)

This mask serves two critical functions in the retake pipeline:

  • Frames inside the mask — Passed to the diffusion denoiser for regeneration
  • Frames outside the mask — Marked with frozen=True, causing the pipeline to skip them entirely

The mask is instantiated at lines 12–13 of [packages/ltx-pipelines/src/ltx_pipelines/retake.py](packages/ltx-pipelines/src/ltx_pipelines/retake.py) and propagated through the conditioning system.

Modality-Aware Diffusion Processing

The DiffusionStage receives both video_modality_spec and audio_modality_spec objects. During the denoising loop, it evaluates each modality's conditionings list:

  • Latents carrying a TemporalRegionMask are routed through the full denoising process
  • Latents without masking—or with frozen=True—are copied unchanged from the input latent

This design ensures computational cost scales with the length of the edited segment, not the entire video duration. A 5-second edit in a 60-second video requires only ~8% of the diffusion steps that full regeneration would demand.

Independent Video and Audio Control

The public __call__ signature at lines 64–66 of [retake.py](packages/ltx-pipelines/src/ltx_pipelines/retake.py) exposes two independent flags:

Parameter Purpose
regenerate_video Enable/disable diffusion on the video track
regenerate_audio Enable/disable diffusion on the audio track

Both modalities can share the same TemporalRegionMask instance, or each can carry its own. This decoupling allows scenarios such as:

  • Video-only retake — Regenerate visual content while preserving original audio
  • Audio-only retake — Replace sound design without altering frames
  • Synchronized retake — Update both modalities across the same time window

CLI and Python Usage Examples

Command-Line Retake

Regenerate the 10-second to 15-second segment using the retake CLI:

python -m ltx_pipelines.retake \
    --video_path input.mp4 \
    --prompt "A sunrise over mountains" \
    --start_time 10 \
    --end_time 15 \
    --seed 42 \
    --output_path output.mp4

The --start_time and --end_time arguments are parsed in [packages/ltx-pipelines/utils/args.py](packages/ltx-pipelines/utils/args.py) and validated before pipeline initialization.

Programmatic Retake in Python

For integration into larger workflows, instantiate RetakePipeline directly:

from ltx_pipelines.retake import RetakePipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Initialize with model paths and optional LoRAs

pipeline = RetakePipeline(
    model_paths=ModelPaths(),
    loras=[],
    distilled=True,  # Use distilled sigma schedule for faster inference

)

# Execute selective regeneration

video_iter, audio, tiling_cfg = pipeline(
    video_path="input.mp4",
    prompt="A sunrise over mountains",
    start_time=10.0,
    end_time=15.0,
    seed=42,
    regenerate_video=True,
    regenerate_audio=False,  # Preserve original audio track

)

# Encode final output using media utilities

from ltx_pipelines.utils.media_io import encode_video, get_videostream_metadata

source_meta = get_videostream_metadata("input.mp4")
encode_video(
    video=video_iter,
    fps=int(source_meta.fps),
    audio=audio,
    output_path="output.mp4",
    video_chunks_number=1,
    color_space="SRGB_LINEAR",
)

The ModalitySpec and OffloadMode types—defined in [packages/ltx-pipelines/utils/types.py](packages/ltx-pipelines/utils/types.py)—orchestrate how masking and freezing semantics are communicated between pipeline stages.

Key Source Files in the Retake Pipeline

File Role
[packages/ltx-pipelines/src/ltx_pipelines/retake.py](packages/ltx-pipelines/src/ltx_pipelines/retake.py) Core RetakePipeline class, __call__ implementation, and CLI entry point
[ltx_core/conditioning/types/noise_mask_cond.py](ltx_core/conditioning/types/noise_mask_cond.py) TemporalRegionMask definition for time-window isolation
[packages/ltx-pipelines/utils/types.py](packages/ltx-pipelines/utils/types.py) ModalitySpec and OffloadMode type definitions
[packages/ltx-pipelines/utils/args.py](packages/ltx-pipelines/utils/args.py) CLI argument parsing for --start_time and --end_time
[packages/ltx-pipelines/utils/media_io.py](packages/ltx-pipelines/utils/media_io.py) Video metadata extraction and output encoding helpers

Summary

  • TemporalRegionMask isolates the target time window and marks surrounding frames as frozen
  • ModalitySpec objects carry masking instructions separately for video and audio tracks
  • Independent regenerate_video/regenerate_audio flags enable selective editing of either modality
  • Frozen latents bypass diffusion entirely, reducing computation to proportional-to-segment cost
  • Seamless blending is achieved by preserving original frames outside the masked region

Frequently Asked Questions

How does the RetakePipeline prevent changes outside the specified time region?

The pipeline constructs a TemporalRegionMask with start_time and end_time, then sets frozen=not regenerate_video on the modality spec. Frozen latents are never passed to the denoiser—instead, they are copied unchanged from the input. This ensures pixels outside the mask remain identical to the source video.

Can I regenerate audio without touching the video, or vice versa?

Yes. The __call__ method accepts independent boolean flags for each modality. Set regenerate_video=True, regenerate_audio=False to update only frames, or invert them for audio-only retake. Both modalities reference the same TemporalRegionMask by default but can be configured separately.

What file formats and color spaces does the retake pipeline support?

According to [media_io.py](packages/ltx-pipelines/utils/media_io.py), the pipeline reads standard video formats through get_videostream_metadata() and outputs to formats supported by encode_video(), including SRGB_LINEAR color space. EXR folder detection is also available for high-bit-depth workflows.

Where is the time window validation performed in the codebase?

Time boundaries are parsed as CLI arguments in [args.py](packages/ltx-pipelines/utils/args.py), then passed directly to TemporalRegionMask. The mask class validates that start_time < end_time and converts temporal coordinates to frame indices using the fps parameter, ensuring precise alignment with the latent timeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →