How DFRPipeline Handles Keyframe Rendering with Spatial Detailing and Temporal Upsampling

The DFRPipeline class in Lightricks/LTX-2 implements a two-stage diffusion pipeline that first generates keyframe slots at half resolution, then applies spatial detailing via upsampling and IC-LoRA conditioning, and optionally performs multi-round temporal upsampling with tiled anchor-keyframe refinement.

The DFRPipeline (Diffusion Fidelity Rendering) is the core video generation orchestrator in the Lightricks/LTX-2 repository. It transforms text prompts into high-fidelity video through a carefully structured three-phase process: base generation with keyframe slot creation, spatial detailing with resolution upscaling, and optional temporal upsampling for higher frame rates. All implementation details below are derived from packages/ltx-pipelines/src/ltx_pipelines/dfr_pipeline.py.

Phase 1: Base Generation and Keyframe Slot Creation

The pipeline begins by establishing the temporal canvas and generating initial keyframe slots at reduced resolution.

Canvas resolution and validation

In dfr_pipeline.py:L17-L34, the resolve_canvas helper determines the total frame count and computes keyframe positions (positions)—the specific pixel-frame indices where the model must emit latent slots. The pipeline then validates that the underlying diffusion stage supports generated keyframes via self.stage.assert_generated_keyframes_supported().

Half-resolution diffusion pass

The first diffusion stage runs with spatial dimensions halved:

stage_1_h = height // 2
stage_1_w = width // 2

The conditioning list includes VideoGeneratedKeyframeSlots with pixel_frame_indices=positions, instructing the model to output latent slots at the specified temporal positions. After stage_1_sigmas completes, the video_state object contains:

  • video_state.latent — the base half-resolution video latent
  • video_state.generated_keyframes — the slot latents extracted at keyframe positions

This phase is implemented in dfr_pipeline.py:L30-L34.

Phase 2: Spatial Detailing and Upscaling

The second phase transforms half-resolution outputs into full-resolution detailed video through spatial upsampling and IC-LoRA conditioning.

Keyframe and video upsampling

As shown in dfr_pipeline.py:L48-L50, the pipeline applies self.upsampler to both the slot keyframes and the reserved half-resolution video:

upsampled_slot_keyframes = self.upsampler(video_state.generated_keyframes)
upscaled_video_latent = self.upsampler(reserved_half_res_video)

The VideoUpsampler (defined in utils/blocks.py) performs the spatial upsampling operation, typically with a 2× factor.

Full-resolution diffusion with detailing conditioning

In dfr_pipeline.py:L62-L74, the pipeline constructs the Stage 2 conditioning set:

  1. Base image conditioning — for the full-resolution canvas dimensions
  2. Up-sampled keyframe slots — via VideoGeneratedKeyframeSlots(..., initial_keyframes=upsampled_slot_keyframes), reusing the spatially-upscaled latents from Stage 1
  3. Optional IC-LoRA detailing reference — if detailing_lora is provided, a VideoConditionByReferenceLatent injects the reserved half-resolution video latent (reserved_half_res_video), enabling pixel-spatial upsampling through the detailing LoRA

The self.stage_detailing diffusion stage then denoises the full-resolution latent while respecting these keyframe constraints, producing the final high-fidelity output.


# Conceptual structure of Stage 2 conditioning assembly

conditionings = [
    image_conditioning_full_res,
    VideoGeneratedKeyframeSlots(
        pixel_frame_indices=positions,
        initial_keyframes=upsampled_slot_keyframes,
    ),
]
if detailing_lora:
    conditionings.append(
        VideoConditionByReferenceLatent(
            latent=reserved_half_res_video,
            strength=detailing_strength,
        )
    )

Phase 3: Temporal Upsampling with Anchor Keyframes

When temporal_upsample_rounds > 0, the pipeline enters a multi-round temporal refinement loop that doubles the frame count each iteration.

Loop structure and temporal upsampling

The entry point at dfr_pipeline.py:L84-L88 validates configuration and begins iterating:

for round_idx in range(1, temporal_upsample_rounds + 1):
    # Each round: 2× frame count

The core temporal upsampling logic in dfr_pipeline.py:L102-L138 uses self.temporal_upsampler to interpolate the video latent, transforming frame count T → 2T-1 (effectively doubling).

Tiled processing with anchor and slot keyframes

For temporal coherence at high frame counts, the pipeline employs tiled diffusion:

  1. Canvas tiling: The temporal axis is divided into 2**round_idx tiles via tile_ranges (see dfr_layout.py for coordinate transformation helpers)

  2. Per-tile conditioning (dfr_pipeline.py:L160-L242):

    • Image conditioning for the tile's temporal extent
    • Anchor keyframes — keyframes from previous rounds re-attached with strength _ANCHOR_KEYFRAME_STRENGTH (0.95) via _keyframe_conditionings_from_latents
    • Slot keyframes — newly created VideoGeneratedKeyframeSlots for mid-segment positions within each tile
  3. Diffusion and stitching: Each tile processes independently with EulerAncestralDiffusionStep, using a unique seed per tile for diversity. Results are merged via stitch_tile_latents.

Keyframe carry-forward mechanism

After each round, _merge_carry_forward_keyframes (in dfr_pipeline.py:L27-L41) rebuilds the anchor bag by merging:

  • Previous-round anchors (now down-weighted)
  • Newly generated slot keyframes from current-round tiles

This preserves temporal consistency across rounds while allowing new detail insertion.

Final frame count formula

The temporal upsampling produces frame counts following:


final_frames = (requested_frames - 1) * 2**rounds + 1

With temporal_upsample_rounds=2, a 25-frame request becomes 97 frames (effectively 4× temporal resolution).

Complete Implementation Example

Below is a runnable configuration demonstrating keyframe rendering, spatial detailing, and temporal upsampling:

from ltx_pipelines.dfr_pipeline import DFRPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Configure model checkpoints

paths = ModelPaths(
    transformer="path/to/transformer.ckpt",
    video_vae="path/to/video_vae.ckpt",
    audio_vae="path/to/audio_vae.ckpt",
    duration_head_path="path/to/duration_head.ckpt",
)

# Initialize pipeline with all enhancement paths

pipeline = DFRPipeline(
    model_paths=paths,
    distilled_lora=[],
    spatial_upsampler_path="path/to/spatial_up.ckpt",
    loras=[],
    detailing_lora=[("path/to/detailing_lora.ckpt", 1.0)],  # IC-LoRA spatial detailing

    temporal_upsampler_path="path/to/temporal_up.ckpt",   # Temporal upsampler

)

# Execute generation with temporal refinement

video, audio, num_frames, tiling_cfg = pipeline(
    prompt="A sunrise over a mountain lake",
    seed=42,
    height=720,
    width=1280,
    frame_rate=30.0,
    images=[],
    temporal_upsample_rounds=2,  # Two rounds = 4× frame rate

)

The resulting video tensor has shape (B, C, T, H, W) where T reflects the temporally-upsampled frame count.

Key Architectural Components

Component Source File Purpose
DFRPipeline dfr_pipeline.py Main orchestrator for two-stage generation and upsampling
VideoUpsampler utils/blocks.py Spatial 2× upsampling of latents and keyframes
VideoGeneratedKeyframeSlots Conditioning type Declares temporal positions for forced keyframe generation
VideoConditionByReferenceLatent Conditioning type Injects reference latents for IC-LoRA detailing
_keyframe_conditionings_from_latents dfr_pipeline.py Re-attaches anchor keyframes with controlled strength
stitch_tile_latents dfr_layout.py Merges independently processed tiles into coherent video
EulerAncestralDiffusionStep Sampling module Per-tile diffusion with ancestral noise

Summary

  • DFRPipeline implements a two-stage diffusion architecture: Stage 1 generates keyframe slots at half resolution; Stage 2 refines with full-resolution spatial detailing.
  • Spatial detailing combines VideoUpsampler for resolution scaling and optional IC-LoRA conditioning via VideoConditionByReferenceLatent for pixel-level fidelity enhancement.
  • Temporal upsampling operates in configurable rounds, each doubling frame count through temporal_upsampler and tiled diffusion with anchor-keyframe preservation.
  • Keyframe anchoring at strength 0.95 (_ANCHOR_KEYFRAME_STRENGTH) ensures temporal consistency across tiles and rounds.
  • All pipeline logic resides in packages/ltx-pipelines/src/ltx_pipelines/dfr_pipeline.py with layout utilities in dfr_layout.py and block definitions in utils/blocks.py.

Frequently Asked Questions

What is the purpose of keyframe slots in DFRPipeline?

Keyframe slots are forced latent outputs at specific temporal positions that the pipeline extracts during Stage 1 generation. According to the source code in dfr_pipeline.py:L30-L34, these generated_keyframes serve as anchor points for both spatial upsampling (Phase 2) and temporal refinement (Phase 3), ensuring structural consistency across resolution and frame-rate changes.

How does spatial detailing differ from basic upsampling?

Basic upsampling in dfr_pipeline.py:L48-L50 uses the VideoUpsampler for pure resolution scaling. Spatial detailing (lines 62-74) adds an IC-LoRA conditioning path via VideoConditionByReferenceLatent that injects the half-resolution source latent as a reference, enabling the detailing LoRA to hallucinate fine details rather than merely interpolating pixels.

Why does temporal upsampling use tiled processing?

Tiled processing in dfr_pipeline.py:L160-L242 addresses computational and memory constraints when generating high frame counts. By splitting the temporal axis into 2**round_idx tiles and processing each with independent seeds, the pipeline achieves diversity while the anchor keyframe mechanism (strength 0.95) maintains cross-tile coherence. The stitch_tile_latents function then merges results without visible boundaries.

What frame rate multiplier does temporal upsampling provide?

Each temporal upsampling round doubles the effective frame count according to the formula (requested_frames - 1) * 2**rounds + 1 (see dfr_pipeline.py:L102-L138). With temporal_upsample_rounds=2, a 30 fps source becomes approximately 120 fps effective output. The actual playback rate depends on the frame_rate parameter passed to the pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →