How Multi-GPU Tiled Data Parallel and Sequence Parallel Work in the LTX-2 Framework

LTX-2 scales video generation across multiple GPUs through two complementary strategies: Tiled Data Parallel (TDP) distributes spatial tiles across devices, while Sequence Parallel (SP) splits the temporal frame dimension with custom attention kernels.

The LTX-2 framework supports distributed transformer inference through two distinct data-parallel strategies. Understanding how multi-GPU tiled data parallel and sequence parallel work is essential for optimizing memory usage and throughput when generating high-resolution or long-duration videos. Both approaches wrap the underlying transformer model at build time and handle all GPU communication transparently through NCCL.

Tiled Data Parallel: Spatial Distribution

Tiled Data Parallel distributes spatial tiles of the video latent tensor across GPUs. Each device processes a distinct region of the height × width dimensions, reducing per-GPU memory proportional to the number of tiles.

Builder Architecture

The TiledDataParallelBuilder in ltx_pipelines/multigpu/tdp_builder.py conforms to the ModelBuilderProtocol and injects tile-aware operations into the model construction pipeline:

class TiledDataParallelBuilder(...):
    def __init__(self, inner, group, tiling, registry, tracker, normalize_positions=True):
        cuda_device = torch.device(f"cuda:{torch.cuda.current_device()}")
        self._inner = inner.with_registry(registry).with_lora_load_device(cuda_device)
        self._tracker = tracker
        self._group = group
        self._tiling = tiling
        self._normalize_positions = normalize_positions

The builder's build() method constructs the base transformer through the weight tracker, then wraps it with TiledDataParallelModelWrapper:

model = self._tracker.build(self._inner, device=device, dtype=dtype, **_kwargs)
return TiledDataParallelModelWrapper(
    model,
    video_tools=video_tools,
    tiling=self._tiling,
    group=self._group,
    normalize_positions=self._normalize_positions,
)

Runtime Execution in TiledDataParallelModelWrapper

The wrapper class in ltx_core/multigpu/transformer/tiled_data_parallel.py performs three critical operations at runtime:

  1. Tile calculation — VideoLatentTools.tile_for_rank(rank, tiling) determines which spatial region belongs to each GPU
  2. Scatter/Gather — the full latent is scattered before forward pass; tiles are gathered after denoising completes
  3. NCCL synchronization — all communication uses the supplied dist.ProcessGroup for deterministic all-to-all operations

Pipeline Integration

In ti2vid_two_stages_mgpu.py (line 111), the TDP builder replaces the stage 2 transformer:

pipeline.stage_2._transformer_builder = TiledDataParallelBuilder(
    inner=builder,
    group=process_group,
    tiling=tile_config,
    registry=registry,
    tracker=weight_tracker,
)

This configuration leaves stage 1 on a single GPU while scaling stage 2 across multiple devices via spatial tiling.

Sequence Parallel: Temporal Distribution

Sequence Parallel distributes the temporal sequence (frames) across GPUs, enabling processing of longer videos than single-device memory allows. Unlike TDP, SP requires custom attention kernels that operate on distributed sequences.

Builder with Custom Attention Ops

The SequenceParallelBuilder in ltx_pipelines/multigpu/sp_builder.py injects specialized self-attention operations:

class SequenceParallelBuilder(...):
    def __init__(self, inner, attn_mgr, registry, tracker):
        sp_ops = create_video_self_attention_module_ops(attn_mgr)
        self._inner = inner.with_registry(registry).with_lora_load_device(cuda_device)\
                       .with_module_ops((*inner.module_ops, sp_ops))
        self._tracker = tracker
        self._attn_mgr = attn_mgr

The build step wraps the constructed model similarly:

model = self._tracker.build(self._inner, device=device, dtype=dtype, **kwargs)
return SequenceParallelModelWrapper(model, self._attn_mgr)

Runtime Execution in SequenceParallelModelWrapper

Located in ltx_core/multigpu/transformer/sequence_parallel.py, this wrapper:

  1. Splits the temporal axis — each rank receives seq = latent[:, :, start:end, ...], a contiguous frame slice
  2. Uses custom attention kernels — create_video_self_attention_module_ops replaces default self-attention, enabling global attention across ranks via NCCL all-to-all
  3. Enforces communication timeouts — AttentionManager exposes all2all_timeout_seconds for configurable synchronization boundaries

Pipeline Usage

Sequence parallel integrates identically to TDP by swapping the builder:

pipeline.stage_2._transformer_builder = SequenceParallelBuilder(
    inner=builder,
    attn_mgr=attention_manager,
    registry=registry,
    tracker=weight_tracker,
)

Shared Infrastructure: Weight Tracking and GPU Residency

Both builders rely on TransformerWeightTracker (ltx_pipelines/multigpu/weight_tracker.py) to maintain GPU-resident parameters across model rebuilds. This eliminates redundant loading and explains why keeps_gpu_resident_weights returns True in both tdp_builder.py (lines 55-56) and sp_builder.py (lines 45-46).

The tracker operates as a shared cache: weights are loaded once and rebound to new model instances as pipeline configurations change.

Comparing the Two Parallelism Strategies

Aspect Tiled Data Parallel Sequence Parallel
Parallelized dimension Spatial (H × W tiles) Temporal (frame sequence)
Core wrapper TiledDataParallelModelWrapper SequenceParallelModelWrapper
Custom kernels required No Yes (attention ops)
Communication pattern Scatter/Gather per-step All-to-all within attention
Best for High-resolution frames Long-duration videos

Interaction and Composition

A single transformer uses either TDP or SP, not both simultaneously. The pipeline selects the strategy by configuring the appropriate builder for stage 2.

Advanced users can compose both approaches spatially and temporally—tiling frames first, then applying sequence parallelism within each tile—though this requires a custom wrapper. The foundational components (TileCountConfig, AttentionManager, both wrapper classes) support such extensions.

Complete Example: Tiled Data Parallel Setup

import torch
import torch.distributed as dist
from ltx_pipelines.ti2vid_two_stages_mgpu import TI2VidTwoStagesMGPU
from ltx_pipelines.multigpu.tdp_builder import TiledDataParallelBuilder
from ltx_pipelines.multigpu.weight_tracker import TransformerWeightTracker
from ltx_core.tiling import TileCountConfig
from ltx_core.utils import VideoLatentTools

# Initialize distributed world

dist.init_process_group(backend="nccl")
rank = dist.get_rank()
group = dist.group.WORLD

# Configure 2×2 spatial tiling

tiling = TileCountConfig(num_tiles_h=2, num_tiles_w=2)

# GPU-resident weight tracking

tracker = TransformerWeightTracker()

# Build tiled-parallel wrapper

tdp_builder = TiledDataParallelBuilder(
    inner=single_gpu_builder,
    group=group,
    tiling=tiling,
    registry=registry,
    tracker=tracker,
)

# Inject into pipeline stage 2

pipeline = TI2VidTwoStagesMGPU(...)
pipeline.stage_2._transformer_builder = tdp_builder

# Run with automatic scatter/gather

video, audio = pipeline(prompt, video_tools=VideoLatentTools(rank, tiling))

Complete Example: Sequence Parallel Setup

from ltx_pipelines.multigpu.sp_builder import SequenceParallelBuilder
from ltx_core.multigpu.transformer.attention import AttentionManager

# Attention manager with 30-second all-to-all timeout

attn_mgr = AttentionManager(group=group, timeout=30.0)

sp_builder = SequenceParallelBuilder(
    inner=single_gpu_builder,
    attn_mgr=attn_mgr,
    registry=registry,
    tracker=tracker,
)

pipeline.stage_2._transformer_builder = sp_builder

# Inference proceeds identically; temporal splitting is transparent

Key Source Files

File Purpose
ltx_pipelines/multigpu/tdp_builder.py Constructs TiledDataParallelModelWrapper
ltx_pipelines/multigpu/sp_builder.py Constructs SequenceParallelModelWrapper
ltx_core/multigpu/transformer/tiled_data_parallel.py Spatial tile scatter/gather logic
ltx_core/multigpu/transformer/sequence_parallel.py Temporal split and custom attention
ltx_pipelines/multigpu/weight_tracker.py GPU-resident parameter caching
ltx_pipelines/ti2vid_two_stages_mgpu.py Example MGPU pipeline integration

Summary

  • Tiled Data Parallel in LTX-2 splits video latents spatially using TiledDataParallelBuilder and TiledDataParallelModelWrapper, with automatic scatter/gather via the supplied ProcessGroup.

  • Sequence Parallel splits the temporal dimension using SequenceParallelBuilder and SequenceParallelModelWrapper, requiring custom attention kernels from create_video_self_attention_module_ops.

  • Both strategies use TransformerWeightTracker to keep model weights GPU-resident across builds, avoiding reload overhead.

  • The pipeline selects parallelism by assigning the appropriate builder to stage_2._transformer_builder in multi-GPU pipeline scripts.

  • NCCL provides all inter-GPU synchronization; both wrappers accept a dist.ProcessGroup for deterministic communication ordering.

Frequently Asked Questions

Can I use Tiled Data Parallel and Sequence Parallel together on the same model?

No—a single transformer instance wraps with either TiledDataParallelModelWrapper or SequenceParallelModelWrapper. However, the underlying components (TileCountConfig, AttentionManager, wrapper classes) support building a custom combined wrapper that tiles spatially then applies sequence parallelism within each tile.

How does LTX-2 handle synchronization between GPUs?

Both wrappers use the dist.ProcessGroup passed to their builders. TDP performs scatter/gather operations around the forward pass, while SP uses all-to-all communication within custom attention kernels. The AttentionManager in SP exposes all2all_timeout_seconds for fine-grained timeout control.

What determines whether to choose TDP or Sequence Parallel?

Choose Tiled Data Parallel when generating high-resolution frames where spatial dimensions exceed single-GPU memory. Choose Sequence Parallel when processing long videos where the temporal dimension dominates memory usage. For extremely large generations, consider composing both approaches.

Where are the model weights stored during multi-GPU inference?

TransformerWeightTracker keeps weights GPU-resident across builds, as confirmed by keeps_gpu_resident_weights returning True in both builders. This avoids repeated CPU-to-GPU transfers when pipeline configurations change.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →