How Multi-GPU Tiled Data Parallel and Sequence Parallel Work in the LTX-2 Framework
LTX-2 scales video generation across multiple GPUs through two complementary strategies: Tiled Data Parallel (TDP) distributes spatial tiles across devices, while Sequence Parallel (SP) splits the temporal frame dimension with custom attention kernels.
The LTX-2 framework supports distributed transformer inference through two distinct data-parallel strategies. Understanding how multi-GPU tiled data parallel and sequence parallel work is essential for optimizing memory usage and throughput when generating high-resolution or long-duration videos. Both approaches wrap the underlying transformer model at build time and handle all GPU communication transparently through NCCL.
Tiled Data Parallel: Spatial Distribution
Tiled Data Parallel distributes spatial tiles of the video latent tensor across GPUs. Each device processes a distinct region of the height × width dimensions, reducing per-GPU memory proportional to the number of tiles.
Builder Architecture
The TiledDataParallelBuilder in ltx_pipelines/multigpu/tdp_builder.py conforms to the ModelBuilderProtocol and injects tile-aware operations into the model construction pipeline:
class TiledDataParallelBuilder(...):
def __init__(self, inner, group, tiling, registry, tracker, normalize_positions=True):
cuda_device = torch.device(f"cuda:{torch.cuda.current_device()}")
self._inner = inner.with_registry(registry).with_lora_load_device(cuda_device)
self._tracker = tracker
self._group = group
self._tiling = tiling
self._normalize_positions = normalize_positions
The builder's build() method constructs the base transformer through the weight tracker, then wraps it with TiledDataParallelModelWrapper:
model = self._tracker.build(self._inner, device=device, dtype=dtype, **_kwargs)
return TiledDataParallelModelWrapper(
model,
video_tools=video_tools,
tiling=self._tiling,
group=self._group,
normalize_positions=self._normalize_positions,
)
Runtime Execution in TiledDataParallelModelWrapper
The wrapper class in ltx_core/multigpu/transformer/tiled_data_parallel.py performs three critical operations at runtime:
- Tile calculation —
VideoLatentTools.tile_for_rank(rank, tiling)determines which spatial region belongs to each GPU - Scatter/Gather — the full latent is scattered before forward pass; tiles are gathered after denoising completes
- NCCL synchronization — all communication uses the supplied
dist.ProcessGroupfor deterministic all-to-all operations
Pipeline Integration
In ti2vid_two_stages_mgpu.py (line 111), the TDP builder replaces the stage 2 transformer:
pipeline.stage_2._transformer_builder = TiledDataParallelBuilder(
inner=builder,
group=process_group,
tiling=tile_config,
registry=registry,
tracker=weight_tracker,
)
This configuration leaves stage 1 on a single GPU while scaling stage 2 across multiple devices via spatial tiling.
Sequence Parallel: Temporal Distribution
Sequence Parallel distributes the temporal sequence (frames) across GPUs, enabling processing of longer videos than single-device memory allows. Unlike TDP, SP requires custom attention kernels that operate on distributed sequences.
Builder with Custom Attention Ops
The SequenceParallelBuilder in ltx_pipelines/multigpu/sp_builder.py injects specialized self-attention operations:
class SequenceParallelBuilder(...):
def __init__(self, inner, attn_mgr, registry, tracker):
sp_ops = create_video_self_attention_module_ops(attn_mgr)
self._inner = inner.with_registry(registry).with_lora_load_device(cuda_device)\
.with_module_ops((*inner.module_ops, sp_ops))
self._tracker = tracker
self._attn_mgr = attn_mgr
The build step wraps the constructed model similarly:
model = self._tracker.build(self._inner, device=device, dtype=dtype, **kwargs)
return SequenceParallelModelWrapper(model, self._attn_mgr)
Runtime Execution in SequenceParallelModelWrapper
Located in ltx_core/multigpu/transformer/sequence_parallel.py, this wrapper:
- Splits the temporal axis — each rank receives
seq = latent[:, :, start:end, ...], a contiguous frame slice - Uses custom attention kernels —
create_video_self_attention_module_opsreplaces default self-attention, enabling global attention across ranks via NCCL all-to-all - Enforces communication timeouts —
AttentionManagerexposesall2all_timeout_secondsfor configurable synchronization boundaries
Pipeline Usage
Sequence parallel integrates identically to TDP by swapping the builder:
pipeline.stage_2._transformer_builder = SequenceParallelBuilder(
inner=builder,
attn_mgr=attention_manager,
registry=registry,
tracker=weight_tracker,
)
Shared Infrastructure: Weight Tracking and GPU Residency
Both builders rely on TransformerWeightTracker (ltx_pipelines/multigpu/weight_tracker.py) to maintain GPU-resident parameters across model rebuilds. This eliminates redundant loading and explains why keeps_gpu_resident_weights returns True in both tdp_builder.py (lines 55-56) and sp_builder.py (lines 45-46).
The tracker operates as a shared cache: weights are loaded once and rebound to new model instances as pipeline configurations change.
Comparing the Two Parallelism Strategies
| Aspect | Tiled Data Parallel | Sequence Parallel |
|---|---|---|
| Parallelized dimension | Spatial (H × W tiles) | Temporal (frame sequence) |
| Core wrapper | TiledDataParallelModelWrapper |
SequenceParallelModelWrapper |
| Custom kernels required | No | Yes (attention ops) |
| Communication pattern | Scatter/Gather per-step | All-to-all within attention |
| Best for | High-resolution frames | Long-duration videos |
Interaction and Composition
A single transformer uses either TDP or SP, not both simultaneously. The pipeline selects the strategy by configuring the appropriate builder for stage 2.
Advanced users can compose both approaches spatially and temporally—tiling frames first, then applying sequence parallelism within each tile—though this requires a custom wrapper. The foundational components (TileCountConfig, AttentionManager, both wrapper classes) support such extensions.
Complete Example: Tiled Data Parallel Setup
import torch
import torch.distributed as dist
from ltx_pipelines.ti2vid_two_stages_mgpu import TI2VidTwoStagesMGPU
from ltx_pipelines.multigpu.tdp_builder import TiledDataParallelBuilder
from ltx_pipelines.multigpu.weight_tracker import TransformerWeightTracker
from ltx_core.tiling import TileCountConfig
from ltx_core.utils import VideoLatentTools
# Initialize distributed world
dist.init_process_group(backend="nccl")
rank = dist.get_rank()
group = dist.group.WORLD
# Configure 2×2 spatial tiling
tiling = TileCountConfig(num_tiles_h=2, num_tiles_w=2)
# GPU-resident weight tracking
tracker = TransformerWeightTracker()
# Build tiled-parallel wrapper
tdp_builder = TiledDataParallelBuilder(
inner=single_gpu_builder,
group=group,
tiling=tiling,
registry=registry,
tracker=tracker,
)
# Inject into pipeline stage 2
pipeline = TI2VidTwoStagesMGPU(...)
pipeline.stage_2._transformer_builder = tdp_builder
# Run with automatic scatter/gather
video, audio = pipeline(prompt, video_tools=VideoLatentTools(rank, tiling))
Complete Example: Sequence Parallel Setup
from ltx_pipelines.multigpu.sp_builder import SequenceParallelBuilder
from ltx_core.multigpu.transformer.attention import AttentionManager
# Attention manager with 30-second all-to-all timeout
attn_mgr = AttentionManager(group=group, timeout=30.0)
sp_builder = SequenceParallelBuilder(
inner=single_gpu_builder,
attn_mgr=attn_mgr,
registry=registry,
tracker=tracker,
)
pipeline.stage_2._transformer_builder = sp_builder
# Inference proceeds identically; temporal splitting is transparent
Key Source Files
| File | Purpose |
|---|---|
ltx_pipelines/multigpu/tdp_builder.py |
Constructs TiledDataParallelModelWrapper |
ltx_pipelines/multigpu/sp_builder.py |
Constructs SequenceParallelModelWrapper |
ltx_core/multigpu/transformer/tiled_data_parallel.py |
Spatial tile scatter/gather logic |
ltx_core/multigpu/transformer/sequence_parallel.py |
Temporal split and custom attention |
ltx_pipelines/multigpu/weight_tracker.py |
GPU-resident parameter caching |
ltx_pipelines/ti2vid_two_stages_mgpu.py |
Example MGPU pipeline integration |
Summary
-
Tiled Data Parallel in LTX-2 splits video latents spatially using
TiledDataParallelBuilderandTiledDataParallelModelWrapper, with automatic scatter/gather via the suppliedProcessGroup. -
Sequence Parallel splits the temporal dimension using
SequenceParallelBuilderandSequenceParallelModelWrapper, requiring custom attention kernels fromcreate_video_self_attention_module_ops. -
Both strategies use
TransformerWeightTrackerto keep model weights GPU-resident across builds, avoiding reload overhead. -
The pipeline selects parallelism by assigning the appropriate builder to
stage_2._transformer_builderin multi-GPU pipeline scripts. -
NCCL provides all inter-GPU synchronization; both wrappers accept a
dist.ProcessGroupfor deterministic communication ordering.
Frequently Asked Questions
Can I use Tiled Data Parallel and Sequence Parallel together on the same model?
No—a single transformer instance wraps with either TiledDataParallelModelWrapper or SequenceParallelModelWrapper. However, the underlying components (TileCountConfig, AttentionManager, wrapper classes) support building a custom combined wrapper that tiles spatially then applies sequence parallelism within each tile.
How does LTX-2 handle synchronization between GPUs?
Both wrappers use the dist.ProcessGroup passed to their builders. TDP performs scatter/gather operations around the forward pass, while SP uses all-to-all communication within custom attention kernels. The AttentionManager in SP exposes all2all_timeout_seconds for fine-grained timeout control.
What determines whether to choose TDP or Sequence Parallel?
Choose Tiled Data Parallel when generating high-resolution frames where spatial dimensions exceed single-GPU memory. Choose Sequence Parallel when processing long videos where the temporal dimension dominates memory usage. For extremely large generations, consider composing both approaches.
Where are the model weights stored during multi-GPU inference?
TransformerWeightTracker keeps weights GPU-resident across builds, as confirmed by keeps_gpu_resident_weights returning True in both builders. This avoids repeated CPU-to-GPU transfers when pipeline configurations change.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →