# How Multi-GPU Tiled Data Parallel and Sequence Parallel Work in the LTX-2 Framework

> Understand how LTX-2 uses multi-GPU Tiled Data Parallel and Sequence Parallel for efficient video generation by distributing spatial tiles and temporal frames across devices.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: internals
- Published: 2026-08-18

---

**LTX-2 scales video generation across multiple GPUs through two complementary strategies: Tiled Data Parallel (TDP) distributes spatial tiles across devices, while Sequence Parallel (SP) splits the temporal frame dimension with custom attention kernels.**

The LTX-2 framework supports distributed transformer inference through two distinct data-parallel strategies. Understanding how **multi-GPU tiled data parallel and sequence parallel** work is essential for optimizing memory usage and throughput when generating high-resolution or long-duration videos. Both approaches wrap the underlying transformer model at build time and handle all GPU communication transparently through NCCL.

## Tiled Data Parallel: Spatial Distribution

Tiled Data Parallel distributes **spatial tiles** of the video latent tensor across GPUs. Each device processes a distinct region of the height × width dimensions, reducing per-GPU memory proportional to the number of tiles.

### Builder Architecture

The `TiledDataParallelBuilder` in [`ltx_pipelines/multigpu/tdp_builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/multigpu/tdp_builder.py) conforms to the `ModelBuilderProtocol` and injects tile-aware operations into the model construction pipeline:

```python
class TiledDataParallelBuilder(...):
    def __init__(self, inner, group, tiling, registry, tracker, normalize_positions=True):
        cuda_device = torch.device(f"cuda:{torch.cuda.current_device()}")
        self._inner = inner.with_registry(registry).with_lora_load_device(cuda_device)
        self._tracker = tracker
        self._group = group
        self._tiling = tiling
        self._normalize_positions = normalize_positions

```

The builder's `build()` method constructs the base transformer through the weight tracker, then wraps it with `TiledDataParallelModelWrapper`:

```python
model = self._tracker.build(self._inner, device=device, dtype=dtype, **_kwargs)
return TiledDataParallelModelWrapper(
    model,
    video_tools=video_tools,
    tiling=self._tiling,
    group=self._group,
    normalize_positions=self._normalize_positions,
)

```

### Runtime Execution in TiledDataParallelModelWrapper

The wrapper class in [`ltx_core/multigpu/transformer/tiled_data_parallel.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/multigpu/transformer/tiled_data_parallel.py) performs three critical operations at runtime:

1. **Tile calculation** — `VideoLatentTools.tile_for_rank(rank, tiling)` determines which spatial region belongs to each GPU
2. **Scatter/Gather** — the full latent is scattered before forward pass; tiles are gathered after denoising completes
3. **NCCL synchronization** — all communication uses the supplied `dist.ProcessGroup` for deterministic all-to-all operations

### Pipeline Integration

In [`ti2vid_two_stages_mgpu.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages_mgpu.py) (line 111), the TDP builder replaces the stage 2 transformer:

```python
pipeline.stage_2._transformer_builder = TiledDataParallelBuilder(
    inner=builder,
    group=process_group,
    tiling=tile_config,
    registry=registry,
    tracker=weight_tracker,
)

```

This configuration leaves stage 1 on a single GPU while scaling stage 2 across multiple devices via spatial tiling.

## Sequence Parallel: Temporal Distribution

Sequence Parallel distributes the **temporal sequence** (frames) across GPUs, enabling processing of longer videos than single-device memory allows. Unlike TDP, SP requires custom attention kernels that operate on distributed sequences.

### Builder with Custom Attention Ops

The `SequenceParallelBuilder` in [`ltx_pipelines/multigpu/sp_builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/multigpu/sp_builder.py) injects specialized self-attention operations:

```python
class SequenceParallelBuilder(...):
    def __init__(self, inner, attn_mgr, registry, tracker):
        sp_ops = create_video_self_attention_module_ops(attn_mgr)
        self._inner = inner.with_registry(registry).with_lora_load_device(cuda_device)\
                       .with_module_ops((*inner.module_ops, sp_ops))
        self._tracker = tracker
        self._attn_mgr = attn_mgr

```

The build step wraps the constructed model similarly:

```python
model = self._tracker.build(self._inner, device=device, dtype=dtype, **kwargs)
return SequenceParallelModelWrapper(model, self._attn_mgr)

```

### Runtime Execution in SequenceParallelModelWrapper

Located in [`ltx_core/multigpu/transformer/sequence_parallel.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/multigpu/transformer/sequence_parallel.py), this wrapper:

1. **Splits the temporal axis** — each rank receives `seq = latent[:, :, start:end, ...]`, a contiguous frame slice
2. **Uses custom attention kernels** — `create_video_self_attention_module_ops` replaces default self-attention, enabling global attention across ranks via NCCL all-to-all
3. **Enforces communication timeouts** — `AttentionManager` exposes `all2all_timeout_seconds` for configurable synchronization boundaries

### Pipeline Usage

Sequence parallel integrates identically to TDP by swapping the builder:

```python
pipeline.stage_2._transformer_builder = SequenceParallelBuilder(
    inner=builder,
    attn_mgr=attention_manager,
    registry=registry,
    tracker=weight_tracker,
)

```

## Shared Infrastructure: Weight Tracking and GPU Residency

Both builders rely on `TransformerWeightTracker` ([`ltx_pipelines/multigpu/weight_tracker.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/multigpu/weight_tracker.py)) to maintain **GPU-resident parameters** across model rebuilds. This eliminates redundant loading and explains why `keeps_gpu_resident_weights` returns `True` in both [`tdp_builder.py`](https://github.com/Lightricks/LTX-2/blob/main/tdp_builder.py) (lines 55-56) and [`sp_builder.py`](https://github.com/Lightricks/LTX-2/blob/main/sp_builder.py) (lines 45-46).

The tracker operates as a shared cache: weights are loaded once and rebound to new model instances as pipeline configurations change.

## Comparing the Two Parallelism Strategies

| Aspect | Tiled Data Parallel | Sequence Parallel |
|--------|---------------------|-------------------|
| **Parallelized dimension** | Spatial (H × W tiles) | Temporal (frame sequence) |
| **Core wrapper** | `TiledDataParallelModelWrapper` | `SequenceParallelModelWrapper` |
| **Custom kernels required** | No | Yes (attention ops) |
| **Communication pattern** | Scatter/Gather per-step | All-to-all within attention |
| **Best for** | High-resolution frames | Long-duration videos |

## Interaction and Composition

A single transformer uses **either** TDP or SP, not both simultaneously. The pipeline selects the strategy by configuring the appropriate builder for stage 2.

Advanced users can compose both approaches spatially and temporally—tiling frames first, then applying sequence parallelism within each tile—though this requires a custom wrapper. The foundational components (`TileCountConfig`, `AttentionManager`, both wrapper classes) support such extensions.

## Complete Example: Tiled Data Parallel Setup

```python
import torch
import torch.distributed as dist
from ltx_pipelines.ti2vid_two_stages_mgpu import TI2VidTwoStagesMGPU
from ltx_pipelines.multigpu.tdp_builder import TiledDataParallelBuilder
from ltx_pipelines.multigpu.weight_tracker import TransformerWeightTracker
from ltx_core.tiling import TileCountConfig
from ltx_core.utils import VideoLatentTools

# Initialize distributed world

dist.init_process_group(backend="nccl")
rank = dist.get_rank()
group = dist.group.WORLD

# Configure 2×2 spatial tiling

tiling = TileCountConfig(num_tiles_h=2, num_tiles_w=2)

# GPU-resident weight tracking

tracker = TransformerWeightTracker()

# Build tiled-parallel wrapper

tdp_builder = TiledDataParallelBuilder(
    inner=single_gpu_builder,
    group=group,
    tiling=tiling,
    registry=registry,
    tracker=tracker,
)

# Inject into pipeline stage 2

pipeline = TI2VidTwoStagesMGPU(...)
pipeline.stage_2._transformer_builder = tdp_builder

# Run with automatic scatter/gather

video, audio = pipeline(prompt, video_tools=VideoLatentTools(rank, tiling))

```

## Complete Example: Sequence Parallel Setup

```python
from ltx_pipelines.multigpu.sp_builder import SequenceParallelBuilder
from ltx_core.multigpu.transformer.attention import AttentionManager

# Attention manager with 30-second all-to-all timeout

attn_mgr = AttentionManager(group=group, timeout=30.0)

sp_builder = SequenceParallelBuilder(
    inner=single_gpu_builder,
    attn_mgr=attn_mgr,
    registry=registry,
    tracker=tracker,
)

pipeline.stage_2._transformer_builder = sp_builder

# Inference proceeds identically; temporal splitting is transparent

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`ltx_pipelines/multigpu/tdp_builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/multigpu/tdp_builder.py) | Constructs `TiledDataParallelModelWrapper` |
| [`ltx_pipelines/multigpu/sp_builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/multigpu/sp_builder.py) | Constructs `SequenceParallelModelWrapper` |
| [`ltx_core/multigpu/transformer/tiled_data_parallel.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/multigpu/transformer/tiled_data_parallel.py) | Spatial tile scatter/gather logic |
| [`ltx_core/multigpu/transformer/sequence_parallel.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/multigpu/transformer/sequence_parallel.py) | Temporal split and custom attention |
| [`ltx_pipelines/multigpu/weight_tracker.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/multigpu/weight_tracker.py) | GPU-resident parameter caching |
| [`ltx_pipelines/ti2vid_two_stages_mgpu.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/ti2vid_two_stages_mgpu.py) | Example MGPU pipeline integration |

## Summary

- **Tiled Data Parallel** in LTX-2 splits video latents spatially using `TiledDataParallelBuilder` and `TiledDataParallelModelWrapper`, with automatic scatter/gather via the supplied `ProcessGroup`.

- **Sequence Parallel** splits the temporal dimension using `SequenceParallelBuilder` and `SequenceParallelModelWrapper`, requiring custom attention kernels from `create_video_self_attention_module_ops`.

- Both strategies use `TransformerWeightTracker` to keep model weights GPU-resident across builds, avoiding reload overhead.

- The pipeline selects parallelism by assigning the appropriate builder to `stage_2._transformer_builder` in multi-GPU pipeline scripts.

- NCCL provides all inter-GPU synchronization; both wrappers accept a `dist.ProcessGroup` for deterministic communication ordering.

## Frequently Asked Questions

### Can I use Tiled Data Parallel and Sequence Parallel together on the same model?

No—a single transformer instance wraps with either `TiledDataParallelModelWrapper` or `SequenceParallelModelWrapper`. However, the underlying components (`TileCountConfig`, `AttentionManager`, wrapper classes) support building a custom combined wrapper that tiles spatially then applies sequence parallelism within each tile.

### How does LTX-2 handle synchronization between GPUs?

Both wrappers use the `dist.ProcessGroup` passed to their builders. TDP performs scatter/gather operations around the forward pass, while SP uses all-to-all communication within custom attention kernels. The `AttentionManager` in SP exposes `all2all_timeout_seconds` for fine-grained timeout control.

### What determines whether to choose TDP or Sequence Parallel?

Choose **Tiled Data Parallel** when generating high-resolution frames where spatial dimensions exceed single-GPU memory. Choose **Sequence Parallel** when processing long videos where the temporal dimension dominates memory usage. For extremely large generations, consider composing both approaches.

### Where are the model weights stored during multi-GPU inference?

`TransformerWeightTracker` keeps weights GPU-resident across builds, as confirmed by `keeps_gpu_resident_weights` returning `True` in both builders. This avoids repeated CPU-to-GPU transfers when pipeline configurations change.