# Multi-GPU Pipeline Runners in LTX-2: Available Runners and Work Distribution Strategies

> Discover LTX-2 multi-GPU runners like TI2VidTwoStagesRunner and DistilledRunner. Learn how they use sequence parallelism, tiled data parallelism, and distributed VAE decoding to speed up video generation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-20

---

**LTX-2 provides two multi-GPU inference runners—`TI2VidTwoStagesRunner` and `DistilledRunner`—that distribute work via sequence parallelism, tiled data parallelism, and distributed VAE decoding to accelerate single-video generation latency.**

LTX-2's multi-GPU pipeline runners in the `ltx-pipelines` package enable high-throughput video generation by parallelizing computation across available GPUs. Both runners operate under the `MGPUController`, which manages worker processes and enforces synchronized execution. Understanding these runners and their distribution strategies is essential for optimizing inference performance on multi-GPU systems.

## Available Multi-GPU Pipeline Runners

LTX-2 ships with two concrete implementations of the abstract `MGPURunner` class. Each targets a specific pipeline architecture and implements distinct work distribution patterns.

### TI2VidTwoStagesRunner

The `TI2VidTwoStagesRunner` (defined in [`packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_mgpu.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_mgpu.py)) executes the two-stage text-to-video pipeline. This runner handles the full generation workflow: a base resolution pass followed by a spatial upsampling stage.

Work distribution follows a four-stage strategy:

- **Sequence Parallelism (SP)** — The first stage's token sequence is split across GPUs via `SequenceParallelBuilder` using NCCL all-to-all attention kernels. This reduces per-step latency by distributing the compute-heavy attention operations.

- **Tiled Data Parallelism (TDP)** — The second spatial upsampling stage processes height × width tiles on a per-GPU basis through `TiledDataParallelBuilder`. Each GPU handles distinct spatial regions independently.

- **Accelerate Gemma** — The Gemma text encoder is sharded or replicated across GPUs using `AccelerateGemmaBuilder` from [`packages/ltx-pipelines/src/ltx_pipelines/multigpu/gemma_builders.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/gemma_builders.py).

- **Distributed VAE Decoder** — Latent tiles are decoded in parallel via `DistributedDecoderBuilder` in [`packages/ltx-pipelines/src/ltx_pipelines/multigpu/vae_builders.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/vae_builders.py). The driver (rank 0) assembles tiles and encodes the final video.

Launch via CLI:

```bash
python -m ltx_pipelines.ti2vid_two_stages_mgpu \
    --checkpoint-path path/to/ltx_checkpoint.safetensors \
    --distilled-lora path/to/distilled_lora.safetensors 1.0 \
    --spatial-upsampler-path path/to/upsampler.safetensors \
    --gemma-root path/to/gemma \
    --prompt "A futuristic city skyline at sunrise" \
    --output-path result.mp4

```

### DistilledRunner

The `DistilledRunner` (defined in [`packages/ltx-pipelines/src/ltx_pipelines/distilled_mgpu.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/distilled_mgpu.py)) targets the single-stage distilled pipeline. This runner uses a shared diffusion stage for both low-resolution and full-resolution passes, eliminating the need for separate stage-specific distribution strategies.

Work distribution uses three components:

- **Shared Stage – Sequence Parallelism** — The same diffusion stage (applied at both resolutions) is wrapped by `SequenceParallelBuilder`, splitting token sequences across GPUs identically for each pass.

- **Accelerate Gemma** — Identical to the two-stage runner, using `AccelerateGemmaBuilder` for encoder distribution.

- **Distributed VAE Decoder** — Tile-wise VAE decoding via `DistributedDecoderBuilder`, with the driver assembling final output.

Launch via CLI:

```bash
python -m ltx_pipelines.distilled_mgpu \
    --checkpoint-path path/to/ltx_checkpoint.safetensors \
    --spatial-upsampler-path path/to/upsampler.safetensors \
    --gemma-root path/to/gemma \
    --prompt "A calm forest with mist" \
    --output-path forest.mp4

```

## How Work Distribution Works

Both LTX-2 multi-GPU pipeline runners rely on a common controller-based architecture with specialized builder components that transform single-GPU pipeline stages into distributed variants.

### MGPUController and Worker Management

The `MGPUController` ([`packages/ltx-pipelines/docs/multigpu/controller.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/docs/multigpu/controller.md)):

1. Spawns one worker process per visible GPU using `torch.multiprocessing`
2. Creates the selected runner instance in each worker
3. Streams generation jobs with serialized parameters
4. Collects intermediate VAE tiles via `SimpleQueue`
5. Shuts down the fleet gracefully on completion or error

The controller enforces **single-job-at-a-time execution**. Workers perform only forward passes; the driver rank (`_DRIVER_RANK = 0`) handles final video encoding and output assembly.

### Weight Management Strategy

The `TransformerWeightTracker` ([`packages/ltx-pipelines/src/ltx_pipelines/multigpu/weight_tracker.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/weight_tracker.py)) coordinates memory layout:

- **Working copy** — Full transformer weights replicated on every GPU (mutable, LoRA-fused)
- **Clean weights** — Immutable base weights sharded via `ShardedSD` for efficient hot-swap operations

This design optimizes for **latency reduction** rather than memory savings. Each generation step completes faster through parallel computation, not through fitting larger models into limited memory.

### Error Handling

Errors raised on any worker rank are converted to `RunnerError` or `SymmetricRunnerError` exceptions, allowing the host process to recover without zombie processes.

## Using MGPUController Programmatically

For custom integrations, instantiate `MGPUController` directly with your chosen runner:

```python
from ltx_pipelines.multigpu.controller import MGPUController
from ltx_pipelines.ti2vid_two_stages_mgpu import TI2VidTwoStagesRunner
import torch.multiprocessing as mp

# Create inter-process queue for VAE tile streaming

vae_queue = mp.get_context("spawn").SimpleQueue()

# Initialize controller with runner class

controller = MGPUController(TI2VidTwoStagesRunner)

# Start worker fleet

controller.start(
    model_paths=ModelPaths.from_monolith("path/to/checkpoint.safetensors"),
    prompt_enhancer_gemma_root="path/to/gemma",
    spatial_upsampler_path="path/to/upsampler.safetensors",
    vae_queue=vae_queue,
)

# Stream generation jobs

for _ in controller.stream(
    output_path="out.mp4",
    prompt="An astronaut floating in space",
    seed=42,
    height=1024,
    width=1536,
    frame_rate=30.0,
    num_frames=120,
    video_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
    audio_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
):
    pass  # yields final output path once driver finishes encoding

# Graceful shutdown

controller.shutdown()

```

The `stream()` method yields control after each job completion, enabling sequential batch processing without restarting worker processes.

## Key Source Files

| File | Purpose |
|------|---------|
| [`packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_mgpu.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_mgpu.py) | `TI2VidTwoStagesRunner` implementation |
| [`packages/ltx-pipelines/src/ltx_pipelines/distilled_mgpu.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/distilled_mgpu.py) | `DistilledRunner` implementation |
| [`packages/ltx-pipelines/src/ltx_pipelines/multigpu/runner.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/runner.py) | Abstract `MGPURunner` base class |
| [`packages/ltx-pipelines/src/ltx_pipelines/multigpu/sp_builder.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/sp_builder.py) | `SequenceParallelBuilder` (token-dimension split) |
| [`packages/ltx-pipelines/src/ltx_pipelines/multigpu/tdp_builder.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/tdp_builder.py) | `TiledDataParallelBuilder` (spatial tiling) |
| [`packages/ltx-pipelines/src/ltx_pipelines/multigpu/vae_builders.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/vae_builders.py) | `DistributedDecoderBuilder` (parallel VAE decode) |
| [`packages/ltx-pipelines/src/ltx_pipelines/multigpu/gemma_builders.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/gemma_builders.py) | `AccelerateGemmaBuilder` (encoder sharding) |
| [`packages/ltx-pipelines/src/ltx_pipelines/multigpu/weight_tracker.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/multigpu/weight_tracker.py) | `TransformerWeightTracker` (LoRA hot-swap coordination) |
| [`packages/ltx-pipelines/docs/multigpu/controller.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/docs/multigpu/controller.md) | `MGPUController` API documentation |

## Summary

- **Two multi-GPU pipeline runners** are available in LTX-2: `TI2VidTwoStagesRunner` for two-stage text-to-video and `DistilledRunner` for single-stage distilled generation.

- **Work distribution strategies** include sequence parallelism (token-split), tiled data parallelism (spatial-split), Accelerate-based Gemma sharding, and distributed VAE decoding.

- **Full weight replication** on each GPU optimizes latency over memory efficiency, with sharded clean weights enabling fast LoRA hot-swapping.

- **MGPUController** manages process lifecycle, error handling, and driver-rank aggregation for both runners.

- **CLI entry points** at `ltx_pipelines.ti2vid_two_stages_mgpu` and `ltx_pipelines.distilled_mgpu` provide immediate multi-GPU execution without code changes.

## Frequently Asked Questions

### What is the difference between TI2VidTwoStagesRunner and DistilledRunner?

`TI2VidTwoStagesRunner` executes a base generation stage followed by a spatial upsampling stage, using sequence parallelism for the first stage and tiled data parallelism for the second. `DistilledRunner` uses a single shared diffusion stage for both resolutions, applying sequence parallelism throughout. Choose the former for maximum quality with two-stage refinement; choose the latter for faster generation with distilled model efficiency.

### Does multi-GPU inference in LTX-2 reduce memory requirements per GPU?

No. LTX-2's multi-GPU pipeline runners replicate full transformer weights on every GPU. The distribution strategies target **latency reduction**—faster single-video generation—not memory optimization. Each GPU holds a complete working copy of weights, with only immutable clean weights sharded for efficient LoRA swapping.

### How do I select which GPUs to use for inference?

The `MGPUController` automatically detects all visible CUDA devices. Control GPU visibility through standard environment variables before launching: `CUDA_VISIBLE_DEVICES=0,1 python -m ltx_pipelines.distilled_mgpu ...`. The controller spawns exactly one worker per visible device.

### What happens if a worker process crashes during generation?

The controller converts worker errors to `RunnerError` (asymmetric failures) or `SymmetricRunnerError` (synchronized failures across ranks). These exceptions propagate to the host process, triggering automatic worker fleet shutdown. The design prevents partial deadlocks and allows graceful recovery without manual process cleanup.