Multi-GPU Pipeline Runners in LTX-2: Available Runners and Work Distribution Strategies

LTX-2 provides two multi-GPU inference runners—TI2VidTwoStagesRunner and DistilledRunner—that distribute work via sequence parallelism, tiled data parallelism, and distributed VAE decoding to accelerate single-video generation latency.

LTX-2's multi-GPU pipeline runners in the ltx-pipelines package enable high-throughput video generation by parallelizing computation across available GPUs. Both runners operate under the MGPUController, which manages worker processes and enforces synchronized execution. Understanding these runners and their distribution strategies is essential for optimizing inference performance on multi-GPU systems.

Available Multi-GPU Pipeline Runners

LTX-2 ships with two concrete implementations of the abstract MGPURunner class. Each targets a specific pipeline architecture and implements distinct work distribution patterns.

TI2VidTwoStagesRunner

The TI2VidTwoStagesRunner (defined in packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_mgpu.py) executes the two-stage text-to-video pipeline. This runner handles the full generation workflow: a base resolution pass followed by a spatial upsampling stage.

Work distribution follows a four-stage strategy:

  • Sequence Parallelism (SP) — The first stage's token sequence is split across GPUs via SequenceParallelBuilder using NCCL all-to-all attention kernels. This reduces per-step latency by distributing the compute-heavy attention operations.

  • Tiled Data Parallelism (TDP) — The second spatial upsampling stage processes height × width tiles on a per-GPU basis through TiledDataParallelBuilder. Each GPU handles distinct spatial regions independently.

  • Accelerate Gemma — The Gemma text encoder is sharded or replicated across GPUs using AccelerateGemmaBuilder from packages/ltx-pipelines/src/ltx_pipelines/multigpu/gemma_builders.py.

  • Distributed VAE Decoder — Latent tiles are decoded in parallel via DistributedDecoderBuilder in packages/ltx-pipelines/src/ltx_pipelines/multigpu/vae_builders.py. The driver (rank 0) assembles tiles and encodes the final video.

Launch via CLI:

python -m ltx_pipelines.ti2vid_two_stages_mgpu \
    --checkpoint-path path/to/ltx_checkpoint.safetensors \
    --distilled-lora path/to/distilled_lora.safetensors 1.0 \
    --spatial-upsampler-path path/to/upsampler.safetensors \
    --gemma-root path/to/gemma \
    --prompt "A futuristic city skyline at sunrise" \
    --output-path result.mp4

DistilledRunner

The DistilledRunner (defined in packages/ltx-pipelines/src/ltx_pipelines/distilled_mgpu.py) targets the single-stage distilled pipeline. This runner uses a shared diffusion stage for both low-resolution and full-resolution passes, eliminating the need for separate stage-specific distribution strategies.

Work distribution uses three components:

  • Shared Stage – Sequence Parallelism — The same diffusion stage (applied at both resolutions) is wrapped by SequenceParallelBuilder, splitting token sequences across GPUs identically for each pass.

  • Accelerate Gemma — Identical to the two-stage runner, using AccelerateGemmaBuilder for encoder distribution.

  • Distributed VAE Decoder — Tile-wise VAE decoding via DistributedDecoderBuilder, with the driver assembling final output.

Launch via CLI:

python -m ltx_pipelines.distilled_mgpu \
    --checkpoint-path path/to/ltx_checkpoint.safetensors \
    --spatial-upsampler-path path/to/upsampler.safetensors \
    --gemma-root path/to/gemma \
    --prompt "A calm forest with mist" \
    --output-path forest.mp4

How Work Distribution Works

Both LTX-2 multi-GPU pipeline runners rely on a common controller-based architecture with specialized builder components that transform single-GPU pipeline stages into distributed variants.

MGPUController and Worker Management

The MGPUController (packages/ltx-pipelines/docs/multigpu/controller.md):

  1. Spawns one worker process per visible GPU using torch.multiprocessing
  2. Creates the selected runner instance in each worker
  3. Streams generation jobs with serialized parameters
  4. Collects intermediate VAE tiles via SimpleQueue
  5. Shuts down the fleet gracefully on completion or error

The controller enforces single-job-at-a-time execution. Workers perform only forward passes; the driver rank (_DRIVER_RANK = 0) handles final video encoding and output assembly.

Weight Management Strategy

The TransformerWeightTracker (packages/ltx-pipelines/src/ltx_pipelines/multigpu/weight_tracker.py) coordinates memory layout:

  • Working copy — Full transformer weights replicated on every GPU (mutable, LoRA-fused)
  • Clean weights — Immutable base weights sharded via ShardedSD for efficient hot-swap operations

This design optimizes for latency reduction rather than memory savings. Each generation step completes faster through parallel computation, not through fitting larger models into limited memory.

Error Handling

Errors raised on any worker rank are converted to RunnerError or SymmetricRunnerError exceptions, allowing the host process to recover without zombie processes.

Using MGPUController Programmatically

For custom integrations, instantiate MGPUController directly with your chosen runner:

from ltx_pipelines.multigpu.controller import MGPUController
from ltx_pipelines.ti2vid_two_stages_mgpu import TI2VidTwoStagesRunner
import torch.multiprocessing as mp

# Create inter-process queue for VAE tile streaming

vae_queue = mp.get_context("spawn").SimpleQueue()

# Initialize controller with runner class

controller = MGPUController(TI2VidTwoStagesRunner)

# Start worker fleet

controller.start(
    model_paths=ModelPaths.from_monolith("path/to/checkpoint.safetensors"),
    prompt_enhancer_gemma_root="path/to/gemma",
    spatial_upsampler_path="path/to/upsampler.safetensors",
    vae_queue=vae_queue,
)

# Stream generation jobs

for _ in controller.stream(
    output_path="out.mp4",
    prompt="An astronaut floating in space",
    seed=42,
    height=1024,
    width=1536,
    frame_rate=30.0,
    num_frames=120,
    video_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
    audio_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
):
    pass  # yields final output path once driver finishes encoding

# Graceful shutdown

controller.shutdown()

The stream() method yields control after each job completion, enabling sequential batch processing without restarting worker processes.

Key Source Files

File Purpose
packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_mgpu.py TI2VidTwoStagesRunner implementation
packages/ltx-pipelines/src/ltx_pipelines/distilled_mgpu.py DistilledRunner implementation
packages/ltx-pipelines/src/ltx_pipelines/multigpu/runner.py Abstract MGPURunner base class
packages/ltx-pipelines/src/ltx_pipelines/multigpu/sp_builder.py SequenceParallelBuilder (token-dimension split)
packages/ltx-pipelines/src/ltx_pipelines/multigpu/tdp_builder.py TiledDataParallelBuilder (spatial tiling)
packages/ltx-pipelines/src/ltx_pipelines/multigpu/vae_builders.py DistributedDecoderBuilder (parallel VAE decode)
packages/ltx-pipelines/src/ltx_pipelines/multigpu/gemma_builders.py AccelerateGemmaBuilder (encoder sharding)
packages/ltx-pipelines/src/ltx_pipelines/multigpu/weight_tracker.py TransformerWeightTracker (LoRA hot-swap coordination)
packages/ltx-pipelines/docs/multigpu/controller.md MGPUController API documentation

Summary

  • Two multi-GPU pipeline runners are available in LTX-2: TI2VidTwoStagesRunner for two-stage text-to-video and DistilledRunner for single-stage distilled generation.

  • Work distribution strategies include sequence parallelism (token-split), tiled data parallelism (spatial-split), Accelerate-based Gemma sharding, and distributed VAE decoding.

  • Full weight replication on each GPU optimizes latency over memory efficiency, with sharded clean weights enabling fast LoRA hot-swapping.

  • MGPUController manages process lifecycle, error handling, and driver-rank aggregation for both runners.

  • CLI entry points at ltx_pipelines.ti2vid_two_stages_mgpu and ltx_pipelines.distilled_mgpu provide immediate multi-GPU execution without code changes.

Frequently Asked Questions

What is the difference between TI2VidTwoStagesRunner and DistilledRunner?

TI2VidTwoStagesRunner executes a base generation stage followed by a spatial upsampling stage, using sequence parallelism for the first stage and tiled data parallelism for the second. DistilledRunner uses a single shared diffusion stage for both resolutions, applying sequence parallelism throughout. Choose the former for maximum quality with two-stage refinement; choose the latter for faster generation with distilled model efficiency.

Does multi-GPU inference in LTX-2 reduce memory requirements per GPU?

No. LTX-2's multi-GPU pipeline runners replicate full transformer weights on every GPU. The distribution strategies target latency reduction—faster single-video generation—not memory optimization. Each GPU holds a complete working copy of weights, with only immutable clean weights sharded for efficient LoRA swapping.

How do I select which GPUs to use for inference?

The MGPUController automatically detects all visible CUDA devices. Control GPU visibility through standard environment variables before launching: CUDA_VISIBLE_DEVICES=0,1 python -m ltx_pipelines.distilled_mgpu .... The controller spawns exactly one worker per visible device.

What happens if a worker process crashes during generation?

The controller converts worker errors to RunnerError (asymmetric failures) or SymmetricRunnerError (synchronized failures across ranks). These exceptions propagate to the host process, triggering automatic worker fleet shutdown. The design prevents partial deadlocks and allows graceful recovery without manual process cleanup.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →