Multi-GPU Pipeline Runners in LTX-2: Available Runners and Work Distribution Strategies
LTX-2 provides two multi-GPU inference runners—TI2VidTwoStagesRunner and DistilledRunner—that distribute work via sequence parallelism, tiled data parallelism, and distributed VAE decoding to accelerate single-video generation latency.
LTX-2's multi-GPU pipeline runners in the ltx-pipelines package enable high-throughput video generation by parallelizing computation across available GPUs. Both runners operate under the MGPUController, which manages worker processes and enforces synchronized execution. Understanding these runners and their distribution strategies is essential for optimizing inference performance on multi-GPU systems.
Available Multi-GPU Pipeline Runners
LTX-2 ships with two concrete implementations of the abstract MGPURunner class. Each targets a specific pipeline architecture and implements distinct work distribution patterns.
TI2VidTwoStagesRunner
The TI2VidTwoStagesRunner (defined in packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_mgpu.py) executes the two-stage text-to-video pipeline. This runner handles the full generation workflow: a base resolution pass followed by a spatial upsampling stage.
Work distribution follows a four-stage strategy:
-
Sequence Parallelism (SP) — The first stage's token sequence is split across GPUs via
SequenceParallelBuilderusing NCCL all-to-all attention kernels. This reduces per-step latency by distributing the compute-heavy attention operations. -
Tiled Data Parallelism (TDP) — The second spatial upsampling stage processes height × width tiles on a per-GPU basis through
TiledDataParallelBuilder. Each GPU handles distinct spatial regions independently. -
Accelerate Gemma — The Gemma text encoder is sharded or replicated across GPUs using
AccelerateGemmaBuilderfrompackages/ltx-pipelines/src/ltx_pipelines/multigpu/gemma_builders.py. -
Distributed VAE Decoder — Latent tiles are decoded in parallel via
DistributedDecoderBuilderinpackages/ltx-pipelines/src/ltx_pipelines/multigpu/vae_builders.py. The driver (rank 0) assembles tiles and encodes the final video.
Launch via CLI:
python -m ltx_pipelines.ti2vid_two_stages_mgpu \
--checkpoint-path path/to/ltx_checkpoint.safetensors \
--distilled-lora path/to/distilled_lora.safetensors 1.0 \
--spatial-upsampler-path path/to/upsampler.safetensors \
--gemma-root path/to/gemma \
--prompt "A futuristic city skyline at sunrise" \
--output-path result.mp4
DistilledRunner
The DistilledRunner (defined in packages/ltx-pipelines/src/ltx_pipelines/distilled_mgpu.py) targets the single-stage distilled pipeline. This runner uses a shared diffusion stage for both low-resolution and full-resolution passes, eliminating the need for separate stage-specific distribution strategies.
Work distribution uses three components:
-
Shared Stage – Sequence Parallelism — The same diffusion stage (applied at both resolutions) is wrapped by
SequenceParallelBuilder, splitting token sequences across GPUs identically for each pass. -
Accelerate Gemma — Identical to the two-stage runner, using
AccelerateGemmaBuilderfor encoder distribution. -
Distributed VAE Decoder — Tile-wise VAE decoding via
DistributedDecoderBuilder, with the driver assembling final output.
Launch via CLI:
python -m ltx_pipelines.distilled_mgpu \
--checkpoint-path path/to/ltx_checkpoint.safetensors \
--spatial-upsampler-path path/to/upsampler.safetensors \
--gemma-root path/to/gemma \
--prompt "A calm forest with mist" \
--output-path forest.mp4
How Work Distribution Works
Both LTX-2 multi-GPU pipeline runners rely on a common controller-based architecture with specialized builder components that transform single-GPU pipeline stages into distributed variants.
MGPUController and Worker Management
The MGPUController (packages/ltx-pipelines/docs/multigpu/controller.md):
- Spawns one worker process per visible GPU using
torch.multiprocessing - Creates the selected runner instance in each worker
- Streams generation jobs with serialized parameters
- Collects intermediate VAE tiles via
SimpleQueue - Shuts down the fleet gracefully on completion or error
The controller enforces single-job-at-a-time execution. Workers perform only forward passes; the driver rank (_DRIVER_RANK = 0) handles final video encoding and output assembly.
Weight Management Strategy
The TransformerWeightTracker (packages/ltx-pipelines/src/ltx_pipelines/multigpu/weight_tracker.py) coordinates memory layout:
- Working copy — Full transformer weights replicated on every GPU (mutable, LoRA-fused)
- Clean weights — Immutable base weights sharded via
ShardedSDfor efficient hot-swap operations
This design optimizes for latency reduction rather than memory savings. Each generation step completes faster through parallel computation, not through fitting larger models into limited memory.
Error Handling
Errors raised on any worker rank are converted to RunnerError or SymmetricRunnerError exceptions, allowing the host process to recover without zombie processes.
Using MGPUController Programmatically
For custom integrations, instantiate MGPUController directly with your chosen runner:
from ltx_pipelines.multigpu.controller import MGPUController
from ltx_pipelines.ti2vid_two_stages_mgpu import TI2VidTwoStagesRunner
import torch.multiprocessing as mp
# Create inter-process queue for VAE tile streaming
vae_queue = mp.get_context("spawn").SimpleQueue()
# Initialize controller with runner class
controller = MGPUController(TI2VidTwoStagesRunner)
# Start worker fleet
controller.start(
model_paths=ModelPaths.from_monolith("path/to/checkpoint.safetensors"),
prompt_enhancer_gemma_root="path/to/gemma",
spatial_upsampler_path="path/to/upsampler.safetensors",
vae_queue=vae_queue,
)
# Stream generation jobs
for _ in controller.stream(
output_path="out.mp4",
prompt="An astronaut floating in space",
seed=42,
height=1024,
width=1536,
frame_rate=30.0,
num_frames=120,
video_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
audio_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
):
pass # yields final output path once driver finishes encoding
# Graceful shutdown
controller.shutdown()
The stream() method yields control after each job completion, enabling sequential batch processing without restarting worker processes.
Key Source Files
| File | Purpose |
|---|---|
packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_mgpu.py |
TI2VidTwoStagesRunner implementation |
packages/ltx-pipelines/src/ltx_pipelines/distilled_mgpu.py |
DistilledRunner implementation |
packages/ltx-pipelines/src/ltx_pipelines/multigpu/runner.py |
Abstract MGPURunner base class |
packages/ltx-pipelines/src/ltx_pipelines/multigpu/sp_builder.py |
SequenceParallelBuilder (token-dimension split) |
packages/ltx-pipelines/src/ltx_pipelines/multigpu/tdp_builder.py |
TiledDataParallelBuilder (spatial tiling) |
packages/ltx-pipelines/src/ltx_pipelines/multigpu/vae_builders.py |
DistributedDecoderBuilder (parallel VAE decode) |
packages/ltx-pipelines/src/ltx_pipelines/multigpu/gemma_builders.py |
AccelerateGemmaBuilder (encoder sharding) |
packages/ltx-pipelines/src/ltx_pipelines/multigpu/weight_tracker.py |
TransformerWeightTracker (LoRA hot-swap coordination) |
packages/ltx-pipelines/docs/multigpu/controller.md |
MGPUController API documentation |
Summary
-
Two multi-GPU pipeline runners are available in LTX-2:
TI2VidTwoStagesRunnerfor two-stage text-to-video andDistilledRunnerfor single-stage distilled generation. -
Work distribution strategies include sequence parallelism (token-split), tiled data parallelism (spatial-split), Accelerate-based Gemma sharding, and distributed VAE decoding.
-
Full weight replication on each GPU optimizes latency over memory efficiency, with sharded clean weights enabling fast LoRA hot-swapping.
-
MGPUController manages process lifecycle, error handling, and driver-rank aggregation for both runners.
-
CLI entry points at
ltx_pipelines.ti2vid_two_stages_mgpuandltx_pipelines.distilled_mgpuprovide immediate multi-GPU execution without code changes.
Frequently Asked Questions
What is the difference between TI2VidTwoStagesRunner and DistilledRunner?
TI2VidTwoStagesRunner executes a base generation stage followed by a spatial upsampling stage, using sequence parallelism for the first stage and tiled data parallelism for the second. DistilledRunner uses a single shared diffusion stage for both resolutions, applying sequence parallelism throughout. Choose the former for maximum quality with two-stage refinement; choose the latter for faster generation with distilled model efficiency.
Does multi-GPU inference in LTX-2 reduce memory requirements per GPU?
No. LTX-2's multi-GPU pipeline runners replicate full transformer weights on every GPU. The distribution strategies target latency reduction—faster single-video generation—not memory optimization. Each GPU holds a complete working copy of weights, with only immutable clean weights sharded for efficient LoRA swapping.
How do I select which GPUs to use for inference?
The MGPUController automatically detects all visible CUDA devices. Control GPU visibility through standard environment variables before launching: CUDA_VISIBLE_DEVICES=0,1 python -m ltx_pipelines.distilled_mgpu .... The controller spawns exactly one worker per visible device.
What happens if a worker process crashes during generation?
The controller converts worker errors to RunnerError (asymmetric failures) or SymmetricRunnerError (synchronized failures across ranks). These exceptions propagate to the host process, triggering automatic worker fleet shutdown. The design prevents partial deadlocks and allows graceful recovery without manual process cleanup.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →