How to Use `torch.compile` for Faster Inference with LTX-2
Enable torch.compile in LTX-2 by setting mode=reduce-overhead, fullgraph=true, and capture=true in the CompilationConfig dataclass to eliminate per-block launch overhead and maximize GPU throughput.
LTX-2 accelerates transformer-based video generation by compiling each transformer block with PyTorch's torch.compile. This article explains how to configure compilation for peak inference performance using the CompilationConfig system in packages/ltx-core/src/ltx_core/model/transformer/compiling.py.
Understanding the Compilation Configuration
The CompilationConfig dataclass controls all aspects of torch.compile behavior in LTX-2. Here are the key fields that impact inference speed:
| Field | Purpose | Recommended Value |
|---|---|---|
mode |
Optimization level (reduce-overhead, max-autotune, or default) |
reduce-overhead or max-autotune |
backend |
Compilation backend | "inductor" |
fullgraph |
Force complete graph compilation (fail on breaks) | True |
dynamic |
Enable dynamic-shape tracing | None or True |
inductor_config |
Fine-grained Inductor settings | {"max_autotune": True} |
dynamo_config |
Fine-grained Dynamo settings | {"cache_size_limit": 0} |
seq_dim_dynamic |
Mark sequence dimension as dynamic | True (default) |
recompile_perturbed_block |
Re-compile for guidance perturbations | False for inference |
capture |
Record block-loop as CUDA graph | True for peak performance |
CLI Method: Using the --compile Flag
The fastest way to enable torch.compile for LTX-2 inference is through the command-line interface. The --compile flag in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py parses KEY=VALUE pairs into a CompilationConfig instance.
Maximum Performance Command
python -m ltx_pipelines.ti2vid_one_stage \
--checkpoint-path /path/to/model.safetensors \
--prompt "A sunrise over mountains" \
--output-path result.mp4 \
--compile \
mode=reduce-overhead \
fullgraph=true \
capture=true \
recompile_perturbed_block=false \
inductor_config='{"max_autotune": true}' \
dynamo_config='{"cache_size_limit": 0}'
This configuration:
mode=reduce-overhead— minimizes kernel launch latencyfullgraph=true— prevents graph breaks that would fall back to eager modecapture=true— records a CUDA graph of the entire block-looprecompile_perturbed_block=false— avoids re-compilation for guidance masks (unnecessary during inference)max_autotune— enables aggressive kernel searchcache_size_limit=0— prevents Dynamo cache thrashing on variable sequence lengths
Programmatic Method: Direct CompilationConfig Usage
For custom scripts or integration into larger pipelines, instantiate CompilationConfig directly and apply it via compile_transformer().
from ltx_core.model.transformer.compiling import (
CompilationConfig,
compile_transformer
)
from ltx_core.model.transformer.model import LTXModel
# Load your model
model: LTXModel = load_transformer(...)
# Configure for maximum inference speed
fast_cfg = CompilationConfig(
mode="reduce-overhead",
backend="inductor",
fullgraph=True,
dynamic=None,
inductor_config={"max_autotune": True},
dynamo_config={"cache_size_limit": 0},
seq_dim_dynamic=True,
recompile_perturbed_block=False,
capture=True,
)
# Apply compilation (mutates model in-place)
model = compile_transformer(model, fast_cfg)
# Run inference
video, audio = model(video_input, audio_input, perturbations=None)
How LTX-2 Compilation Works Internally
Per-Block Compilation with _compile_blocks
In packages/ltx-core/src/ltx_core/model/transformer/compiling.py, the _compile_blocks function (lines 61-71) iterates through each transformer block and applies torch.compile() with the configured options:
# From compiling.py lines 61-71
compiled_blocks = [
torch.compile(
block,
mode=config.mode,
backend=config.backend,
fullgraph=config.fullgraph,
dynamic=config.dynamic,
)
for block in model.blocks
]
Dynamic Dimension Marking with seq_dim_dynamic
When seq_dim_dynamic=True, the helper _mark_seq_dim_dynamic (lines 67-95) calls torch._dynamo.mark_dynamic() on the sequence dimension. This allows a single compiled artifact to handle variable sequence lengths without re-compilation.
CUDA Graph Capture with capture=True
If capture=True, the block-loop is wrapped by CudaGraphRunner (lines 78-84). This records the compiled operations into a CUDA graph on the first forward pass, then replays the graph on subsequent calls—eliminating CPU overhead entirely after warm-up.
Key Source Files Reference
| File | Purpose |
|---|---|
packages/ltx-core/src/ltx_core/model/transformer/compiling.py |
CompilationConfig definition, _compile_blocks, _mark_seq_dim_dynamic, CUDA graph integration |
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py |
CLI --compile flag parsing (lines 76-105) |
packages/ltx-trainer/src/ltx_trainer/trainer.py |
Compilation status logging and FSDP compatibility warnings (lines 804-815) |
packages/ltx-core/src/ltx_core/model/transformer/model.py |
LTXModel class with transformer blocks |
Troubleshooting Common Issues
Compilation fails with graph breaks
Set fullgraph=False to allow partial compilation, or trace the error to identify unsupported operations (typically custom CUDA kernels or complex control flow).
Variable sequence lengths cause re-compilation
Ensure seq_dim_dynamic=True and dynamic=None or dynamic=True. Verify that torch._dynamo.mark_dynamic is being applied correctly in _mark_seq_dim_dynamic.
Memory usage spikes during first inference
This is expected warm-up behavior. The CUDA graph capture (capture=True) and Inductor autotuning both execute during the first forward pass. Subsequent passes use the cached artifacts.
Summary
torch.compilein LTX-2 is configured through theCompilationConfigdataclass incompiling.py- CLI: Use
--compile mode=reduce-overhead fullgraph=true capture=truefor one-line activation - Programmatic: Build a
CompilationConfigand pass tocompile_transformer(model, config) - Peak performance requires:
mode=reduce-overheadormax-autotune,fullgraph=true,capture=true, andrecompile_perturbed_block=false - Dynamic shapes: Handled automatically via
seq_dim_dynamic=Trueandtorch._dynamo.mark_dynamic
Frequently Asked Questions
What is the fastest torch.compile mode for LTX-2 inference?
reduce-overhead provides the best balance of compilation speed and runtime performance for most GPUs. max-autotune can yield faster kernels but increases compilation time significantly. Both outperform default mode for video generation workloads where the same shapes repeat across many frames.
Why does LTX-2 compile per-block instead of the full model?
Per-block compilation in _compile_blocks allows finer-grained control over which components are optimized and enables selective re-compilation when guidance perturbations change. The optional capture=true flag then fuses the compiled blocks into a single CUDA graph for minimal launch overhead.
Does capture=true work with variable video resolutions?
Yes, but only if seq_dim_dynamic=True (the default). The _mark_seq_dim_dynamic helper marks the sequence dimension as dynamic, allowing the captured CUDA graph to handle varying lengths. However, spatial dimensions (height × width) typically require separate compiled artifacts unless they remain constant across generations.
Can I use torch.compile with multi-GPU training in LTX-2?
Not directly with FSDP. The trainer in packages/ltx-trainer/src/ltx_trainer/trainer.py logs a warning about torch.compile incompatibilities with FSDP at lines 804-815. For multi-GPU inference, consider data-parallel approaches or compile on a single GPU before distributing weights.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →