How to Configure torch.compile for Maximum Inference Performance in LTX-2
Set mode=reduce-overhead, fullgraph=true, capture=true, and recompile_perturbed_block=false via the CompilationConfig dataclass to enable per-block compilation with CUDA-graph capture for peak LTX-2 inference speed.
The LTX-2 video generation model from Lightricks accelerates transformer inference by compiling each transformer block with torch.compile. This article explains how to configure the immutable CompilationConfig dataclass—defined in packages/ltx-core/src/ltx_core/model/transformer/compiling.py—to extract maximum performance from your GPU hardware.
Understanding the CompilationConfig Dataclass
The CompilationConfig class controls every aspect of PyTorch compilation behavior. Here are the fields that matter most for inference optimization:
| Field | Purpose | Recommended Value for Peak Performance |
|---|---|---|
mode |
torch.compile optimization mode |
reduce-overhead or max-autotune |
backend |
Inductor backend selection | "inductor" |
fullgraph |
Abort on graph breaks to force full compilation | True |
dynamic |
Dynamic-shape tracing for variable sequence lengths | None or True |
inductor_config |
Fine-grained torch._inductor.config overrides |
{"max_autotune": True} |
dynamo_config |
Fine-grained torch._dynamo.config overrides |
{"cache_size_limit": 0} |
seq_dim_dynamic |
Mark sequence dimension as dynamic | True (default) |
recompile_perturbed_block |
Separate graph per guidance-perturbed block | False for inference |
capture |
Capture block-loop as single CUDA graph | True for fastest runtime |
The seq_dim_dynamic field is particularly important for video generation workloads. When enabled, the helper _mark_seq_dim_dynamic calls torch._dynamo.mark_dynamic on the sequence dimension, allowing a single compiled artifact to serve any sequence length rather than triggering recompilation.
How LTX-2 Applies torch.compile
The compilation pipeline in _compile_blocks (lines 61-71 of compiling.py) works in three stages:
-
Per-block compilation – Each transformer block receives
torch.compile(m, mode=..., backend=..., fullgraph=..., dynamic=...), producing shape-polymorphic Inductor kernels. -
Dynamic dimension marking – With
seq_dim_dynamic=True, sequence dimensions are marked dynamic via_mark_seq_dim_dynamic(lines 67-95), eliminating recompilation when video lengths vary. -
CUDA-graph capture – When
capture=True, theCudaGraphRunnerwrapper records the entire block-loop as a CUDA graph on first inference, then replays it to eliminate per-block kernel launch overhead.
CLI Configuration for Maximum Speed
The pipelines package exposes compilation settings through the --compile flag, parsed in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 76-105). Use key-value pairs to construct your configuration:
python -m ltx_pipelines.ti2vid_one_stage \
--checkpoint-path /path/to/model.safetensors \
--prompt "A sunrise over mountains" \
--output-path result.mp4 \
--compile \
mode=reduce-overhead \
fullgraph=true \
capture=true \
recompile_perturbed_block=false \
inductor_config='{"max_autotune": true}' \
dynamo_config='{"cache_size_limit": 0}'
This configuration eliminates compilation overhead for guidance perturbations, enables aggressive kernel autotuning, and prevents Dynamo cache thrashing when sequence lengths change between generations.
Programmatic Configuration
For custom inference scripts, instantiate CompilationConfig directly and apply it via compile_transformer:
from ltx_core.model.transformer.compiling import CompilationConfig, compile_transformer
from ltx_core.model.transformer.model import LTXModel
# Load your model
model: LTXModel = load_transformer(...)
# Build maximum-performance configuration
fast_cfg = CompilationConfig(
mode="reduce-overhead",
backend="inductor",
fullgraph=True,
dynamic=None,
inductor_config={"max_autotune": True},
dynamo_config={"cache_size_limit": 0},
seq_dim_dynamic=True,
recompile_perturbed_block=False,
capture=True,
)
# Apply compilation in-place
model = compile_transformer(model, fast_cfg)
# Inference with optimized path
video, audio = model(video_input, audio_input, perturbations=None)
The compile_transformer function mutates the model's transformer blocks according to the configuration, returning the modified model ready for accelerated inference.
Why These Settings Deliver Peak Performance
mode=reduce-overheadgenerates kernels optimized for minimal launch latency, critical for the many small operations in transformer blocks.fullgraph=trueguarantees no Python-side graph breaks, keeping the entire forward pass in compiled Inductor code.capture=truemoves beyond per-block compilation to a single captured CUDA graph, eliminating all Python overhead after warm-up.recompile_perturbed_block=falseavoids creating specialized graphs for guidance-mask variations, which matters during training but wastes compilation time during pure inference.inductor_config.max_autotunespends upfront time searching for optimal fused kernels.dynamo_config.cache_size_limit=0disables caching that could retain stale shape-specialized artifacts.
Verification and Monitoring
When running under the trainer, LTX-2 logs compilation status via the Accelerate Dynamo plugin. Check packages/ltx-trainer/src/ltx_trainer/trainer.py (lines 804-815) to see how compilation health is reported. The trainer also warns about FSDP and torch.compile incompatibilities that could affect distributed inference setups.
Summary
- Use
CompilationConfigto control alltorch.compilebehavior in LTX-2 - Set
mode=reduce-overhead,fullgraph=true,capture=true, andrecompile_perturbed_block=falsefor fastest inference - Enable
seq_dim_dynamicto handle variable video lengths without recompilation - Configure
inductor_configanddynamo_configfor aggressive autotuning and cache management - Apply via
--compileCLI flag orcompile_transformer()programmatically
Frequently Asked Questions
What is the difference between reduce-overhead and max-autotune modes?
reduce-overhead minimizes kernel launch latency and compilation time, making it ideal for inference scenarios with varying inputs. max-autotune spends more time at compilation searching for optimal kernels, which can yield better steady-state performance but increases warmup time. For LTX-2 inference, start with reduce-overhead and benchmark against max-autotune on your specific hardware.
When should I disable capture=True?
Disable CUDA-graph capture when you need memory flexibility or dynamic control flow that cannot be captured in a static graph. The capture step requires a fixed memory allocation for the recorded operations, which may conflict with memory-constrained environments or workflows requiring mid-inference interventions.
Why does recompile_perturbed_block exist if false is faster for inference?
This flag exists to support classifier-free guidance training, where each transformer block may receive perturbed inputs requiring different optimized graphs. During training, true allows specialized compilation per perturbation. During pure inference with fixed guidance scales, false eliminates unnecessary compilation overhead.
Can I use torch.compile with multi-GPU inference?
Multi-GPU inference with FSDP (Fully Sharded Data Parallel) has known incompatibilities with torch.compile. The LTX-2 trainer logs warnings about these conflicts in trainer.py. For tensor-parallel or pipeline-parallel inference, test compilation carefully—some configurations may require disabling fullgraph or capture to maintain correctness.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →