How to Configure torch.compile for Maximum Inference Performance in LTX-2

Set mode=reduce-overhead, fullgraph=true, capture=true, and recompile_perturbed_block=false via the CompilationConfig dataclass to enable per-block compilation with CUDA-graph capture for peak LTX-2 inference speed.

The LTX-2 video generation model from Lightricks accelerates transformer inference by compiling each transformer block with torch.compile. This article explains how to configure the immutable CompilationConfig dataclass—defined in packages/ltx-core/src/ltx_core/model/transformer/compiling.py—to extract maximum performance from your GPU hardware.

Understanding the CompilationConfig Dataclass

The CompilationConfig class controls every aspect of PyTorch compilation behavior. Here are the fields that matter most for inference optimization:

Field Purpose Recommended Value for Peak Performance
mode torch.compile optimization mode reduce-overhead or max-autotune
backend Inductor backend selection "inductor"
fullgraph Abort on graph breaks to force full compilation True
dynamic Dynamic-shape tracing for variable sequence lengths None or True
inductor_config Fine-grained torch._inductor.config overrides {"max_autotune": True}
dynamo_config Fine-grained torch._dynamo.config overrides {"cache_size_limit": 0}
seq_dim_dynamic Mark sequence dimension as dynamic True (default)
recompile_perturbed_block Separate graph per guidance-perturbed block False for inference
capture Capture block-loop as single CUDA graph True for fastest runtime

The seq_dim_dynamic field is particularly important for video generation workloads. When enabled, the helper _mark_seq_dim_dynamic calls torch._dynamo.mark_dynamic on the sequence dimension, allowing a single compiled artifact to serve any sequence length rather than triggering recompilation.

How LTX-2 Applies torch.compile

The compilation pipeline in _compile_blocks (lines 61-71 of compiling.py) works in three stages:

  1. Per-block compilation – Each transformer block receives torch.compile(m, mode=..., backend=..., fullgraph=..., dynamic=...), producing shape-polymorphic Inductor kernels.

  2. Dynamic dimension marking – With seq_dim_dynamic=True, sequence dimensions are marked dynamic via _mark_seq_dim_dynamic (lines 67-95), eliminating recompilation when video lengths vary.

  3. CUDA-graph capture – When capture=True, the CudaGraphRunner wrapper records the entire block-loop as a CUDA graph on first inference, then replays it to eliminate per-block kernel launch overhead.

CLI Configuration for Maximum Speed

The pipelines package exposes compilation settings through the --compile flag, parsed in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 76-105). Use key-value pairs to construct your configuration:

python -m ltx_pipelines.ti2vid_one_stage \
  --checkpoint-path /path/to/model.safetensors \
  --prompt "A sunrise over mountains" \
  --output-path result.mp4 \
  --compile \
    mode=reduce-overhead \
    fullgraph=true \
    capture=true \
    recompile_perturbed_block=false \
    inductor_config='{"max_autotune": true}' \
    dynamo_config='{"cache_size_limit": 0}'

This configuration eliminates compilation overhead for guidance perturbations, enables aggressive kernel autotuning, and prevents Dynamo cache thrashing when sequence lengths change between generations.

Programmatic Configuration

For custom inference scripts, instantiate CompilationConfig directly and apply it via compile_transformer:

from ltx_core.model.transformer.compiling import CompilationConfig, compile_transformer
from ltx_core.model.transformer.model import LTXModel

# Load your model

model: LTXModel = load_transformer(...)

# Build maximum-performance configuration

fast_cfg = CompilationConfig(
    mode="reduce-overhead",
    backend="inductor",
    fullgraph=True,
    dynamic=None,
    inductor_config={"max_autotune": True},
    dynamo_config={"cache_size_limit": 0},
    seq_dim_dynamic=True,
    recompile_perturbed_block=False,
    capture=True,
)

# Apply compilation in-place

model = compile_transformer(model, fast_cfg)

# Inference with optimized path

video, audio = model(video_input, audio_input, perturbations=None)

The compile_transformer function mutates the model's transformer blocks according to the configuration, returning the modified model ready for accelerated inference.

Why These Settings Deliver Peak Performance

  • mode=reduce-overhead generates kernels optimized for minimal launch latency, critical for the many small operations in transformer blocks.
  • fullgraph=true guarantees no Python-side graph breaks, keeping the entire forward pass in compiled Inductor code.
  • capture=true moves beyond per-block compilation to a single captured CUDA graph, eliminating all Python overhead after warm-up.
  • recompile_perturbed_block=false avoids creating specialized graphs for guidance-mask variations, which matters during training but wastes compilation time during pure inference.
  • inductor_config.max_autotune spends upfront time searching for optimal fused kernels.
  • dynamo_config.cache_size_limit=0 disables caching that could retain stale shape-specialized artifacts.

Verification and Monitoring

When running under the trainer, LTX-2 logs compilation status via the Accelerate Dynamo plugin. Check packages/ltx-trainer/src/ltx_trainer/trainer.py (lines 804-815) to see how compilation health is reported. The trainer also warns about FSDP and torch.compile incompatibilities that could affect distributed inference setups.

Summary

  • Use CompilationConfig to control all torch.compile behavior in LTX-2
  • Set mode=reduce-overhead, fullgraph=true, capture=true, and recompile_perturbed_block=false for fastest inference
  • Enable seq_dim_dynamic to handle variable video lengths without recompilation
  • Configure inductor_config and dynamo_config for aggressive autotuning and cache management
  • Apply via --compile CLI flag or compile_transformer() programmatically

Frequently Asked Questions

What is the difference between reduce-overhead and max-autotune modes?

reduce-overhead minimizes kernel launch latency and compilation time, making it ideal for inference scenarios with varying inputs. max-autotune spends more time at compilation searching for optimal kernels, which can yield better steady-state performance but increases warmup time. For LTX-2 inference, start with reduce-overhead and benchmark against max-autotune on your specific hardware.

When should I disable capture=True?

Disable CUDA-graph capture when you need memory flexibility or dynamic control flow that cannot be captured in a static graph. The capture step requires a fixed memory allocation for the recorded operations, which may conflict with memory-constrained environments or workflows requiring mid-inference interventions.

Why does recompile_perturbed_block exist if false is faster for inference?

This flag exists to support classifier-free guidance training, where each transformer block may receive perturbed inputs requiring different optimized graphs. During training, true allows specialized compilation per perturbation. During pure inference with fixed guidance scales, false eliminates unnecessary compilation overhead.

Can I use torch.compile with multi-GPU inference?

Multi-GPU inference with FSDP (Fully Sharded Data Parallel) has known incompatibilities with torch.compile. The LTX-2 trainer logs warnings about these conflicts in trainer.py. For tensor-parallel or pipeline-parallel inference, test compilation carefully—some configurations may require disabling fullgraph or capture to maintain correctness.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →