# How to Configure torch.compile for Maximum Inference Performance in LTX-2

> Achieve maximum LTX-2 inference performance by configuring torch.compile with specific settings for per-block compilation and CUDA-graph capture. Learn the optimal parameters.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: performance
- Published: 2026-08-18

---

**Set `mode=reduce-overhead`, `fullgraph=true`, `capture=true`, and `recompile_perturbed_block=false` via the `CompilationConfig` dataclass to enable per-block compilation with CUDA-graph capture for peak LTX-2 inference speed.**

The LTX-2 video generation model from Lightricks accelerates transformer inference by compiling each transformer block with `torch.compile`. This article explains how to configure the immutable **`CompilationConfig`** dataclass—defined in [`packages/ltx-core/src/ltx_core/model/transformer/compiling.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/compiling.py)—to extract maximum performance from your GPU hardware.

## Understanding the CompilationConfig Dataclass

The `CompilationConfig` class controls every aspect of PyTorch compilation behavior. Here are the fields that matter most for inference optimization:

| Field | Purpose | Recommended Value for Peak Performance |
|-------|---------|----------------------------------------|
| `mode` | `torch.compile` optimization mode | `reduce-overhead` or `max-autotune` |
| `backend` | Inductor backend selection | `"inductor"` |
| `fullgraph` | Abort on graph breaks to force full compilation | `True` |
| `dynamic` | Dynamic-shape tracing for variable sequence lengths | `None` or `True` |
| `inductor_config` | Fine-grained `torch._inductor.config` overrides | `{"max_autotune": True}` |
| `dynamo_config` | Fine-grained `torch._dynamo.config` overrides | `{"cache_size_limit": 0}` |
| `seq_dim_dynamic` | Mark sequence dimension as dynamic | `True` (default) |
| `recompile_perturbed_block` | Separate graph per guidance-perturbed block | `False` for inference |
| `capture` | Capture block-loop as single CUDA graph | `True` for fastest runtime |

The `seq_dim_dynamic` field is particularly important for video generation workloads. When enabled, the helper `_mark_seq_dim_dynamic` calls `torch._dynamo.mark_dynamic` on the sequence dimension, allowing a single compiled artifact to serve any sequence length rather than triggering recompilation.

## How LTX-2 Applies torch.compile

The compilation pipeline in `_compile_blocks` (lines 61-71 of [`compiling.py`](https://github.com/Lightricks/LTX-2/blob/main/compiling.py)) works in three stages:

1. **Per-block compilation** – Each transformer block receives `torch.compile(m, mode=..., backend=..., fullgraph=..., dynamic=...)`, producing shape-polymorphic Inductor kernels.

2. **Dynamic dimension marking** – With `seq_dim_dynamic=True`, sequence dimensions are marked dynamic via `_mark_seq_dim_dynamic` (lines 67-95), eliminating recompilation when video lengths vary.

3. **CUDA-graph capture** – When `capture=True`, the `CudaGraphRunner` wrapper records the entire block-loop as a CUDA graph on first inference, then replays it to eliminate per-block kernel launch overhead.

## CLI Configuration for Maximum Speed

The pipelines package exposes compilation settings through the `--compile` flag, parsed in [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) (lines 76-105). Use key-value pairs to construct your configuration:

```bash
python -m ltx_pipelines.ti2vid_one_stage \
  --checkpoint-path /path/to/model.safetensors \
  --prompt "A sunrise over mountains" \
  --output-path result.mp4 \
  --compile \
    mode=reduce-overhead \
    fullgraph=true \
    capture=true \
    recompile_perturbed_block=false \
    inductor_config='{"max_autotune": true}' \
    dynamo_config='{"cache_size_limit": 0}'

```

This configuration eliminates compilation overhead for guidance perturbations, enables aggressive kernel autotuning, and prevents Dynamo cache thrashing when sequence lengths change between generations.

## Programmatic Configuration

For custom inference scripts, instantiate `CompilationConfig` directly and apply it via `compile_transformer`:

```python
from ltx_core.model.transformer.compiling import CompilationConfig, compile_transformer
from ltx_core.model.transformer.model import LTXModel

# Load your model

model: LTXModel = load_transformer(...)

# Build maximum-performance configuration

fast_cfg = CompilationConfig(
    mode="reduce-overhead",
    backend="inductor",
    fullgraph=True,
    dynamic=None,
    inductor_config={"max_autotune": True},
    dynamo_config={"cache_size_limit": 0},
    seq_dim_dynamic=True,
    recompile_perturbed_block=False,
    capture=True,
)

# Apply compilation in-place

model = compile_transformer(model, fast_cfg)

# Inference with optimized path

video, audio = model(video_input, audio_input, perturbations=None)

```

The `compile_transformer` function mutates the model's transformer blocks according to the configuration, returning the modified model ready for accelerated inference.

## Why These Settings Deliver Peak Performance

- **`mode=reduce-overhead`** generates kernels optimized for minimal launch latency, critical for the many small operations in transformer blocks.
- **`fullgraph=true`** guarantees no Python-side graph breaks, keeping the entire forward pass in compiled Inductor code.
- **`capture=true`** moves beyond per-block compilation to a single captured CUDA graph, eliminating all Python overhead after warm-up.
- **`recompile_perturbed_block=false`** avoids creating specialized graphs for guidance-mask variations, which matters during training but wastes compilation time during pure inference.
- **`inductor_config.max_autotune`** spends upfront time searching for optimal fused kernels.
- **`dynamo_config.cache_size_limit=0`** disables caching that could retain stale shape-specialized artifacts.

## Verification and Monitoring

When running under the trainer, LTX-2 logs compilation status via the Accelerate Dynamo plugin. Check [`packages/ltx-trainer/src/ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py) (lines 804-815) to see how compilation health is reported. The trainer also warns about FSDP and `torch.compile` incompatibilities that could affect distributed inference setups.

## Summary

- Use **`CompilationConfig`** to control all `torch.compile` behavior in LTX-2
- Set **`mode=reduce-overhead`**, **`fullgraph=true`**, **`capture=true`**, and **`recompile_perturbed_block=false`** for fastest inference
- Enable **`seq_dim_dynamic`** to handle variable video lengths without recompilation
- Configure **`inductor_config`** and **`dynamo_config`** for aggressive autotuning and cache management
- Apply via **`--compile`** CLI flag or **`compile_transformer()`** programmatically

## Frequently Asked Questions

### What is the difference between `reduce-overhead` and `max-autotune` modes?

**`reduce-overhead`** minimizes kernel launch latency and compilation time, making it ideal for inference scenarios with varying inputs. **`max-autotune`** spends more time at compilation searching for optimal kernels, which can yield better steady-state performance but increases warmup time. For LTX-2 inference, start with `reduce-overhead` and benchmark against `max-autotune` on your specific hardware.

### When should I disable `capture=True`?

Disable CUDA-graph capture when you need **memory flexibility** or **dynamic control flow** that cannot be captured in a static graph. The capture step requires a fixed memory allocation for the recorded operations, which may conflict with memory-constrained environments or workflows requiring mid-inference interventions.

### Why does `recompile_perturbed_block` exist if `false` is faster for inference?

This flag exists to support **classifier-free guidance training**, where each transformer block may receive perturbed inputs requiring different optimized graphs. During training, `true` allows specialized compilation per perturbation. During pure inference with fixed guidance scales, `false` eliminates unnecessary compilation overhead.

### Can I use `torch.compile` with multi-GPU inference?

Multi-GPU inference with **FSDP (Fully Sharded Data Parallel)** has known incompatibilities with `torch.compile`. The LTX-2 trainer logs warnings about these conflicts in [`trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/trainer.py). For tensor-parallel or pipeline-parallel inference, test compilation carefully—some configurations may require disabling `fullgraph` or `capture` to maintain correctness.