# How to Use `torch.compile` for Faster Inference with LTX-2

> Accelerate LTX-2 inference speed using torch.compile. Learn how to set CompilationConfig for reduced overhead and maximum GPU throughput.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-20

---

**Enable `torch.compile` in LTX-2 by setting `mode=reduce-overhead`, `fullgraph=true`, and `capture=true` in the `CompilationConfig` dataclass to eliminate per-block launch overhead and maximize GPU throughput.**

LTX-2 accelerates transformer-based video generation by compiling each transformer block with PyTorch's `torch.compile`. This article explains how to configure compilation for peak inference performance using the `CompilationConfig` system in [`packages/ltx-core/src/ltx_core/model/transformer/compiling.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/compiling.py).

## Understanding the Compilation Configuration

The **`CompilationConfig`** dataclass controls all aspects of `torch.compile` behavior in LTX-2. Here are the key fields that impact inference speed:

| Field | Purpose | Recommended Value |
|-------|---------|-------------------|
| `mode` | Optimization level (`reduce-overhead`, `max-autotune`, or `default`) | `reduce-overhead` or `max-autotune` |
| `backend` | Compilation backend | `"inductor"` |
| `fullgraph` | Force complete graph compilation (fail on breaks) | `True` |
| `dynamic` | Enable dynamic-shape tracing | `None` or `True` |
| `inductor_config` | Fine-grained Inductor settings | `{"max_autotune": True}` |
| `dynamo_config` | Fine-grained Dynamo settings | `{"cache_size_limit": 0}` |
| `seq_dim_dynamic` | Mark sequence dimension as dynamic | `True` (default) |
| `recompile_perturbed_block` | Re-compile for guidance perturbations | `False` for inference |
| `capture` | Record block-loop as CUDA graph | `True` for peak performance |

## CLI Method: Using the `--compile` Flag

The fastest way to enable `torch.compile` for LTX-2 inference is through the command-line interface. The `--compile` flag in [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) parses `KEY=VALUE` pairs into a `CompilationConfig` instance.

### Maximum Performance Command

```bash
python -m ltx_pipelines.ti2vid_one_stage \
  --checkpoint-path /path/to/model.safetensors \
  --prompt "A sunrise over mountains" \
  --output-path result.mp4 \
  --compile \
    mode=reduce-overhead \
    fullgraph=true \
    capture=true \
    recompile_perturbed_block=false \
    inductor_config='{"max_autotune": true}' \
    dynamo_config='{"cache_size_limit": 0}'

```

This configuration:
- **`mode=reduce-overhead`** — minimizes kernel launch latency
- **`fullgraph=true`** — prevents graph breaks that would fall back to eager mode
- **`capture=true`** — records a CUDA graph of the entire block-loop
- **`recompile_perturbed_block=false`** — avoids re-compilation for guidance masks (unnecessary during inference)
- **`max_autotune`** — enables aggressive kernel search
- **`cache_size_limit=0`** — prevents Dynamo cache thrashing on variable sequence lengths

## Programmatic Method: Direct `CompilationConfig` Usage

For custom scripts or integration into larger pipelines, instantiate `CompilationConfig` directly and apply it via `compile_transformer()`.

```python
from ltx_core.model.transformer.compiling import (
    CompilationConfig, 
    compile_transformer
)
from ltx_core.model.transformer.model import LTXModel

# Load your model

model: LTXModel = load_transformer(...)

# Configure for maximum inference speed

fast_cfg = CompilationConfig(
    mode="reduce-overhead",
    backend="inductor",
    fullgraph=True,
    dynamic=None,
    inductor_config={"max_autotune": True},
    dynamo_config={"cache_size_limit": 0},
    seq_dim_dynamic=True,
    recompile_perturbed_block=False,
    capture=True,
)

# Apply compilation (mutates model in-place)

model = compile_transformer(model, fast_cfg)

# Run inference

video, audio = model(video_input, audio_input, perturbations=None)

```

## How LTX-2 Compilation Works Internally

### Per-Block Compilation with `_compile_blocks`

In [`packages/ltx-core/src/ltx_core/model/transformer/compiling.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/compiling.py), the `_compile_blocks` function (lines 61-71) iterates through each transformer block and applies `torch.compile()` with the configured options:

```python

# From compiling.py lines 61-71

compiled_blocks = [
    torch.compile(
        block,
        mode=config.mode,
        backend=config.backend,
        fullgraph=config.fullgraph,
        dynamic=config.dynamic,
    )
    for block in model.blocks
]

```

### Dynamic Dimension Marking with `seq_dim_dynamic`

When `seq_dim_dynamic=True`, the helper `_mark_seq_dim_dynamic` (lines 67-95) calls `torch._dynamo.mark_dynamic()` on the sequence dimension. This allows a single compiled artifact to handle variable sequence lengths without re-compilation.

### CUDA Graph Capture with `capture=True`

If `capture=True`, the block-loop is wrapped by `CudaGraphRunner` (lines 78-84). This records the compiled operations into a CUDA graph on the first forward pass, then replays the graph on subsequent calls—eliminating CPU overhead entirely after warm-up.

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`packages/ltx-core/src/ltx_core/model/transformer/compiling.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/compiling.py) | `CompilationConfig` definition, `_compile_blocks`, `_mark_seq_dim_dynamic`, CUDA graph integration |
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) | CLI `--compile` flag parsing (lines 76-105) |
| [`packages/ltx-trainer/src/ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py) | Compilation status logging and FSDP compatibility warnings (lines 804-815) |
| [`packages/ltx-core/src/ltx_core/model/transformer/model.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/model.py) | `LTXModel` class with transformer blocks |

## Troubleshooting Common Issues

**Compilation fails with graph breaks**

Set `fullgraph=False` to allow partial compilation, or trace the error to identify unsupported operations (typically custom CUDA kernels or complex control flow).

**Variable sequence lengths cause re-compilation**

Ensure `seq_dim_dynamic=True` and `dynamic=None` or `dynamic=True`. Verify that `torch._dynamo.mark_dynamic` is being applied correctly in `_mark_seq_dim_dynamic`.

**Memory usage spikes during first inference**

This is expected warm-up behavior. The CUDA graph capture (`capture=True`) and Inductor autotuning both execute during the first forward pass. Subsequent passes use the cached artifacts.

## Summary

- **`torch.compile`** in LTX-2 is configured through the **`CompilationConfig`** dataclass in [`compiling.py`](https://github.com/Lightricks/LTX-2/blob/main/compiling.py)
- **CLI**: Use `--compile mode=reduce-overhead fullgraph=true capture=true` for one-line activation
- **Programmatic**: Build a `CompilationConfig` and pass to `compile_transformer(model, config)`
- **Peak performance requires**: `mode=reduce-overhead` or `max-autotune`, `fullgraph=true`, `capture=true`, and `recompile_perturbed_block=false`
- **Dynamic shapes**: Handled automatically via `seq_dim_dynamic=True` and `torch._dynamo.mark_dynamic`

## Frequently Asked Questions

### What is the fastest `torch.compile` mode for LTX-2 inference?

**`reduce-overhead`** provides the best balance of compilation speed and runtime performance for most GPUs. **`max-autotune`** can yield faster kernels but increases compilation time significantly. Both outperform `default` mode for video generation workloads where the same shapes repeat across many frames.

### Why does LTX-2 compile per-block instead of the full model?

Per-block compilation in `_compile_blocks` allows **finer-grained control** over which components are optimized and enables **selective re-compilation** when guidance perturbations change. The optional `capture=true` flag then fuses the compiled blocks into a single CUDA graph for minimal launch overhead.

### Does `capture=true` work with variable video resolutions?

Yes, but only if `seq_dim_dynamic=True` (the default). The `_mark_seq_dim_dynamic` helper marks the sequence dimension as dynamic, allowing the captured CUDA graph to handle varying lengths. However, **spatial dimensions** (height × width) typically require separate compiled artifacts unless they remain constant across generations.

### Can I use `torch.compile` with multi-GPU training in LTX-2?

**Not directly with FSDP.** The trainer in [`packages/ltx-trainer/src/ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py) logs a warning about `torch.compile` incompatibilities with FSDP at lines 804-815. For multi-GPU inference, consider **data-parallel** approaches or compile on a single GPU before distributing weights.