How to Use `torch.compile` for Faster Inference with LTX-2

Enable torch.compile in LTX-2 by setting mode=reduce-overhead, fullgraph=true, and capture=true in the CompilationConfig dataclass to eliminate per-block launch overhead and maximize GPU throughput.

LTX-2 accelerates transformer-based video generation by compiling each transformer block with PyTorch's torch.compile. This article explains how to configure compilation for peak inference performance using the CompilationConfig system in packages/ltx-core/src/ltx_core/model/transformer/compiling.py.

Understanding the Compilation Configuration

The CompilationConfig dataclass controls all aspects of torch.compile behavior in LTX-2. Here are the key fields that impact inference speed:

Field Purpose Recommended Value
mode Optimization level (reduce-overhead, max-autotune, or default) reduce-overhead or max-autotune
backend Compilation backend "inductor"
fullgraph Force complete graph compilation (fail on breaks) True
dynamic Enable dynamic-shape tracing None or True
inductor_config Fine-grained Inductor settings {"max_autotune": True}
dynamo_config Fine-grained Dynamo settings {"cache_size_limit": 0}
seq_dim_dynamic Mark sequence dimension as dynamic True (default)
recompile_perturbed_block Re-compile for guidance perturbations False for inference
capture Record block-loop as CUDA graph True for peak performance

CLI Method: Using the --compile Flag

The fastest way to enable torch.compile for LTX-2 inference is through the command-line interface. The --compile flag in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py parses KEY=VALUE pairs into a CompilationConfig instance.

Maximum Performance Command

python -m ltx_pipelines.ti2vid_one_stage \
  --checkpoint-path /path/to/model.safetensors \
  --prompt "A sunrise over mountains" \
  --output-path result.mp4 \
  --compile \
    mode=reduce-overhead \
    fullgraph=true \
    capture=true \
    recompile_perturbed_block=false \
    inductor_config='{"max_autotune": true}' \
    dynamo_config='{"cache_size_limit": 0}'

This configuration:

  • mode=reduce-overhead — minimizes kernel launch latency
  • fullgraph=true — prevents graph breaks that would fall back to eager mode
  • capture=true — records a CUDA graph of the entire block-loop
  • recompile_perturbed_block=false — avoids re-compilation for guidance masks (unnecessary during inference)
  • max_autotune — enables aggressive kernel search
  • cache_size_limit=0 — prevents Dynamo cache thrashing on variable sequence lengths

Programmatic Method: Direct CompilationConfig Usage

For custom scripts or integration into larger pipelines, instantiate CompilationConfig directly and apply it via compile_transformer().

from ltx_core.model.transformer.compiling import (
    CompilationConfig, 
    compile_transformer
)
from ltx_core.model.transformer.model import LTXModel

# Load your model

model: LTXModel = load_transformer(...)

# Configure for maximum inference speed

fast_cfg = CompilationConfig(
    mode="reduce-overhead",
    backend="inductor",
    fullgraph=True,
    dynamic=None,
    inductor_config={"max_autotune": True},
    dynamo_config={"cache_size_limit": 0},
    seq_dim_dynamic=True,
    recompile_perturbed_block=False,
    capture=True,
)

# Apply compilation (mutates model in-place)

model = compile_transformer(model, fast_cfg)

# Run inference

video, audio = model(video_input, audio_input, perturbations=None)

How LTX-2 Compilation Works Internally

Per-Block Compilation with _compile_blocks

In packages/ltx-core/src/ltx_core/model/transformer/compiling.py, the _compile_blocks function (lines 61-71) iterates through each transformer block and applies torch.compile() with the configured options:


# From compiling.py lines 61-71

compiled_blocks = [
    torch.compile(
        block,
        mode=config.mode,
        backend=config.backend,
        fullgraph=config.fullgraph,
        dynamic=config.dynamic,
    )
    for block in model.blocks
]

Dynamic Dimension Marking with seq_dim_dynamic

When seq_dim_dynamic=True, the helper _mark_seq_dim_dynamic (lines 67-95) calls torch._dynamo.mark_dynamic() on the sequence dimension. This allows a single compiled artifact to handle variable sequence lengths without re-compilation.

CUDA Graph Capture with capture=True

If capture=True, the block-loop is wrapped by CudaGraphRunner (lines 78-84). This records the compiled operations into a CUDA graph on the first forward pass, then replays the graph on subsequent calls—eliminating CPU overhead entirely after warm-up.

Key Source Files Reference

File Purpose
packages/ltx-core/src/ltx_core/model/transformer/compiling.py CompilationConfig definition, _compile_blocks, _mark_seq_dim_dynamic, CUDA graph integration
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py CLI --compile flag parsing (lines 76-105)
packages/ltx-trainer/src/ltx_trainer/trainer.py Compilation status logging and FSDP compatibility warnings (lines 804-815)
packages/ltx-core/src/ltx_core/model/transformer/model.py LTXModel class with transformer blocks

Troubleshooting Common Issues

Compilation fails with graph breaks

Set fullgraph=False to allow partial compilation, or trace the error to identify unsupported operations (typically custom CUDA kernels or complex control flow).

Variable sequence lengths cause re-compilation

Ensure seq_dim_dynamic=True and dynamic=None or dynamic=True. Verify that torch._dynamo.mark_dynamic is being applied correctly in _mark_seq_dim_dynamic.

Memory usage spikes during first inference

This is expected warm-up behavior. The CUDA graph capture (capture=True) and Inductor autotuning both execute during the first forward pass. Subsequent passes use the cached artifacts.

Summary

  • torch.compile in LTX-2 is configured through the CompilationConfig dataclass in compiling.py
  • CLI: Use --compile mode=reduce-overhead fullgraph=true capture=true for one-line activation
  • Programmatic: Build a CompilationConfig and pass to compile_transformer(model, config)
  • Peak performance requires: mode=reduce-overhead or max-autotune, fullgraph=true, capture=true, and recompile_perturbed_block=false
  • Dynamic shapes: Handled automatically via seq_dim_dynamic=True and torch._dynamo.mark_dynamic

Frequently Asked Questions

What is the fastest torch.compile mode for LTX-2 inference?

reduce-overhead provides the best balance of compilation speed and runtime performance for most GPUs. max-autotune can yield faster kernels but increases compilation time significantly. Both outperform default mode for video generation workloads where the same shapes repeat across many frames.

Why does LTX-2 compile per-block instead of the full model?

Per-block compilation in _compile_blocks allows finer-grained control over which components are optimized and enables selective re-compilation when guidance perturbations change. The optional capture=true flag then fuses the compiled blocks into a single CUDA graph for minimal launch overhead.

Does capture=true work with variable video resolutions?

Yes, but only if seq_dim_dynamic=True (the default). The _mark_seq_dim_dynamic helper marks the sequence dimension as dynamic, allowing the captured CUDA graph to handle varying lengths. However, spatial dimensions (height × width) typically require separate compiled artifacts unless they remain constant across generations.

Can I use torch.compile with multi-GPU training in LTX-2?

Not directly with FSDP. The trainer in packages/ltx-trainer/src/ltx_trainer/trainer.py logs a warning about torch.compile incompatibilities with FSDP at lines 804-815. For multi-GPU inference, consider data-parallel approaches or compile on a single GPU before distributing weights.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →