FlashAttention 3 vs FlashAttention 4 vs PyTorch SDPA: Which LTX-2 Attention Backend Should You Choose?
Use FlashAttention 4 for maximum performance on modern CUDA GPUs, FlashAttention 3 for older environments, and PyTorch SDPA as the CPU/MPS fallback.
The LTX-2 video generation model from Lightricks ships with three interchangeable attention backends. Your choice directly impacts inference speed, memory efficiency, and hardware compatibility. This guide breaks down the selection logic implemented in the source code and provides concrete configuration steps.
How LTX-2 Implements Attention Backend Selection
The backend selection mechanism lives in packages/ltx-core/src/ltx_core/model/transformer/attention.py. The AttentionFunction enum defines the three supported options:
# From attention.py lines 349-350
class AttentionFunction(str, Enum):
FLASH_ATTENTION_3 = "flash_attention_3"
FLASH_ATTENTION_4 = "flash_attention_4"
SDPA = "sdpa"
When you instantiate an LTX-2 model, the code validates your chosen backend against available libraries. If you request FlashAttention 4 but lack the flash-attn 4.x package, the code raises a clear RuntimeError at lines 380-394 directing you to install the missing dependency.
The fallback chain operates as follows: FlashAttention 4 → FlashAttention 3 → PyTorch SDPA. The SDPA backend triggers automatically when CUDA is unavailable or when both FlashAttention packages are missing.
Backend Comparison: Speed, Memory, and Compatibility
FlashAttention 4: The Performance Default
Select FlashAttention 4 when running on NVIDIA GPUs with CUDA 11.8 or newer. This backend delivers the highest throughput because it uses kernels that skip attention mask processing entirely where possible.
- Speed: Fastest option for long sequences
- Memory: Most memory-efficient due to fused kernel design
- Requirements:
pip install flash-attn>=2.4.0(4.x API compatible)
FlashAttention 3: The Compatibility Option
Use FlashAttention 3 when your environment cannot install the 4.x FlashAttention package or when running on older CUDA versions (11.6-11.7). This backend provides nearly identical numerical outputs with slightly reduced kernel optimization.
The LTX-2 codebase maintains both versions explicitly to support heterogeneous deployment environments.
PyTorch SDPA: The Universal Fallback
Configure SDPA for:
- CPU inference
- Apple Silicon (MPS devices)
- AMD ROCm GPUs without FlashAttention support
- Environments where compiled CUDA extensions are prohibited
SDPA uses PyTorch's native scaled_dot_product_attention with automatic kernel selection (math, memory-efficient, or flash-optimized depending on PyTorch version). While functional, expect 2-4x slower throughput compared to FlashAttention 4 on equivalent hardware.
Configuring Your LTX-2 Attention Backend
Step 1: Install the Appropriate Package
# FlashAttention 4 (recommended)
pip install flash-attn --no-build-isolation
# Verify installation
python -c "from flash_attn import flash_attn_func; print('FA4 ready')"
# FlashAttention 3 (legacy environments)
pip install flash-attn==2.3.6 --no-build-isolation
Step 2: Set the Backend in Model Configuration
LTX-2 reads the backend from TransformerArgs, defined in transformer_args.py. Pass your selection during model initialization:
from ltx_core.model.transformer.attention import AttentionFunction
from ltx_core.model.transformer.transformer import LTXVideoTransformer
# Explicit backend selection
model = LTXVideoTransformer.from_pretrained(
"Lightricks/LTX-2",
attention_backend=AttentionFunction.FLASH_ATTENTION_4,
)
Or via configuration dictionary:
config = {
"attention_backend": "flash_attention_4", # or "flash_attention_3", "sdpa"
# ... other transformer args
}
Step 3: Validate Active Backend at Runtime
The selection logic in attention.py performs runtime verification. Add this check to confirm your configuration:
import torch
from ltx_core.model.transformer.attention import get_attention_backend
# After model initialization
backend_info = get_attention_backend(device_type="cuda")
print(f"Active backend: {backend_info.name}")
print(f"Impl type: {backend_info.implementation}")
Troubleshooting Backend Selection Errors
| Error Message | Cause | Resolution |
|---|---|---|
flash_attn is not installed |
Requested FA3/FA4 but package missing | pip install flash-attn |
flash_attn version >= 2.4.0 required |
Installed FA version too old for FA4 path | Upgrade with pip install --upgrade flash-attn |
CUDA not available |
GPU access failed | Verify torch.cuda.is_available() or switch to SDPA |
| Silent SDPA fallback | FlashAttention requested but import failed | Check stderr logs for import errors |
The error handling at lines 380-394 of attention.py surfaces specific guidance. When FlashAttention 4 is requested but unavailable, the code explicitly checks the import and raises:
# From attention.py lines 380-394
if not FLASH_ATTN_4_AVAILABLE:
raise RuntimeError(
"FlashAttention 4 was requested but flash_attn is not installed. "
"Please install with: pip install flash-attn"
)
Performance Benchmark Guidance
Based on the kernel characteristics in the LTX-2 codebase:
- FlashAttention 4: ~15-20% faster than FA3 on 4090/A100 for 720p video generation
- FlashAttention 3: ~40-50% faster than SDPA on CUDA-enabled GPUs
- PyTorch SDPA: Baseline compatibility, no compilation required
For production inference pipelines, the default FLASH_ATTENTION_4 provides optimal throughput without configuration complexity.
Summary
- FlashAttention 4 delivers maximum performance on modern CUDA GPUs and should be your default choice
- FlashAttention 3 serves as a compatibility layer for environments where FA4 packages cannot install
- PyTorch SDPA guarantees functional inference on any PyTorch-supported device without external dependencies
- The
AttentionFunctionenum inattention.pymaps your string configuration to the implementation layer - Runtime errors guide you explicitly when requested backends are unavailable
Frequently Asked Questions
What happens if I specify FlashAttention 4 but the package isn't installed?
LTX-2 raises a RuntimeError with explicit installation instructions rather than silently falling back. This prevents performance surprises in production. Install with pip install flash-attn and restart your process.
Can I switch backends without restarting Python?
No. The attention backend is resolved at model initialization time in attention.py. The get_attention_backend() function caches the implementation selection. Create a fresh model instance to change backends.
Does SDPA produce identical outputs to FlashAttention?
Numerically equivalent within floating-point tolerance, but not bit-identical. The LTX-2 training pipeline validates convergence across all three backends, so inference outputs are visually indistinguishable for video generation tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →