How to Apply Custom LoRAs to the LTX-2 Transformer Model: A Complete Guide
LTX-2 uses the PEFT library to inject Low-Rank Adaptation (LoRA) adapters into its transformer backbone, enabling efficient fine-tuning with minimal memory overhead while supporting advanced inference features like In-Context LoRA (IC-LoRA) with automatic reference video rescaling.
The LTX-2 video generation model from Lightricks supports custom LoRAs through a PEFT-based parameter-efficient fine-tuning pipeline. This article explains how LoRA adapters integrate with LTX-2's architecture, how to train your own adapters, and how to apply them at inference time—including the IC-LoRA workflow for video-to-video transformation.
How LoRA Works in LTX-2
LTX-2 follows a standard PEFT workflow with three core stages:
LoRA Configuration
The trainer builds a LoraConfig from YAML settings in packages/ltx-trainer/src/ltx_trainer/trainer.py (lines 70-78). This configuration specifies:
lora.rank— dimension of the low-rank decompositionlora.alpha— scaling factor for adapter outputslora.target_modules— which transformer layers to adapt (typically attention and FFN)lora.dropout— regularization for adapter training
Model Wrapping
The _setup_lora() method calls get_peft_model(base_transformer, lora_config) to decorate the frozen base transformer with trainable adapter sub-layers. Only the rank-r matrices are updated during back-propagation, reducing trainable parameters by 90%+ compared to full fine-tuning.
Checkpoint Handling
LoRA weights are saved as .safetensors files during training. At inference, set_peft_model_state_dict extracts and applies only the adapter weights without reloading the full transformer.
In-Context LoRA (IC-LoRA) Architecture
LTX-2 extends standard LoRA with reference-conditioned generation. The ICLoraPipeline in packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py automatically manages:
Reference Video Conditioning
During training, ReferenceConditionConfig (in config.py) specifies a reference video with downscale and temporal scale factors. Reference latents are concatenated with target latents, teaching the LoRA to transform the reference into the output.
Automatic Metadata Resolution
When loading a LoRA for inference, the pipeline reads stored scaling factors via read_lora_reference_downscale_factor and read_lora_reference_temporal_scale_factor from iclora_utils.py. It automatically rescales input reference videos to match the training configuration.
Multi-LoRA Validation
If you supply multiple LoRAs with conflicting reference metadata, the constructor raises a clear ValueError at lines 58-65 of ic_lora.py.
Training a Custom LoRA
1. Prepare Your Configuration
Create a YAML config with LoRA parameters:
# configs/v2v_ic_lora.yaml
model:
training_mode: "lora"
lora:
rank: 64
alpha: 64
target_modules: ["to_q", "to_k", "to_v", "to_out"]
dropout: 0.0
reference:
downscale_factor: 2.0
temporal_scale_factor: 1.0
2. Run Training
# train.py
from ltx_trainer.trainer import LtxvTrainer
from ltx_trainer.config import LtxTrainerConfig
cfg = LtxTrainerConfig.parse_file("configs/v2v_ic_lora.yaml")
trainer = LtxvTrainer(cfg)
# _setup_lora() is called automatically; only adapters train
checkpoint_path, stats = trainer.train()
print(f"LoRA saved to: {checkpoint_path}")
Key implementation details:
LtxvTrainer._load_models()loads the base transformer fromltx_core.loader.registry.Registryand freezes it withrequires_grad_(False)LtxvTrainer._collect_trainable_params()gathers only adapter parameters whentraining_mode: "lora"- Checkpoints save as
checkpoints/lora_weights_step_XXXXX.safetensors
Applying LoRAs at Inference
Basic IC-LoRA Pipeline
from ltx_pipelines import ICLoraPipeline, ModelPaths, LoraPathStrengthAndSDOps
# Define model checkpoints
paths = ModelPaths(
transformer="models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
video_vae="models/video_vae.safetensors",
audio_vae="models/audio_vae.safetensors",
spatial_upsampler="models/spatial_upsampler.safetensors",
)
# Load your trained LoRA
lora = LoraPathStrengthAndSDOps(
path="checkpoints/lora_weights_step_20000.safetensors",
strength=1.0,
sdo=None, # no optimizer state needed for inference
)
# Build pipeline with automatic IC-LoRA support
pipe = ICLoraPipeline(
model_paths=paths,
spatial_upsampler_path=paths.spatial_upsampler,
loras=[lora],
device="cuda",
)
# Generate with reference video conditioning
gen = pipe(
prompt="A sunrise over a futuristic city",
seed=42,
height=512,
width=512,
num_frames=24,
frame_rate=24.0,
images=[], # no image conditioning
video_conditioning=[("reference.mp4", 1.0)], # (path, weight)
)
# Output: decoded video tensor with reference-matched conditioning
The pipeline automatically rescales reference.mp4 using the stored reference_downscale_factor metadata before latent concatenation.
Combining Multiple LoRAs
lora_style = LoraPathStrengthAndSDOps(
path="lora_cinematic.safetensors",
strength=1.0,
sdo=None,
)
lora_motion = LoraPathStrengthAndSDOps(
path="lora_slowmotion.safetensors",
strength=0.6,
sdo=None,
)
pipe = ICLoraPipeline(
model_paths=paths,
spatial_upsampler_path=paths.spatial_upsampler,
loras=[lora_style, lora_motion],
device="cuda",
)
Important: All LoRAs must share compatible reference_downscale_factor and reference_temporal_scale_factor values, or the constructor raises ValueError.
Key Source Files
| File | Purpose |
|---|---|
packages/ltx-trainer/src/ltx_trainer/trainer.py |
Core training loop; _setup_lora(), _load_lora_checkpoint() |
packages/ltx-trainer/src/ltx_trainer/config.py |
LoraConfig, ReferenceConditionConfig schemas |
packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py |
ICLoraPipeline with metadata-aware loading |
packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py |
read_lora_reference_downscale_factor(), read_lora_reference_temporal_scale_factor() |
packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py |
V2V IC-LoRA training implementation |
packages/ltx-trainer/src/ltx_trainer/training_strategies/flexible.py |
Custom conditioning combinations |
Summary
- LTX-2 applies custom LoRAs via PEFT, wrapping frozen transformers with trainable low-rank adapters
- Training requires
training_mode: "lora"in YAML;_setup_lora()handles PEFT integration automatically - IC-LoRA enables video-to-video transformation by conditioning on reference videos with automatic resolution matching
- Multi-LoRA inference validates metadata consistency across adapters, raising errors for incompatible configurations
- Checkpoints store both weights and metadata in
.safetensorsformat for portable, self-contained adapters
Frequently Asked Questions
What modules should I target with target_modules?
Target attention projections and FFN layers: ["to_q", "to_k", "to_v", "to_out", "ff.net.0.proj", "ff.net.2"]. These capture style and motion patterns most effectively. Avoid targeting all layers—this defeats the parameter efficiency purpose.
How do I choose LoRA rank and alpha?
Start with rank=64, alpha=64 for style LoRAs; increase to rank=128 for complex motion transfer. Alpha typically matches rank for 1:1 scaling. Higher ranks improve fidelity at the cost of parameter count and inference memory.
Can I use LoRAs trained on different LTX-2 versions?
LoRAs are version-specific due to architecture changes in attention patterns and latent shapes. The ICLoraPipeline validates metadata but cannot detect base model mismatches—always verify your LoRA was trained on the same transformer version you're using for inference.
Why does my reference video look wrong with IC-LoRA?
The pipeline rescales based on stored reference_downscale_factor metadata. If your LoRA was trained with downscale_factor=2 but your input differs, automatic rescaling applies. Check the LoRA metadata with read_lora_reference_downscale_factor() and ensure your training/inference configurations match.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →