How to Train Custom LoRAs Using ltx-trainer: A Complete Guide for LTX Video Diffusion Models
Train custom LoRAs for LTX audio-video models by preprocessing data into VAE latents, writing a YAML config with training_mode: lora, and running scripts/train.py.
ltx-trainer is the official training harness for Lightricks' LTX family of diffusion models, supporting both full fine-tuning and parameter-efficient LoRA adaptation. This guide walks through the complete workflow to train custom LoRAs using ltx-trainer, from data preparation to launch.
Data Preparation: Convert Raw Media to VAE Latents
The ltx-trainer pipeline requires pre-encoded VAE latents, not raw video files. You must run two preprocessing scripts before training.
Generate Reference Latents for IC-LoRA (Video-to-Video)
For IC-LoRA (image-conditioned video generation), first create reference videos at lower resolution:
python packages/ltx-trainer/scripts/compute_reference.py \
--data-root /data/ltx_preproc \
--output-dir /data/ltx_preproc/reference_videos \
--downscale-factor 2
This script (ltx-trainer/scripts/compute_reference.py) reads training videos, downscales them by the specified factor, encodes with the VAE, and stores latents under reference_latents/ [6†L1-L4]. These reference latents are later concatenated with noisy target latents during training [6†L25-L33].
Process the Full Dataset
Convert your raw dataset into the required latent format:
python packages/ltx-trainer/scripts/process_dataset.py \
--manifest my_dataset.tsv \
--data-root /data/ltx_preproc \
--lora-trigger "[lora]"
The process_dataset.py script writes three directories under preprocessed_data_root [5†L119-L131]:
latents/— VAE-encoded target videosconditions/— text prompts and other conditioningreference_latents/— reference latents for IC-LoRA (if applicable)
Configuration: Write a Valid LoRA YAML Config
ltx-trainer uses Pydantic-based validation. Your config must include model.training_mode: "lora" and a non-null lora: block, or LtxTrainerConfig will raise a validation error [4†L45-L51][4†L82-L88].
Minimal LoRA Configuration
model:
model_path: /path/to/ltx2_checkpt.safetensors
training_mode: lora
lora:
rank: 64
alpha: 64
dropout: 0.0
target_modules: ["to_k", "to_q", "to_v", "to_out.0"]
data:
preprocessed_data_root: /data/ltx_preproc
training_strategy:
name: video_to_video
reference_latents_dir: reference_latents
validation:
samples:
- prompt: "A sunrise over a mountain lake"
conditions:
- type: reference
video: /data/ltx_preproc/reference_videos/001.mp4
downscale_factor: 2
Key Config Parameters
| Parameter | Purpose | Source |
|---|---|---|
model.training_mode |
Must be "lora" to enable adapter training |
LtxTrainerConfig validation [4†L45-L51] |
lora.rank / lora.alpha |
LoRA rank and scaling factor | LoraConfig definition [4†L82-L108] |
lora.target_modules |
Transformer layers to inject adapters (attention projections) | PEFT integration [5†L52-L58] |
training_strategy.name |
Use video_to_video for IC-LoRA |
VideoToVideoStrategy [6†L25-L33] |
The VideoToVideoStrategy implements IC-LoRA by concatenating clean reference latents with noisy target latents and building per-token conditioning masks [6†L58-L70][6†L85-L91].
Launch Training: Run the CLI Entry Point
With data and config ready, start training:
python packages/ltx-trainer/scripts/train.py configs/my_lora.yaml
For multi-GPU training, use:
accelerate launch packages/ltx-trainer/scripts/train.py configs/my_lora.yaml
What Happens Internally
The train.py script [5†L23-L30][5†L58-L60]:
- Parses the YAML into
LtxTrainerConfig - Constructs
LtxvTrainerwith your configuration - Calls
self._add_lora_layers()whentraining_mode == "lora"to inject PEFT adapters [5†L52-L58] - Begins the diffusion training loop
The trainer optimizes only LoRA parameters while keeping the base model frozen, dramatically reducing memory requirements compared to full fine-tuning.
Training Strategies: Choose the Right Mode for Your Use Case
| Strategy | Use Case | Implementation |
|---|---|---|
| Standard LoRA | Text-to-video generation | Base training loop with LoRA adapters on transformer attention layers |
| IC-LoRA (Video-to-Video) | Video-to-video transformation using reference conditioning | VideoToVideoStrategy concatenates reference and target latents [6†L25-L33] |
IC-LoRA requires the reference condition type in validation samples and proper reference_latents_dir configuration. The strategy automatically handles the latent concatenation and mask construction [6†L58-L70].
Summary
- Data prep is mandatory: Run
compute_reference.py(for IC-LoRA) andprocess_dataset.pyto generate VAE latents before training [6†L1-L4][5†L119-L131]. - Config validation is strict:
LtxTrainerConfigrequirestraining_mode: "lora"and a completelora:block [4†L45-L51][4†L82-L88]. - PEFT integration:
LtxvTrainer._add_lora_layers()injects adapters into target transformer modules when LoRA mode is active [5†L52-L58]. - IC-LoRA via strategy: Use
training_strategy.name: video_to_videowith reference latents for video-to-video training [6†L25-L33]. - Entry point:
scripts/train.pyis the CLI wrapper that orchestrates the full pipeline [5†L23-L30].
Frequently Asked Questions
What files does ltx-trainer require as input?
ltx-trainer requires preprocessed VAE latents, not raw videos. Run process_dataset.py to generate latents/, conditions/, and optionally reference_latents/ directories. The trainer reads these cached tensors directly during training, avoiding repeated VAE encoding [5†L119-L131].
How do I enable LoRA training instead of full fine-tuning?
Set model.training_mode: "lora" and provide a lora: block with rank, alpha, dropout, and target_modules. The LtxTrainerConfig validator enforces this schema and will reject configs missing either requirement [4†L45-L51][4†L82-L88].
What is IC-LoRA and how do I use it with ltx-trainer?
IC-LoRA (Image-Conditioned LoRA) enables video-to-video generation by conditioning on a reference video. Use training_strategy.name: video_to_video, generate reference latents with compute_reference.py, and include reference conditions in your validation config. The VideoToVideoStrategy handles latent concatenation and masking automatically [6†L25-L33][6†L58-L70].
Can I train on multiple GPUs?
Yes. Use accelerate launch with scripts/train.py for distributed training. The LtxvTrainer uses Accelerate for multi-GPU orchestration and gradient synchronization [5†L23-L30].
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →