How to Train Custom LoRAs Using ltx-trainer: A Complete Guide for LTX Video Diffusion Models

Train custom LoRAs for LTX audio-video models by preprocessing data into VAE latents, writing a YAML config with training_mode: lora, and running scripts/train.py.

ltx-trainer is the official training harness for Lightricks' LTX family of diffusion models, supporting both full fine-tuning and parameter-efficient LoRA adaptation. This guide walks through the complete workflow to train custom LoRAs using ltx-trainer, from data preparation to launch.


Data Preparation: Convert Raw Media to VAE Latents

The ltx-trainer pipeline requires pre-encoded VAE latents, not raw video files. You must run two preprocessing scripts before training.

Generate Reference Latents for IC-LoRA (Video-to-Video)

For IC-LoRA (image-conditioned video generation), first create reference videos at lower resolution:

python packages/ltx-trainer/scripts/compute_reference.py \
    --data-root /data/ltx_preproc \
    --output-dir /data/ltx_preproc/reference_videos \
    --downscale-factor 2

This script (ltx-trainer/scripts/compute_reference.py) reads training videos, downscales them by the specified factor, encodes with the VAE, and stores latents under reference_latents/ [6†L1-L4]. These reference latents are later concatenated with noisy target latents during training [6†L25-L33].

Process the Full Dataset

Convert your raw dataset into the required latent format:

python packages/ltx-trainer/scripts/process_dataset.py \
    --manifest my_dataset.tsv \
    --data-root /data/ltx_preproc \
    --lora-trigger "[lora]"

The process_dataset.py script writes three directories under preprocessed_data_root [5†L119-L131]:

  • latents/ — VAE-encoded target videos
  • conditions/ — text prompts and other conditioning
  • reference_latents/ — reference latents for IC-LoRA (if applicable)

Configuration: Write a Valid LoRA YAML Config

ltx-trainer uses Pydantic-based validation. Your config must include model.training_mode: "lora" and a non-null lora: block, or LtxTrainerConfig will raise a validation error [4†L45-L51][4†L82-L88].

Minimal LoRA Configuration

model:
  model_path: /path/to/ltx2_checkpt.safetensors
  training_mode: lora

lora:
  rank: 64
  alpha: 64
  dropout: 0.0
  target_modules: ["to_k", "to_q", "to_v", "to_out.0"]

data:
  preprocessed_data_root: /data/ltx_preproc

training_strategy:
  name: video_to_video
  reference_latents_dir: reference_latents

validation:
  samples:
    - prompt: "A sunrise over a mountain lake"
      conditions:
        - type: reference
          video: /data/ltx_preproc/reference_videos/001.mp4
          downscale_factor: 2

Key Config Parameters

Parameter Purpose Source
model.training_mode Must be "lora" to enable adapter training LtxTrainerConfig validation [4†L45-L51]
lora.rank / lora.alpha LoRA rank and scaling factor LoraConfig definition [4†L82-L108]
lora.target_modules Transformer layers to inject adapters (attention projections) PEFT integration [5†L52-L58]
training_strategy.name Use video_to_video for IC-LoRA VideoToVideoStrategy [6†L25-L33]

The VideoToVideoStrategy implements IC-LoRA by concatenating clean reference latents with noisy target latents and building per-token conditioning masks [6†L58-L70][6†L85-L91].


Launch Training: Run the CLI Entry Point

With data and config ready, start training:

python packages/ltx-trainer/scripts/train.py configs/my_lora.yaml

For multi-GPU training, use:

accelerate launch packages/ltx-trainer/scripts/train.py configs/my_lora.yaml

What Happens Internally

The train.py script [5†L23-L30][5†L58-L60]:

  1. Parses the YAML into LtxTrainerConfig
  2. Constructs LtxvTrainer with your configuration
  3. Calls self._add_lora_layers() when training_mode == "lora" to inject PEFT adapters [5†L52-L58]
  4. Begins the diffusion training loop

The trainer optimizes only LoRA parameters while keeping the base model frozen, dramatically reducing memory requirements compared to full fine-tuning.


Training Strategies: Choose the Right Mode for Your Use Case

Strategy Use Case Implementation
Standard LoRA Text-to-video generation Base training loop with LoRA adapters on transformer attention layers
IC-LoRA (Video-to-Video) Video-to-video transformation using reference conditioning VideoToVideoStrategy concatenates reference and target latents [6†L25-L33]

IC-LoRA requires the reference condition type in validation samples and proper reference_latents_dir configuration. The strategy automatically handles the latent concatenation and mask construction [6†L58-L70].


Summary

  • Data prep is mandatory: Run compute_reference.py (for IC-LoRA) and process_dataset.py to generate VAE latents before training [6†L1-L4][5†L119-L131].
  • Config validation is strict: LtxTrainerConfig requires training_mode: "lora" and a complete lora: block [4†L45-L51][4†L82-L88].
  • PEFT integration: LtxvTrainer._add_lora_layers() injects adapters into target transformer modules when LoRA mode is active [5†L52-L58].
  • IC-LoRA via strategy: Use training_strategy.name: video_to_video with reference latents for video-to-video training [6†L25-L33].
  • Entry point: scripts/train.py is the CLI wrapper that orchestrates the full pipeline [5†L23-L30].

Frequently Asked Questions

What files does ltx-trainer require as input?

ltx-trainer requires preprocessed VAE latents, not raw videos. Run process_dataset.py to generate latents/, conditions/, and optionally reference_latents/ directories. The trainer reads these cached tensors directly during training, avoiding repeated VAE encoding [5†L119-L131].

How do I enable LoRA training instead of full fine-tuning?

Set model.training_mode: "lora" and provide a lora: block with rank, alpha, dropout, and target_modules. The LtxTrainerConfig validator enforces this schema and will reject configs missing either requirement [4†L45-L51][4†L82-L88].

What is IC-LoRA and how do I use it with ltx-trainer?

IC-LoRA (Image-Conditioned LoRA) enables video-to-video generation by conditioning on a reference video. Use training_strategy.name: video_to_video, generate reference latents with compute_reference.py, and include reference conditions in your validation config. The VideoToVideoStrategy handles latent concatenation and masking automatically [6†L25-L33][6†L58-L70].

Can I train on multiple GPUs?

Yes. Use accelerate launch with scripts/train.py for distributed training. The LtxvTrainer uses Accelerate for multi-GPU orchestration and gradient synchronization [5†L23-L30].

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →