LoRA vs. IC-LoRA vs. Distilled LoRA vs. Detailing IC-LoRA in LTX-2: A Complete Technical Comparison
In LTX-2, these four adapter types serve distinct purposes: regular LoRA enables lightweight fine-tuning; IC-LoRA adds in-context reference conditioning; Distilled LoRA accelerates high-resolution refinement; and Detailing IC-LoRA provides pixel-level spatial upsampling for maximum sharpness.
LTX-2 from Lightricks introduces multiple LoRA variants optimized for different video generation workflows. Understanding when to use each adapter type—and how they interact in multi-stage pipelines—lets you optimize both training efficiency and output quality. This guide breaks down the implementation details found directly in the LTX-2 source code.
Regular LoRA: Lightweight Adapter Fine-Tuning
Regular LoRA (Low-Rank Adaptation) is the foundation: small, trainable rank decomposition matrices inserted into frozen transformer weights.
In packages/ltx-trainer/src/ltx_trainer/trainer.py, the _configure_lora method instantiates PEFT's LoraAdapter at line 471, targeting specific modules via the target_modules configuration. Only these adapter weights update during training—the base LTX-2 parameters remain fixed.
This design enables fast experimentation with minimal storage overhead (typically megabytes versus gigabytes for full fine-tunes). Use regular LoRA for style transfer, inpainting extensions, or any scenario requiring quick, swappable adaptations.
from ltx_core.loader.single_gpu_model_builder import SingleGPUModelBuilder
builder = SingleGPUModelBuilder(
checkpoint_path="model.safetensors",
loras=[("lora.safetensors", 0.8, None)], # (path, strength, optional sd_ops)
)
model = builder.build()
Source: packages/ltx-core/README.md lines 64–80
IC-LoRA: In-Context Conditioning for Reference-Driven Generation
IC-LoRA extends standard LoRA with reference-conditioning capability. Instead of text-only guidance, the model processes an encoded reference video or audio as an additional context token.
The training implementation lives in packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py. Here, reference latents are concatenated at timestep 0 before the target sequence, with configurable downsampling via reference_downscale_factor. This architecture enables:
- Video-to-video transformation (style transfer, weather changes)
- Audio-to-audio modification
- Joint audio-video synchronization tasks
At inference, ICLoraPipeline handles the concatenation automatically—no model architecture changes required beyond loading the adapter.
ltx-pipelines run ic_lora \
--checkpoint-path model.safetensors \
--lora path/to/ic_lora.safetensors 0.75 \
--reference-video ref.mp4 \
--prompt "turn day into night"
Source: packages/ltx-trainer/docs/training-modes.md lines 119–142
Distilled LoRA: Efficient High-Resolution Refinement
Distilled LoRA operates in the second stage of two-stage pipelines, paired with a distilled checkpoint—a compact model variant using a reduced sigma schedule for faster sampling.
Key characteristics from packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 1151–1183):
| Aspect | Configuration |
|---|---|
| Checkpoint loading | --distilled-checkpoint-path |
| LoRA loading | --distilled-lora PATH STRENGTH |
| Sampling | Shorter denoising schedule, no CFG |
| Typical pipelines | ti2vid_two_stages, ti2vid_two_stages_hq, DFRPipeline |
The distilled LoRA refines upsampled outputs efficiently, avoiding the computational cost of running the full model at high resolution.
ltx-pipelines run ti2vid_two_stages \
--checkpoint-path model.safetensors \
--distilled-checkpoint-path model_distilled.safetensors \
--distilled-lora distilled_lora.safetensors 0.5 \
--prompt "a forest in autumn"
Stage 1: Full model at low resolution. Stage 2: 2× upsample + distilled LoRA refinement.
Detailing IC-LoRA: Pixel-Level Spatial Upsampling
Detailing IC-LoRA provides an optional ×2 spatial detailing pass after the distilled LoRA stage. Unlike the general refinement of Distilled LoRA, this variant specifically targets pixel-level sharpness at full resolution.
Invocation requires the --detailing-lora flag in supported pipelines. In packages/ltx-pipelines/src/ltx_pipelines/dfr_pipeline.py at line 576, the argument accepts a path and optional strength. Internally, the pipeline creates a stage_detailing phase with _detailing_downscale_factor set to 2, running the IC-LoRA adapter at full resolution for spatial enhancement.
ltx-pipelines run dfr \
--checkpoint-path model.safetensors \
--distilled-checkpoint-path model_distilled.safetensors \
--distilled-lora distilled_lora.safetensors 0.5 \
--detailing-lora detailing_ic_lora.safetensors 0.7 \
--prompt "high-detail medieval city at sunset"
This three-stage flow (base generation → distilled refinement → spatial detailing) maximizes both efficiency and output fidelity.
Quick Comparison: When to Use Each LoRA Type
| Type | Training Target | Inference Context | Best For |
|---|---|---|---|
| LoRA | Style/concept adaptation | Single-stage generation | Lightweight, swappable modifications |
| IC-LoRA | Reference-conditioned transformation | Video-to-video, audio tasks | Maintaining structure while altering content |
| Distilled LoRA | High-res refinement with compact model | Two-stage pipeline Stage 2 | Fast, efficient quality improvement |
| Detailing IC-LoRA | Spatial upsampling | Optional Stage 2+ in DFR | Maximum fine-grained sharpness |
Implementation Reference: Key Source Files
| Component | File Path | Critical Function/Line |
|---|---|---|
| LoRA configuration | packages/ltx-trainer/src/ltx_trainer/trainer.py |
_configure_lora (line 471) |
| IC-LoRA training | packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py |
Reference latent handling (lines 1–78) |
| IC-LoRA inference | packages/ltx-pipelines/src/ltx_pipelines/ic_lora_pipeline.py |
Pipeline concatenation logic |
| Distilled LoRA args | packages/ltx-pipelines/src/ltx_pipelines/utils/args.py |
--distilled-lora definition (lines 1151–1183) |
| Two-stage pipelines | packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py |
Stage orchestration |
| DFR detailing | packages/ltx-pipelines/src/ltx_pipelines/dfr_pipeline.py |
--detailing-lora flag (line 576) |
| LoRA loading/fusion | packages/ltx-core/src/ltx_core/loader/single_gpu_model_builder.py |
Weight merging at build time |
Summary
-
Regular LoRA enables efficient, frozen-weight fine-tuning through rank decomposition adapters configured in
trainer.py. -
IC-LoRA adds in-context reference conditioning, concatenating reference latents for video-to-video and audio-to-audio tasks via
video_to_video.pytraining strategies. -
Distilled LoRA pairs with compact distilled checkpoints in two-stage pipelines, providing fast high-resolution refinement without full model cost.
-
Detailing IC-LoRA offers optional ×2 spatial upsampling as a final detailing pass in DFR pipelines, triggered by
--detailing-lorafor maximum sharpness.
Frequently Asked Questions
Can IC-LoRA and regular LoRA be used together?
Yes. IC-LoRA is fundamentally a LoRA with additional training objectives for reference conditioning. You can load both types simultaneously through the loras parameter in SingleGPUModelBuilder, applying strength multipliers independently. The IC-LoRA's reference-handling logic activates only when the pipeline provides reference inputs.
Why does Distilled LoRA require a separate checkpoint?
The distilled checkpoint uses a reduced noise schedule and compressed architecture that enables faster sampling. As defined in utils/args.py, the --distilled-checkpoint-path loads this compact model while --distilled-lora provides task-specific adaptations. This separation allows the same distilled base to serve multiple specialized LoRAs without redundant storage.
Is Detailing IC-LoRA always beneficial, or does it add artifacts?
Detailing IC-LoRA applies aggressive spatial upsampling that can amplify noise or over-sharpen if the preceding stages produce inconsistent outputs. The dfr_pipeline.py implementation makes this stage optional via --detailing-lora precisely because it trades compute for sharpness—test with your specific content to determine optimal strength values, typically 0.6–0.8.
How do I train my own IC-LoRA for a custom video transformation?
Configure the video_to_video training strategy in packages/ltx-trainer, specifying reference video paths in your dataset. The strategy automatically handles ref_latents preparation and reference_downscale_factor application. Output adapters load directly into ICLoraPipeline without additional conversion.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →