How to Use and Fuse Multiple LoRA Adapters in LTX-2
LTX-2 fuses multiple LoRA adapters by loading each checkpoint with a strength factor, aggregating their low-rank deltas through aggregate_lora_products, and applying the combined update via FuseRule policies, all orchestrated through the SingleGPUModelBuilder API or the lower-level apply_loras function.
The LTX-2 video generation framework from Lightricks provides a sophisticated adapter fusion system that allows you to combine multiple LoRA (Low-Rank Adaptation) checkpoints in a single inference pass. Understanding how to use and fuse multiple LoRA adapters in LTX-2 enables efficient style mixing and concept combination without loading separate model instances. This guide examines the core implementation in packages/ltx-core/src/ltx_core/loader/fuse_loras.py and demonstrates three practical approaches to adapter fusion.
Understanding the LoRA Fusion Architecture
The fusion workflow operates through three distinct phases. First, LoraStateDictWithStrength pairs each LoRA checkpoint with a scaling factor. Second, aggregate_lora_products computes the combined delta by summing the matrix products of all adapters. Third, a FuseRule merges the aggregated delta into the base weights.
At the heart of the system lies the LoraProduct class defined in packages/ltx-core/src/ltx_core/loader/fuse_loras.py, which holds the A and B tensors alongside the strength multiplier. The aggregate_lora_products function efficiently computes the sum of (B * strength) @ A across all registered adapters, ensuring memory efficiency by processing contributions without simultaneously allocating all LoRA tensors on GPU.
Method 1: Using the SingleGPUModelBuilder API (Recommended)
The SingleGPUModelBuilder class provides an immutable builder pattern that automates the entire fusion pipeline. Located in packages/ltx-core/src/ltx_core/loader/single_gpu_model_builder.py, this API handles device placement, loading, and rule application automatically.
The .lora() method registers each adapter path and strength, returning a new builder instance. When you invoke .build(), the system loads the base model, constructs LoraStateDictWithStrength objects for each adapter, and calls apply_loras with the default bf16_fuse_rule.
from pathlib import Path
import torch
from ltx_core.loader.single_gpu_model_builder import SingleGPUModelBuilder
from ltx_core.model.model_protocol import ModelConfigurator
# Define the base model configurator (replace with the actual model you need)
class MyModelConfigurator(ModelConfigurator):
... # model-specific implementation (omitted for brevity)
# Instantiate the builder with the base checkpoint
builder = SingleGPUModelBuilder(
model_class_configurator=MyModelConfigurator,
model_path=Path("checkpoints/base_model.safetensors"),
)
# Register several LoRA adapters (path + strength)
builder = builder.lora("loras/style_a.safetensors", strength=0.7, sd_ops=None)
builder = builder.lora("loras/style_b.safetensors", strength=0.3, sd_ops=None)
# Build the model – LoRAs are fused automatically on the GPU
model = builder.build(device=torch.device("cuda"), dtype=torch.float16)
# model is ready for inference
Under the hood, build() invokes _load_model_weights, which loads LoRA checkpoints on the CPU (configurable via lora_load_device), prepares the state dictionaries, and executes the fusion before moving the final weights to the target device.
Method 2: Manual Fusion with apply_loras
For scenarios requiring custom preprocessing or explicit control over the state dictionary lifecycle, use the apply_loras function directly from packages/ltx-core/src/ltx_core/loader/fuse_loras.py. This approach requires manually constructing LoraStateDictWithStrength objects and selecting a fuse rule.
import torch
from ltx_core.loader.fuse_loras import apply_loras, bf16_fuse_rule
from ltx_core.loader.primitives import LoraStateDictWithStrength
from ltx_core.loader.helpers import load_state_dict
# Load the base model state dict (any loader that implements StateDictLoader works)
base_sd = load_state_dict(
paths=["checkpoints/base_model.safetensors"],
loader=SafetensorsModelStateDictLoader(),
registry=DummyRegistry(),
device=torch.device("cpu"),
)
# Load two LoRA adapters
lora_paths = ["loras/style_a.safetensors", "loras/style_b.safetensors"]
lora_strengths = [0.7, 0.3]
lora_sd_and_strengths = [
LoraStateDictWithStrength(
sd=load_state_dict([p], SafetensorsModelStateDictLoader(), DummyRegistry(), torch.device("cpu")),
strength=s,
)
for p, s in zip(lora_paths, lora_strengths)
]
# Fuse – the result is a new StateDict with the LoRA deltas applied
fused_sd = apply_loras(
model_sd=base_sd,
lora_sd_and_strengths=lora_sd_and_strengths,
fuse_rule=bf16_fuse_rule,
)
# Convert back to a model (pseudo-code, depends on your model class)
model = MyModelConfigurator().create_model()
model.load_state_dict(fused_sd.sd)
model.to("cuda")
This method exposes the full flexibility of the fusion engine, allowing you to inspect intermediate state dictionaries or modify the aggregation process before final weight application.
Method 3: Implementing Custom Fuse Rules
Advanced use cases may require non-standard precision or scaling strategies. The FuseRule callable encapsulates the policy for merging deltas into weights. You can implement custom rules by defining a function that accepts (key, weight, deltas, model_sd) and returns a dictionary of updated tensors.
from ltx_core.loader.fuse_loras import FuseRule, FuseFn
import torch
def fp8_fuse(key: str, weight: torch.Tensor, deltas: torch.Tensor, model_sd):
# Example: quantize the fused weight to fp8 (pseudo-code)
new_weight = (weight + deltas).to(torch.float8_e4m3fn)
return {key: new_weight}
fp8_rule = FuseRule(aggregation_dtype=torch.bfloat16, fuse_fn=fp8_fuse)
# Use the rule with the builder or apply_loras
model = builder.with_fuse_rule(fp8_rule).build(...)
The FuseRule constructor accepts an aggregation_dtype parameter that controls the precision during the delta summation phase, independent of the final weight precision defined in your custom fuse function.
Key Implementation Files
Understanding the codebase structure helps when debugging or extending the fusion system:
packages/ltx-core/src/ltx_core/loader/fuse_loras.py: Containsapply_loras,aggregate_lora_products,FuseRule, andLoraProductimplementations.packages/ltx-core/src/ltx_core/loader/single_gpu_model_builder.py: Houses theSingleGPUModelBuilderclass with the.lora()and.build()methods.packages/ltx-core/src/ltx_core/loader/primitives.py: DefinesLoraStateDictWithStrengthandStateDicttype definitions.packages/ltx-trainer/configs/*_lora.yaml: Example configurations demonstrating LoRA adapter registries for specific training pipelines.packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py: Utility functions for extracting LoRA metadata and conditioning parameters.
Summary
Fusing multiple LoRA adapters in LTX-2 follows a clear, memory-efficient pipeline:
- LoraStateDictWithStrength pairs checkpoints with scaling factors to control individual adapter influence.
- aggregate_lora_products computes the combined delta by summing matrix products across all adapters without loading full tensors to GPU simultaneously.
- SingleGPUModelBuilder provides the recommended immutable API for automatic fusion during model instantiation.
- apply_loras offers direct access to the fusion engine for custom workflows requiring explicit state dictionary manipulation.
- FuseRule enables custom precision and scaling policies through a callable interface.
Frequently Asked Questions
How many LoRA adapters can I fuse simultaneously in LTX-2?
LTX-2 imposes no hardcoded limit on the number of adapters. The system processes LoRAs sequentially during aggregation in aggregate_lora_products, keeping peak memory constant regardless of adapter count. Memory constraints depend only on the base model size and the temporary storage required for the aggregated delta tensors.
What is the default fuse rule and data type used in LTX-2?
The default configuration uses bf16_fuse_rule, which performs aggregation in bfloat16 precision before adding deltas to the base weights. This rule is automatically applied when using SingleGPUModelBuilder unless explicitly overridden via .with_fuse_rule().
Can I use different strengths for each LoRA adapter?
Yes. Each call to builder.lora() or each LoraStateDictWithStrength instantiation accepts an independent strength parameter. The aggregation algorithm scales each adapter's contribution by its respective strength factor before computing the final delta, enabling fine-grained control over style mixing ratios.
Where does LTX-2 load LoRA weights during fusion?
By default, the SingleGPUModelBuilder loads LoRA checkpoints on the CPU using the lora_load_device parameter. This design minimizes GPU memory pressure during the aggregation phase. Weights transfer to the GPU only after fusion completes, ensuring efficient memory utilization when combining multiple large adapters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →