How to Use LoRA Adaptation with LongLive 2.0 Generator Models: A Complete Guide
To use LoRA adaptation with LongLive 2.0 generator models, enable the adapter block in your YAML configuration, wrap the transformer using configure_lora_for_model from utils/lora_utils.py, and let the DistillationTrainer handle training and checkpointing of the low-rank weights.
LongLive 2.0 implements LoRA (Low-Rank Adaptation) as a modular plug-in that allows efficient fine-tuning of diffusion generators without updating the full weight matrix. This approach significantly reduces memory consumption and training time while maintaining adaptation quality. The following sections detail the complete workflow from configuration to deployment, based on the actual implementation in the NVlabs/LongLive repository.
Configuring the LoRA Adapter
The first step is declaring the adapter in your configuration file. LongLive 2.0 detects LoRA settings through the top-level adapter field, which accepts parameters compatible with the PEFT library.
Create or modify your YAML configuration (e.g., configs/inference.yaml or configs/train_dmd.yaml) to include:
adapter:
type: lora # currently only "lora" is supported
rank: 16 # low-rank dimension (r)
alpha: 16 # scaling factor (defaults to rank if omitted)
dropout: 0.0 # dropout probability for LoRA layers
The system checks for this configuration using getattr(config, "adapter", None). When present, the training and inference pipelines automatically trigger LoRA wrapping routines.
Wrapping the Transformer Model
The core utility function configure_lora_for_model in utils/lora_utils.py (lines 19–68) handles the mechanical process of converting a standard transformer into a PEFT-enhanced model. This function constructs a peft.LoraConfig from your YAML values and attaches the adapter using peft.get_peft_model.
To manually wrap a generator:
from utils.lora_utils import configure_lora_for_model
generator.model = configure_lora_for_model(
transformer=generator.model,
model_name="generator",
lora_config=config.adapter,
)
This call injects trainable low-rank matrices into the attention and feed-forward layers while freezing the original base weights. The method returns the modified model ready for efficient fine-tuning.
Training with LoRA Adaptation
During the training loop, the trainer/distillation.py module manages LoRA initialization automatically. The DistillationTrainer class checks the is_lora_enabled property and, when true, wraps the generator before training begins (lines 274–284).
The trainer executes:
if self.is_lora_enabled:
self.model.generator.model = self._configure_lora_for_model(
self.model.generator.model, "generator")
To initiate training with LoRA:
python train.py \
--config configs/train_dmd.yaml \
--output_dir ./outputs/lora_run
The DistillationTrainer automatically:
- Freezes base model parameters
- Optimizes only the LoRA weights
- Saves adapter checkpoints under the
generator_lorakey
You may also apply LoRA to the critic by enabling the apply_to_critic flag in your configuration, allowing simultaneous adaptation of both networks.
Saving and Loading LoRA Weights
LongLive 2.0 separates LoRA parameters from base model weights to enable modular checkpointing. After training, the system extracts only the adapter state using gather_lora_state_dict (also located in utils/lora_utils.py).
Checkpoint files store LoRA weights under specific keys:
generator_lora: Contains the generator's adapter matricescritic_lora: Contains the critic's adapter matrices (if enabled)
During inference or resuming training, the pipeline loads these keys via load_lora_checkpoint and injects them into the wrapped model using peft.set_peft_model_state_dict.
Merging LoRA Weights for Production Deployment
For deployment scenarios requiring maximum inference speed, LongLive 2.0 provides scripts/merge_lora_generator.py. This utility merges trained LoRA weights back into the base generator, producing a single checkpoint that requires no adapter overhead at runtime.
The script performs:
- Loading the base generator checkpoint
- Re-attaching the LoRA adapter via
configure_lora_for_model - Injecting saved LoRA weights using
peft.set_peft_model_state_dict - Saving the merged model (lines 82–92)
Execute the merge process:
python scripts/merge_lora_generator.py \
--generator_ckpt path/to/base_generator.pt \
--lora_ckpt path/to/lora_checkpoint.pt \
--output path/to/merged_generator.pt
The resulting checkpoint can be used with standard inference pipelines without configuration changes or PEFT dependencies.
Running Inference with LoRA
The inference entry points inference.py and inference_sp.py detect config.adapter automatically. If present, they wrap the generator on-the-fly before loading weights. For merged checkpoints, the wrapper initializes but remains inactive, ensuring compatibility with both adapter-based and merged models.
Example inference setup:
import torch
from omegaconf import OmegaConf
from pipeline import CausalDiffusionInferencePipeline
from utils.config import normalize_config
from utils.inference_utils import load_generator_checkpoint, setup_nvfp4_pipeline
# Load configuration containing adapter block
cfg = normalize_config(OmegaConf.load("configs/inference.yaml"))
device = torch.device("cuda")
# Initialize pipeline
pipe = CausalDiffusionInferencePipeline(cfg, device=device)
# For NVFP4 models: automatically handles LoRA wrapping
setup_nvfp4_pipeline(pipe, cfg, device)
# For standard BF16 models:
# load_generator_checkpoint(pipe.generator, "path/to/merged.pt")
Summary
- Enable LoRA by adding an
adapterblock withtype: lora,rank, andalphavalues to your YAML configuration. - Wrap models using
configure_lora_for_modelfromutils/lora_utils.pyto inject PEFT adapters into the transformer. - Train efficiently via
DistillationTrainer, which automatically handles LoRA wrapping and checkpoints adapter weights separately undergenerator_lora. - Merge for deployment using
scripts/merge_lora_generator.pyto bake adapters into the base model for streamlined inference. - Maintain compatibility—the inference pipeline detects LoRA configurations automatically and handles both merged and adapter-based checkpoints.
Frequently Asked Questions
What is the difference between merged and unmerged LoRA checkpoints?
Unmerged checkpoints contain base model weights plus separate generator_lora adapter matrices, requiring PEFT and configuration files to load. Merged checkpoints combine these into a single weight file using scripts/merge_lora_generator.py, eliminating adapter overhead and dependencies for production inference.
Can I apply LoRA to the critic model as well as the generator?
Yes. Set the apply_to_critic: true flag in your configuration. The DistillationTrainer will then wrap both networks, and checkpoints will include critic_lora weights alongside generator_lora for comprehensive low-rank adaptation.
How do I choose appropriate rank and alpha values?
The rank parameter controls the dimensionality of the low-rank matrices—values between 4 and 64 typically balance quality and efficiency. The alpha parameter scales the adapter contribution; setting alpha equal to rank (the default behavior when omitted) provides standard scaling. Higher ranks capture more fine-grained adaptations but increase memory usage.
Does LongLive 2.0 require the PEFT library to use LoRA?
Yes. The configure_lora_for_model function relies on peft.LoraConfig and peft.get_peft_model to construct adapters. Ensure the PEFT library is installed in your environment. Merged checkpoints, however, remove this dependency for inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →