How to Use LoRA Adaptation with LongLive 2.0 Generator Models: A Complete Guide

To use LoRA adaptation with LongLive 2.0 generator models, enable the adapter block in your YAML configuration, wrap the transformer using configure_lora_for_model from utils/lora_utils.py, and let the DistillationTrainer handle training and checkpointing of the low-rank weights.

LongLive 2.0 implements LoRA (Low-Rank Adaptation) as a modular plug-in that allows efficient fine-tuning of diffusion generators without updating the full weight matrix. This approach significantly reduces memory consumption and training time while maintaining adaptation quality. The following sections detail the complete workflow from configuration to deployment, based on the actual implementation in the NVlabs/LongLive repository.

Configuring the LoRA Adapter

The first step is declaring the adapter in your configuration file. LongLive 2.0 detects LoRA settings through the top-level adapter field, which accepts parameters compatible with the PEFT library.

Create or modify your YAML configuration (e.g., configs/inference.yaml or configs/train_dmd.yaml) to include:

adapter:
  type: lora          # currently only "lora" is supported

  rank: 16            # low-rank dimension (r)

  alpha: 16           # scaling factor (defaults to rank if omitted)

  dropout: 0.0        # dropout probability for LoRA layers

The system checks for this configuration using getattr(config, "adapter", None). When present, the training and inference pipelines automatically trigger LoRA wrapping routines.

Wrapping the Transformer Model

The core utility function configure_lora_for_model in utils/lora_utils.py (lines 19–68) handles the mechanical process of converting a standard transformer into a PEFT-enhanced model. This function constructs a peft.LoraConfig from your YAML values and attaches the adapter using peft.get_peft_model.

To manually wrap a generator:

from utils.lora_utils import configure_lora_for_model

generator.model = configure_lora_for_model(
    transformer=generator.model,
    model_name="generator",
    lora_config=config.adapter,
)

This call injects trainable low-rank matrices into the attention and feed-forward layers while freezing the original base weights. The method returns the modified model ready for efficient fine-tuning.

Training with LoRA Adaptation

During the training loop, the trainer/distillation.py module manages LoRA initialization automatically. The DistillationTrainer class checks the is_lora_enabled property and, when true, wraps the generator before training begins (lines 274–284).

The trainer executes:

if self.is_lora_enabled:
    self.model.generator.model = self._configure_lora_for_model(
        self.model.generator.model, "generator")

To initiate training with LoRA:

python train.py \
  --config configs/train_dmd.yaml \
  --output_dir ./outputs/lora_run

The DistillationTrainer automatically:

  • Freezes base model parameters
  • Optimizes only the LoRA weights
  • Saves adapter checkpoints under the generator_lora key

You may also apply LoRA to the critic by enabling the apply_to_critic flag in your configuration, allowing simultaneous adaptation of both networks.

Saving and Loading LoRA Weights

LongLive 2.0 separates LoRA parameters from base model weights to enable modular checkpointing. After training, the system extracts only the adapter state using gather_lora_state_dict (also located in utils/lora_utils.py).

Checkpoint files store LoRA weights under specific keys:

  • generator_lora: Contains the generator's adapter matrices
  • critic_lora: Contains the critic's adapter matrices (if enabled)

During inference or resuming training, the pipeline loads these keys via load_lora_checkpoint and injects them into the wrapped model using peft.set_peft_model_state_dict.

Merging LoRA Weights for Production Deployment

For deployment scenarios requiring maximum inference speed, LongLive 2.0 provides scripts/merge_lora_generator.py. This utility merges trained LoRA weights back into the base generator, producing a single checkpoint that requires no adapter overhead at runtime.

The script performs:

  1. Loading the base generator checkpoint
  2. Re-attaching the LoRA adapter via configure_lora_for_model
  3. Injecting saved LoRA weights using peft.set_peft_model_state_dict
  4. Saving the merged model (lines 82–92)

Execute the merge process:

python scripts/merge_lora_generator.py \
  --generator_ckpt path/to/base_generator.pt \
  --lora_ckpt path/to/lora_checkpoint.pt \
  --output path/to/merged_generator.pt

The resulting checkpoint can be used with standard inference pipelines without configuration changes or PEFT dependencies.

Running Inference with LoRA

The inference entry points inference.py and inference_sp.py detect config.adapter automatically. If present, they wrap the generator on-the-fly before loading weights. For merged checkpoints, the wrapper initializes but remains inactive, ensuring compatibility with both adapter-based and merged models.

Example inference setup:

import torch
from omegaconf import OmegaConf
from pipeline import CausalDiffusionInferencePipeline
from utils.config import normalize_config
from utils.inference_utils import load_generator_checkpoint, setup_nvfp4_pipeline

# Load configuration containing adapter block

cfg = normalize_config(OmegaConf.load("configs/inference.yaml"))
device = torch.device("cuda")

# Initialize pipeline

pipe = CausalDiffusionInferencePipeline(cfg, device=device)

# For NVFP4 models: automatically handles LoRA wrapping

setup_nvfp4_pipeline(pipe, cfg, device)

# For standard BF16 models:

# load_generator_checkpoint(pipe.generator, "path/to/merged.pt")

Summary

  • Enable LoRA by adding an adapter block with type: lora, rank, and alpha values to your YAML configuration.
  • Wrap models using configure_lora_for_model from utils/lora_utils.py to inject PEFT adapters into the transformer.
  • Train efficiently via DistillationTrainer, which automatically handles LoRA wrapping and checkpoints adapter weights separately under generator_lora.
  • Merge for deployment using scripts/merge_lora_generator.py to bake adapters into the base model for streamlined inference.
  • Maintain compatibility—the inference pipeline detects LoRA configurations automatically and handles both merged and adapter-based checkpoints.

Frequently Asked Questions

What is the difference between merged and unmerged LoRA checkpoints?

Unmerged checkpoints contain base model weights plus separate generator_lora adapter matrices, requiring PEFT and configuration files to load. Merged checkpoints combine these into a single weight file using scripts/merge_lora_generator.py, eliminating adapter overhead and dependencies for production inference.

Can I apply LoRA to the critic model as well as the generator?

Yes. Set the apply_to_critic: true flag in your configuration. The DistillationTrainer will then wrap both networks, and checkpoints will include critic_lora weights alongside generator_lora for comprehensive low-rank adaptation.

How do I choose appropriate rank and alpha values?

The rank parameter controls the dimensionality of the low-rank matrices—values between 4 and 64 typically balance quality and efficiency. The alpha parameter scales the adapter contribution; setting alpha equal to rank (the default behavior when omitted) provides standard scaling. Higher ranks capture more fine-grained adaptations but increase memory usage.

Does LongLive 2.0 require the PEFT library to use LoRA?

Yes. The configure_lora_for_model function relies on peft.LoraConfig and peft.get_peft_model to construct adapters. Ensure the PEFT library is installed in your environment. Merged checkpoints, however, remove this dependency for inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →