# How to Use LoRA Adaptation with LongLive 2.0 Generator Models: A Complete Guide

> Master LoRA adaptation with LongLive 2.0 generator models. Configure adapters, wrap transformers, and let DistillationTrainer manage low-rank weights for efficient training.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: how-to-guide
- Published: 2026-05-24

---

**To use LoRA adaptation with LongLive 2.0 generator models, enable the `adapter` block in your YAML configuration, wrap the transformer using `configure_lora_for_model` from [`utils/lora_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/lora_utils.py), and let the `DistillationTrainer` handle training and checkpointing of the low-rank weights.**

LongLive 2.0 implements LoRA (Low-Rank Adaptation) as a modular plug-in that allows efficient fine-tuning of diffusion generators without updating the full weight matrix. This approach significantly reduces memory consumption and training time while maintaining adaptation quality. The following sections detail the complete workflow from configuration to deployment, based on the actual implementation in the NVlabs/LongLive repository.

## Configuring the LoRA Adapter

The first step is declaring the adapter in your configuration file. LongLive 2.0 detects LoRA settings through the top-level `adapter` field, which accepts parameters compatible with the PEFT library.

Create or modify your YAML configuration (e.g., [`configs/inference.yaml`](https://github.com/NVlabs/LongLive/blob/main/configs/inference.yaml) or [`configs/train_dmd.yaml`](https://github.com/NVlabs/LongLive/blob/main/configs/train_dmd.yaml)) to include:

```yaml
adapter:
  type: lora          # currently only "lora" is supported

  rank: 16            # low-rank dimension (r)

  alpha: 16           # scaling factor (defaults to rank if omitted)

  dropout: 0.0        # dropout probability for LoRA layers

```

The system checks for this configuration using `getattr(config, "adapter", None)`. When present, the training and inference pipelines automatically trigger LoRA wrapping routines.

## Wrapping the Transformer Model

The core utility function `configure_lora_for_model` in **[`utils/lora_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/lora_utils.py)** (lines 19–68) handles the mechanical process of converting a standard transformer into a PEFT-enhanced model. This function constructs a `peft.LoraConfig` from your YAML values and attaches the adapter using `peft.get_peft_model`.

To manually wrap a generator:

```python
from utils.lora_utils import configure_lora_for_model

generator.model = configure_lora_for_model(
    transformer=generator.model,
    model_name="generator",
    lora_config=config.adapter,
)

```

This call injects trainable low-rank matrices into the attention and feed-forward layers while freezing the original base weights. The method returns the modified model ready for efficient fine-tuning.

## Training with LoRA Adaptation

During the training loop, the **[`trainer/distillation.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/distillation.py)** module manages LoRA initialization automatically. The `DistillationTrainer` class checks the `is_lora_enabled` property and, when true, wraps the generator before training begins (lines 274–284).

The trainer executes:

```python
if self.is_lora_enabled:
    self.model.generator.model = self._configure_lora_for_model(
        self.model.generator.model, "generator")

```

To initiate training with LoRA:

```bash
python train.py \
  --config configs/train_dmd.yaml \
  --output_dir ./outputs/lora_run

```

The `DistillationTrainer` automatically:
- Freezes base model parameters
- Optimizes only the LoRA weights
- Saves adapter checkpoints under the `generator_lora` key

You may also apply LoRA to the critic by enabling the `apply_to_critic` flag in your configuration, allowing simultaneous adaptation of both networks.

## Saving and Loading LoRA Weights

LongLive 2.0 separates LoRA parameters from base model weights to enable modular checkpointing. After training, the system extracts only the adapter state using `gather_lora_state_dict` (also located in [`utils/lora_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/lora_utils.py)).

Checkpoint files store LoRA weights under specific keys:
- `generator_lora`: Contains the generator's adapter matrices
- `critic_lora`: Contains the critic's adapter matrices (if enabled)

During inference or resuming training, the pipeline loads these keys via `load_lora_checkpoint` and injects them into the wrapped model using `peft.set_peft_model_state_dict`.

## Merging LoRA Weights for Production Deployment

For deployment scenarios requiring maximum inference speed, LongLive 2.0 provides **[`scripts/merge_lora_generator.py`](https://github.com/NVlabs/LongLive/blob/main/scripts/merge_lora_generator.py)**. This utility merges trained LoRA weights back into the base generator, producing a single checkpoint that requires no adapter overhead at runtime.

The script performs:
1. Loading the base generator checkpoint
2. Re-attaching the LoRA adapter via `configure_lora_for_model`
3. Injecting saved LoRA weights using `peft.set_peft_model_state_dict`
4. Saving the merged model (lines 82–92)

Execute the merge process:

```bash
python scripts/merge_lora_generator.py \
  --generator_ckpt path/to/base_generator.pt \
  --lora_ckpt path/to/lora_checkpoint.pt \
  --output path/to/merged_generator.pt

```

The resulting checkpoint can be used with standard inference pipelines without configuration changes or PEFT dependencies.

## Running Inference with LoRA

The inference entry points [`inference.py`](https://github.com/NVlabs/LongLive/blob/main/inference.py) and [`inference_sp.py`](https://github.com/NVlabs/LongLive/blob/main/inference_sp.py) detect `config.adapter` automatically. If present, they wrap the generator on-the-fly before loading weights. For merged checkpoints, the wrapper initializes but remains inactive, ensuring compatibility with both adapter-based and merged models.

Example inference setup:

```python
import torch
from omegaconf import OmegaConf
from pipeline import CausalDiffusionInferencePipeline
from utils.config import normalize_config
from utils.inference_utils import load_generator_checkpoint, setup_nvfp4_pipeline

# Load configuration containing adapter block

cfg = normalize_config(OmegaConf.load("configs/inference.yaml"))
device = torch.device("cuda")

# Initialize pipeline

pipe = CausalDiffusionInferencePipeline(cfg, device=device)

# For NVFP4 models: automatically handles LoRA wrapping

setup_nvfp4_pipeline(pipe, cfg, device)

# For standard BF16 models:

# load_generator_checkpoint(pipe.generator, "path/to/merged.pt")

```

## Summary

- **Enable LoRA** by adding an `adapter` block with `type: lora`, `rank`, and `alpha` values to your YAML configuration.
- **Wrap models** using `configure_lora_for_model` from [`utils/lora_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/lora_utils.py) to inject PEFT adapters into the transformer.
- **Train efficiently** via `DistillationTrainer`, which automatically handles LoRA wrapping and checkpoints adapter weights separately under `generator_lora`.
- **Merge for deployment** using [`scripts/merge_lora_generator.py`](https://github.com/NVlabs/LongLive/blob/main/scripts/merge_lora_generator.py) to bake adapters into the base model for streamlined inference.
- **Maintain compatibility**—the inference pipeline detects LoRA configurations automatically and handles both merged and adapter-based checkpoints.

## Frequently Asked Questions

### What is the difference between merged and unmerged LoRA checkpoints?

Unmerged checkpoints contain base model weights plus separate `generator_lora` adapter matrices, requiring PEFT and configuration files to load. Merged checkpoints combine these into a single weight file using [`scripts/merge_lora_generator.py`](https://github.com/NVlabs/LongLive/blob/main/scripts/merge_lora_generator.py), eliminating adapter overhead and dependencies for production inference.

### Can I apply LoRA to the critic model as well as the generator?

Yes. Set the `apply_to_critic: true` flag in your configuration. The `DistillationTrainer` will then wrap both networks, and checkpoints will include `critic_lora` weights alongside `generator_lora` for comprehensive low-rank adaptation.

### How do I choose appropriate rank and alpha values?

The `rank` parameter controls the dimensionality of the low-rank matrices—values between 4 and 64 typically balance quality and efficiency. The `alpha` parameter scales the adapter contribution; setting `alpha` equal to `rank` (the default behavior when omitted) provides standard scaling. Higher ranks capture more fine-grained adaptations but increase memory usage.

### Does LongLive 2.0 require the PEFT library to use LoRA?

Yes. The `configure_lora_for_model` function relies on `peft.LoraConfig` and `peft.get_peft_model` to construct adapters. Ensure the PEFT library is installed in your environment. Merged checkpoints, however, remove this dependency for inference.