# Score Distillation vs Diffusion Training in LongLive: Technical Differences Explained

> Explore score distillation vs diffusion training in NVlabs LongLive. Understand key technical differences: KL-gradient matching versus direct MSE flow prediction for your generative models.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: deep-dive
- Published: 2026-05-24

---

**LongLive uses score-distillation mode to train a student network via KL-gradient matching against a frozen teacher, while standard diffusion mode trains the generator directly through MSE-based flow prediction without external supervision.**

LongLive, NVIDIA's open-source framework for long-form video generation, provides two distinct training pipelines selected by the `distribution_loss` configuration flag. Understanding the difference between score distillation and diffusion training modes helps developers optimize for computational resources, data availability, and model performance.

## How Training Modes Are Selected

The training behavior is determined by the `distribution_loss` parameter in your configuration file, as parsed in [`utils/config.py`](https://github.com/NVlabs/LongLive/blob/main/utils/config.py). Setting `distribution_loss = "dmd"` activates the score-distillation pipeline, while omitting this flag or using a non-DMD value triggers standard diffusion training.

- **Score-Distillation (DMD)**: Uses [`trainer/distillation.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/distillation.py) with the `DMD` model class
- **Standard Diffusion**: Uses [`trainer/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/diffusion.py) with the `CausalDiffusion` model class

## Score-Distillation (DMD) Mode

Score-distillation implements Distribution Matching Distillation (DMD) to transfer knowledge from a pre-trained teacher model to a student network.

### Model Architecture

The `DMD` class in [`model/dmd.py`](https://github.com/NVlabs/LongLive/blob/main/model/dmd.py) wraps three components:

1. **Generator**: Produces video/image outputs through backward simulation
2. **Fake Score** (`fake_score`): The learnable student network initialized from data or checkpoints
3. **Real Score** (`real_score`): The frozen teacher network loaded from `real_score_ckpt`

### Training Objective

The core loss function `compute_distribution_matching_loss` (lines 55-84 in [`model/dmd.py`](https://github.com/NVlabs/LongLive/blob/main/model/dmd.py)) computes the KL-gradient between the fake and real score networks. During `Trainer.train_one_step()` in [`trainer/distillation.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/distillation.py), the pipeline executes:

1. Generate video using backward simulation via `self.model._run_generator()`
2. Evaluate the fake score on noisy video samples
3. Evaluate the real score (teacher) on identical samples
4. Compute KL-gradient in `DMD._compute_kl_grad` (lines 92-120)
5. Apply guidance scaling via `real_guidance_scale` and `fake_guidance_scale` parameters

This approach allows training with minimal data because the frozen teacher provides the supervisory signal.

## Standard Diffusion Training Mode

Standard diffusion training follows the conventional flow-matching paradigm without external teacher networks.

### Model Architecture

The `CausalDiffusion` class in [`model/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/model/diffusion.py) contains only:

- **Generator** (`self.generator`): The primary learnable diffusion model
- **VAE and text encoder**: For latent encoding and conditioning
- **Scheduler**: Manages noise schedules and training targets

### Training Objective

The `generator_loss` method (lines 71-89) computes standard diffusion loss through:

1. Adding noise to clean latents using `self.scheduler.add_noise()`
2. Predicting flow (`flow_pred`) and initial state (`x0_pred`) via the generator
3. Calculating MSE between `flow_pred` and the scheduler's training target

Unlike distillation, this mode supports optional error-recycling logic that injects historic errors into the latent stream, but maintains no teacher-student relationship.

## Key Architectural Differences

| Aspect | Score-Distillation | Standard Diffusion |
|--------|-------------------|-------------------|
| **Loss Function** | `compute_distribution_matching_loss` (KL-gradient between scores) | `torch.nn.functional.mse_loss` on flow predictions |
| **Teacher Model** | `real_score` (frozen, loaded from checkpoint) | None |
| **Student Model** | `fake_score` learns to mimic teacher | Generator learns from scratch |
| **Training File** | [`trainer/distillation.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/distillation.py) | [`trainer/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/diffusion.py) |
| **Primary Model** | [`model/dmd.py`](https://github.com/NVlabs/LongLive/blob/main/model/dmd.py) | [`model/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/model/diffusion.py) |
| **Config Flags** | `distribution_loss = "dmd"`, `real_score_ckpt`, `fake_score_ckpt` | `distribution_loss` omitted, `generator_ckpt` for warm-starts |

## Practical Training Implications

**Memory Requirements**: Score-distillation requires loading both `fake_score` and `real_score` networks simultaneously, significantly increasing GPU memory footprint compared to standard diffusion which only loads the generator.

**Data Efficiency**: The distillation mode can train effectively with limited or no training data because the frozen teacher (`real_score`) provides synthetic supervision. Standard diffusion requires substantial datasets to learn the denoising objective from scratch.

**Training Speed**: Standard diffusion executes faster per step by avoiding the dual forward passes required for teacher and student evaluation. Distillation involves additional computation for `fake_score` and `real_score` inference plus KL-gradient calculation.

**Model Output**: Score-distillation produces a compact student model that approximates teacher behavior, ideal for model compression. Standard diffusion yields a full-capability diffusion model trained directly on the data distribution.

## Code Implementation Comparison

### Score-Distillation Training Step

```python

# Inside trainer/distillation.py → Trainer.train_one_step()

# 1. Generate video with backward simulation

pred_image, _, _, _ = self.model._run_generator(...)

# 2. Compute DMD loss (KL-gradient between fake and real scores)

dmd_loss, dmd_log = self.model.compute_distribution_matching_loss(
    image_or_video=pred_image,
    conditional_dict=cond,
    unconditional_dict=uncond,
)

# 3. Back-propagate and step optimizer

dmd_loss.backward()
self.generator_optimizer.step()

```

The `compute_distribution_matching_loss` function in [`model/dmd.py`](https://github.com/NVlabs/LongLive/blob/main/model/dmd.py) internally calls `_compute_kl_grad` to calculate the gradient matching between the student and teacher predictions.

### Standard Diffusion Training Step

```python

# Inside trainer/diffusion.py → Trainer.train_one_step()

# 1. Add noise to clean latents

noisy_latents = self.scheduler.add_noise(clean_latent, noise, timestep)

# 2. Predict flow and x0 from generator

flow_pred, x0_pred = self.model.generator(
    noisy_image_or_video=noisy_latents,
    conditional_dict=cond,
    timestep=timestep,
)

# 3. Compute diffusion loss (MSE)

loss = torch.nn.functional.mse_loss(
    flow_pred.float(), 
    self.scheduler.training_target(...).float()
)

# 4. Back-propagate and step optimizer

loss.backward()
self.generator_optimizer.step()

```

This implementation in [`model/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/model/diffusion.py) uses pure flow-matching without external supervision, making it the default choice for training from scratch.

## Summary

- **Score-distillation** employs a teacher-student paradigm with `real_score` (frozen) supervising `fake_score` (learnable) via KL-gradient matching in [`model/dmd.py`](https://github.com/NVlabs/LongLive/blob/main/model/dmd.py)
- **Standard diffusion** trains the generator directly using MSE loss on flow predictions in [`model/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/model/diffusion.py) without external teachers
- Selection occurs through the `distribution_loss` config flag parsed in [`utils/config.py`](https://github.com/NVlabs/LongLive/blob/main/utils/config.py)
- Distillation trades memory and speed for data efficiency and model compression
- Standard diffusion requires more data but offers simpler implementation and lower resource usage

## Frequently Asked Questions

### When should I use score distillation instead of standard diffusion training?

Use score distillation when you have limited training data but access to a high-quality pre-trained teacher model, or when you need to compress a large diffusion model into a smaller student network. According to the NVlabs/LongLive source code, this mode leverages the frozen `real_score` network to provide supervision, eliminating the need for massive datasets required by standard diffusion training in [`trainer/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/diffusion.py).

### How does the DMD loss function differ from the standard diffusion loss?

The DMD loss implemented in [`model/dmd.py`](https://github.com/NVlabs/LongLive/blob/main/model/dmd.py) computes a KL-gradient between the fake score network's predictions and the real score network's predictions using `_compute_kl_grad`, effectively measuring how well the student mimics the teacher. In contrast, the standard diffusion loss in [`model/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/model/diffusion.py) simply calculates MSE between the predicted flow (`flow_pred`) and the ground-truth training target from the scheduler.

### What are the memory implications of using score distillation mode?

Score-distillation mode requires approximately twice the GPU memory of standard diffusion because [`trainer/distillation.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/distillation.py) maintains both the `fake_score` student network and the frozen `real_score` teacher network in memory simultaneously, along with the generator. Standard diffusion training only loads the generator and scheduler, making it feasible for hardware-constrained environments.

### Can I switch between training modes during LongLive training?

While both modes share the same underlying generator architecture, switching between them requires restarting training with a different `distribution_loss` configuration value and potentially loading different checkpoints (`real_score_ckpt` for distillation versus `generator_ckpt` for diffusion). The [`utils/config.py`](https://github.com/NVlabs/LongLive/blob/main/utils/config.py) file parses this flag at initialization, and the `Trainer` classes in [`trainer/distillation.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/distillation.py) and [`trainer/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/diffusion.py) implement incompatible loss computation logic that cannot be toggled mid-training without code modification.