Score Distillation vs Diffusion Training in LongLive: Technical Differences Explained

LongLive uses score-distillation mode to train a student network via KL-gradient matching against a frozen teacher, while standard diffusion mode trains the generator directly through MSE-based flow prediction without external supervision.

LongLive, NVIDIA's open-source framework for long-form video generation, provides two distinct training pipelines selected by the distribution_loss configuration flag. Understanding the difference between score distillation and diffusion training modes helps developers optimize for computational resources, data availability, and model performance.

How Training Modes Are Selected

The training behavior is determined by the distribution_loss parameter in your configuration file, as parsed in utils/config.py. Setting distribution_loss = "dmd" activates the score-distillation pipeline, while omitting this flag or using a non-DMD value triggers standard diffusion training.

Score-Distillation (DMD) Mode

Score-distillation implements Distribution Matching Distillation (DMD) to transfer knowledge from a pre-trained teacher model to a student network.

Model Architecture

The DMD class in model/dmd.py wraps three components:

  1. Generator: Produces video/image outputs through backward simulation
  2. Fake Score (fake_score): The learnable student network initialized from data or checkpoints
  3. Real Score (real_score): The frozen teacher network loaded from real_score_ckpt

Training Objective

The core loss function compute_distribution_matching_loss (lines 55-84 in model/dmd.py) computes the KL-gradient between the fake and real score networks. During Trainer.train_one_step() in trainer/distillation.py, the pipeline executes:

  1. Generate video using backward simulation via self.model._run_generator()
  2. Evaluate the fake score on noisy video samples
  3. Evaluate the real score (teacher) on identical samples
  4. Compute KL-gradient in DMD._compute_kl_grad (lines 92-120)
  5. Apply guidance scaling via real_guidance_scale and fake_guidance_scale parameters

This approach allows training with minimal data because the frozen teacher provides the supervisory signal.

Standard Diffusion Training Mode

Standard diffusion training follows the conventional flow-matching paradigm without external teacher networks.

Model Architecture

The CausalDiffusion class in model/diffusion.py contains only:

  • Generator (self.generator): The primary learnable diffusion model
  • VAE and text encoder: For latent encoding and conditioning
  • Scheduler: Manages noise schedules and training targets

Training Objective

The generator_loss method (lines 71-89) computes standard diffusion loss through:

  1. Adding noise to clean latents using self.scheduler.add_noise()
  2. Predicting flow (flow_pred) and initial state (x0_pred) via the generator
  3. Calculating MSE between flow_pred and the scheduler's training target

Unlike distillation, this mode supports optional error-recycling logic that injects historic errors into the latent stream, but maintains no teacher-student relationship.

Key Architectural Differences

Aspect Score-Distillation Standard Diffusion
Loss Function compute_distribution_matching_loss (KL-gradient between scores) torch.nn.functional.mse_loss on flow predictions
Teacher Model real_score (frozen, loaded from checkpoint) None
Student Model fake_score learns to mimic teacher Generator learns from scratch
Training File trainer/distillation.py trainer/diffusion.py
Primary Model model/dmd.py model/diffusion.py
Config Flags distribution_loss = "dmd", real_score_ckpt, fake_score_ckpt distribution_loss omitted, generator_ckpt for warm-starts

Practical Training Implications

Memory Requirements: Score-distillation requires loading both fake_score and real_score networks simultaneously, significantly increasing GPU memory footprint compared to standard diffusion which only loads the generator.

Data Efficiency: The distillation mode can train effectively with limited or no training data because the frozen teacher (real_score) provides synthetic supervision. Standard diffusion requires substantial datasets to learn the denoising objective from scratch.

Training Speed: Standard diffusion executes faster per step by avoiding the dual forward passes required for teacher and student evaluation. Distillation involves additional computation for fake_score and real_score inference plus KL-gradient calculation.

Model Output: Score-distillation produces a compact student model that approximates teacher behavior, ideal for model compression. Standard diffusion yields a full-capability diffusion model trained directly on the data distribution.

Code Implementation Comparison

Score-Distillation Training Step


# Inside trainer/distillation.py → Trainer.train_one_step()

# 1. Generate video with backward simulation

pred_image, _, _, _ = self.model._run_generator(...)

# 2. Compute DMD loss (KL-gradient between fake and real scores)

dmd_loss, dmd_log = self.model.compute_distribution_matching_loss(
    image_or_video=pred_image,
    conditional_dict=cond,
    unconditional_dict=uncond,
)

# 3. Back-propagate and step optimizer

dmd_loss.backward()
self.generator_optimizer.step()

The compute_distribution_matching_loss function in model/dmd.py internally calls _compute_kl_grad to calculate the gradient matching between the student and teacher predictions.

Standard Diffusion Training Step


# Inside trainer/diffusion.py → Trainer.train_one_step()

# 1. Add noise to clean latents

noisy_latents = self.scheduler.add_noise(clean_latent, noise, timestep)

# 2. Predict flow and x0 from generator

flow_pred, x0_pred = self.model.generator(
    noisy_image_or_video=noisy_latents,
    conditional_dict=cond,
    timestep=timestep,
)

# 3. Compute diffusion loss (MSE)

loss = torch.nn.functional.mse_loss(
    flow_pred.float(), 
    self.scheduler.training_target(...).float()
)

# 4. Back-propagate and step optimizer

loss.backward()
self.generator_optimizer.step()

This implementation in model/diffusion.py uses pure flow-matching without external supervision, making it the default choice for training from scratch.

Summary

  • Score-distillation employs a teacher-student paradigm with real_score (frozen) supervising fake_score (learnable) via KL-gradient matching in model/dmd.py
  • Standard diffusion trains the generator directly using MSE loss on flow predictions in model/diffusion.py without external teachers
  • Selection occurs through the distribution_loss config flag parsed in utils/config.py
  • Distillation trades memory and speed for data efficiency and model compression
  • Standard diffusion requires more data but offers simpler implementation and lower resource usage

Frequently Asked Questions

When should I use score distillation instead of standard diffusion training?

Use score distillation when you have limited training data but access to a high-quality pre-trained teacher model, or when you need to compress a large diffusion model into a smaller student network. According to the NVlabs/LongLive source code, this mode leverages the frozen real_score network to provide supervision, eliminating the need for massive datasets required by standard diffusion training in trainer/diffusion.py.

How does the DMD loss function differ from the standard diffusion loss?

The DMD loss implemented in model/dmd.py computes a KL-gradient between the fake score network's predictions and the real score network's predictions using _compute_kl_grad, effectively measuring how well the student mimics the teacher. In contrast, the standard diffusion loss in model/diffusion.py simply calculates MSE between the predicted flow (flow_pred) and the ground-truth training target from the scheduler.

What are the memory implications of using score distillation mode?

Score-distillation mode requires approximately twice the GPU memory of standard diffusion because trainer/distillation.py maintains both the fake_score student network and the frozen real_score teacher network in memory simultaneously, along with the generator. Standard diffusion training only loads the generator and scheduler, making it feasible for hardware-constrained environments.

Can I switch between training modes during LongLive training?

While both modes share the same underlying generator architecture, switching between them requires restarting training with a different distribution_loss configuration value and potentially loading different checkpoints (real_score_ckpt for distillation versus generator_ckpt for diffusion). The utils/config.py file parses this flag at initialization, and the Trainer classes in trainer/distillation.py and trainer/diffusion.py implement incompatible loss computation logic that cannot be toggled mid-training without code modification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →