Score Distillation vs Diffusion Training in LongLive: Technical Differences Explained
LongLive uses score-distillation mode to train a student network via KL-gradient matching against a frozen teacher, while standard diffusion mode trains the generator directly through MSE-based flow prediction without external supervision.
LongLive, NVIDIA's open-source framework for long-form video generation, provides two distinct training pipelines selected by the distribution_loss configuration flag. Understanding the difference between score distillation and diffusion training modes helps developers optimize for computational resources, data availability, and model performance.
How Training Modes Are Selected
The training behavior is determined by the distribution_loss parameter in your configuration file, as parsed in utils/config.py. Setting distribution_loss = "dmd" activates the score-distillation pipeline, while omitting this flag or using a non-DMD value triggers standard diffusion training.
- Score-Distillation (DMD): Uses
trainer/distillation.pywith theDMDmodel class - Standard Diffusion: Uses
trainer/diffusion.pywith theCausalDiffusionmodel class
Score-Distillation (DMD) Mode
Score-distillation implements Distribution Matching Distillation (DMD) to transfer knowledge from a pre-trained teacher model to a student network.
Model Architecture
The DMD class in model/dmd.py wraps three components:
- Generator: Produces video/image outputs through backward simulation
- Fake Score (
fake_score): The learnable student network initialized from data or checkpoints - Real Score (
real_score): The frozen teacher network loaded fromreal_score_ckpt
Training Objective
The core loss function compute_distribution_matching_loss (lines 55-84 in model/dmd.py) computes the KL-gradient between the fake and real score networks. During Trainer.train_one_step() in trainer/distillation.py, the pipeline executes:
- Generate video using backward simulation via
self.model._run_generator() - Evaluate the fake score on noisy video samples
- Evaluate the real score (teacher) on identical samples
- Compute KL-gradient in
DMD._compute_kl_grad(lines 92-120) - Apply guidance scaling via
real_guidance_scaleandfake_guidance_scaleparameters
This approach allows training with minimal data because the frozen teacher provides the supervisory signal.
Standard Diffusion Training Mode
Standard diffusion training follows the conventional flow-matching paradigm without external teacher networks.
Model Architecture
The CausalDiffusion class in model/diffusion.py contains only:
- Generator (
self.generator): The primary learnable diffusion model - VAE and text encoder: For latent encoding and conditioning
- Scheduler: Manages noise schedules and training targets
Training Objective
The generator_loss method (lines 71-89) computes standard diffusion loss through:
- Adding noise to clean latents using
self.scheduler.add_noise() - Predicting flow (
flow_pred) and initial state (x0_pred) via the generator - Calculating MSE between
flow_predand the scheduler's training target
Unlike distillation, this mode supports optional error-recycling logic that injects historic errors into the latent stream, but maintains no teacher-student relationship.
Key Architectural Differences
| Aspect | Score-Distillation | Standard Diffusion |
|---|---|---|
| Loss Function | compute_distribution_matching_loss (KL-gradient between scores) |
torch.nn.functional.mse_loss on flow predictions |
| Teacher Model | real_score (frozen, loaded from checkpoint) |
None |
| Student Model | fake_score learns to mimic teacher |
Generator learns from scratch |
| Training File | trainer/distillation.py |
trainer/diffusion.py |
| Primary Model | model/dmd.py |
model/diffusion.py |
| Config Flags | distribution_loss = "dmd", real_score_ckpt, fake_score_ckpt |
distribution_loss omitted, generator_ckpt for warm-starts |
Practical Training Implications
Memory Requirements: Score-distillation requires loading both fake_score and real_score networks simultaneously, significantly increasing GPU memory footprint compared to standard diffusion which only loads the generator.
Data Efficiency: The distillation mode can train effectively with limited or no training data because the frozen teacher (real_score) provides synthetic supervision. Standard diffusion requires substantial datasets to learn the denoising objective from scratch.
Training Speed: Standard diffusion executes faster per step by avoiding the dual forward passes required for teacher and student evaluation. Distillation involves additional computation for fake_score and real_score inference plus KL-gradient calculation.
Model Output: Score-distillation produces a compact student model that approximates teacher behavior, ideal for model compression. Standard diffusion yields a full-capability diffusion model trained directly on the data distribution.
Code Implementation Comparison
Score-Distillation Training Step
# Inside trainer/distillation.py → Trainer.train_one_step()
# 1. Generate video with backward simulation
pred_image, _, _, _ = self.model._run_generator(...)
# 2. Compute DMD loss (KL-gradient between fake and real scores)
dmd_loss, dmd_log = self.model.compute_distribution_matching_loss(
image_or_video=pred_image,
conditional_dict=cond,
unconditional_dict=uncond,
)
# 3. Back-propagate and step optimizer
dmd_loss.backward()
self.generator_optimizer.step()
The compute_distribution_matching_loss function in model/dmd.py internally calls _compute_kl_grad to calculate the gradient matching between the student and teacher predictions.
Standard Diffusion Training Step
# Inside trainer/diffusion.py → Trainer.train_one_step()
# 1. Add noise to clean latents
noisy_latents = self.scheduler.add_noise(clean_latent, noise, timestep)
# 2. Predict flow and x0 from generator
flow_pred, x0_pred = self.model.generator(
noisy_image_or_video=noisy_latents,
conditional_dict=cond,
timestep=timestep,
)
# 3. Compute diffusion loss (MSE)
loss = torch.nn.functional.mse_loss(
flow_pred.float(),
self.scheduler.training_target(...).float()
)
# 4. Back-propagate and step optimizer
loss.backward()
self.generator_optimizer.step()
This implementation in model/diffusion.py uses pure flow-matching without external supervision, making it the default choice for training from scratch.
Summary
- Score-distillation employs a teacher-student paradigm with
real_score(frozen) supervisingfake_score(learnable) via KL-gradient matching inmodel/dmd.py - Standard diffusion trains the generator directly using MSE loss on flow predictions in
model/diffusion.pywithout external teachers - Selection occurs through the
distribution_lossconfig flag parsed inutils/config.py - Distillation trades memory and speed for data efficiency and model compression
- Standard diffusion requires more data but offers simpler implementation and lower resource usage
Frequently Asked Questions
When should I use score distillation instead of standard diffusion training?
Use score distillation when you have limited training data but access to a high-quality pre-trained teacher model, or when you need to compress a large diffusion model into a smaller student network. According to the NVlabs/LongLive source code, this mode leverages the frozen real_score network to provide supervision, eliminating the need for massive datasets required by standard diffusion training in trainer/diffusion.py.
How does the DMD loss function differ from the standard diffusion loss?
The DMD loss implemented in model/dmd.py computes a KL-gradient between the fake score network's predictions and the real score network's predictions using _compute_kl_grad, effectively measuring how well the student mimics the teacher. In contrast, the standard diffusion loss in model/diffusion.py simply calculates MSE between the predicted flow (flow_pred) and the ground-truth training target from the scheduler.
What are the memory implications of using score distillation mode?
Score-distillation mode requires approximately twice the GPU memory of standard diffusion because trainer/distillation.py maintains both the fake_score student network and the frozen real_score teacher network in memory simultaneously, along with the generator. Standard diffusion training only loads the generator and scheduler, making it feasible for hardware-constrained environments.
Can I switch between training modes during LongLive training?
While both modes share the same underlying generator architecture, switching between them requires restarting training with a different distribution_loss configuration value and potentially loading different checkpoints (real_score_ckpt for distillation versus generator_ckpt for diffusion). The utils/config.py file parses this flag at initialization, and the Trainer classes in trainer/distillation.py and trainer/diffusion.py implement incompatible loss computation logic that cannot be toggled mid-training without code modification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →