How to Configure Multi-Shot Video Generation with Scene Cut Prefixes in LongLive
To enable multi-shot video generation in LongLive, set scene_cut_prefix in your inference configuration, activate multi-shot RoPE offsets via inference.multi_shot.enable: true, and ensure the dataset prepends the prefix to new shots, which allows the model to pin the KV cache across scene boundaries.
Generating coherent long-form videos requires efficient handling of scene transitions without recomputing attention over the entire sequence. In NVlabs/LongLive, multi-shot video generation with scene cut prefixes enables the model to produce videos as sequences of shots sharing a common KV cache, using special tokens as hard boundaries that trigger cache pinning.
How Scene-Cut Prefixes Work
LongLive treats the scene-cut prefix as a hard boundary in the token stream. When the model encounters this token, it pins the KV cache for the following shot, creating an attention sink that enables context sharing across shots. This mechanism avoids the overhead of re-encoding the full prefix for each new shot while maintaining temporal coherence.
Step-by-Step Configuration
Step 1: Enable Scene-Cut Prefixes in the Config
Set scene_cut_prefix in your inference or training configuration. The default value is "[SCENE_CUT]", defined in pipeline/self_forcing_training.py at line 27. Override this globally via config.inference.scene_cut_prefix as referenced in pipeline/causal_diffusion_inference.py at line 71.
Step 2: Configure Dataset Prefix Injection
The dataset must prepend the prefix whenever a new shot begins. In utils/dataset.py, lines 109-125 and 251-252 handle this logic: when scene_cut_prefix is non-empty, each shot's first caption is concatenated with the prefix before tokenization.
Step 3: Activate Multi-Shot RoPE Offsets
Multi-shot attention requires adjusted positional encodings per shot. Set inference.multi_shot.enable: true in your YAML to activate the RoPE offset computation. The offset logic resides in utils/position_embedding_utils.py (see the header comment on line 10 and the helper function on line 66), which adds shot-specific offsets to the rotary positional encodings.
Step 4: Pin the KV Cache on Scene Cuts
During inference, the pipeline calls pin_kv_on_scene_cut to preserve the cache across shots. This function appears in pipeline/self_forcing_training.py at line 598 and pipeline/causal_diffusion_inference.py at line 547. The pinned region serves as the attention sink for the next shot, enabling efficient cross-shot generation.
Step 5: (Optional) Control Shot Length
Limit the maximum number of shot chunks via max_shot_chunks in your config. In utils/dataset.py at line 714, this parameter forces a scene cut when the limit is reached, ensuring manageable memory usage and automatic shot boundary creation.
Configuration YAML and Code Examples
Minimal Inference Configuration
Create or modify your configs/inference.yaml:
# configs/inference.yaml
inference:
scene_cut_prefix: "[SCENE_CUT]"
multi_shot:
enable: true
offset: 0
max_shot_chunks: 64
Running Multi-Shot Inference
import torch
from omegaconf import OmegaConf
from pipeline import CausalDiffusionInferencePipeline
from utils.config import normalize_config
from utils.inference_utils import (
load_generator_checkpoint,
place_vae_for_streaming,
prepare_single_prompt_inputs,
save_video,
)
# Load configuration
config = normalize_config(OmegaConf.load("configs/inference.yaml"))
device = torch.device("cuda")
torch.set_grad_enabled(False)
# Initialize pipeline
pipe = CausalDiffusionInferencePipeline(config, device=device)
load_generator_checkpoint(pipe.generator, "LongLive-2.0-5B/model_bf16.pt")
pipe = pipe.to(device=device, dtype=torch.bfloat16)
# Optional: enable streaming VAE
place_vae_for_streaming(pipe, config)
# Prepare initial shot prompt
noise, prompts = prepare_single_prompt_inputs(
config,
"A sunrise over a mountain range.",
device
)
# Generate video with automatic scene-cut handling
video = pipe.inference(noise=noise, text_prompts=prompts)
save_video(video[0], "output/multi_shot_demo.mp4", fps=24)
Custom Prefix in Training
from utils.dataset import TextOnlyDataset
# Use custom scene-cut prefix
dataset = TextOnlyDataset(
data_path="path/to/video_captions.json",
scene_cut_prefix="[NEW_SCENE]",
# additional arguments...
)
Manual Scene Cut in Prompts
prompt = "[SCENE_CUT] A bustling city at night."
# The pipeline treats this as a new shot boundary regardless of max_shot_chunks
Key Source Files for Multi-Shot Generation
| File | Role | Key Lines |
|---|---|---|
pipeline/self_forcing_training.py |
Training-time pinning logic and default prefix | L27 (default), L598 (pin_kv_on_scene_cut) |
pipeline/causal_diffusion_inference.py |
Inference-time scene-cut handling and config override | L71 (config), L547 (pin_kv_on_scene_cut) |
utils/dataset.py |
Prefix injection and shot chunk limiting | L109-125, L251-252 (prefix), L714 (max_shot_chunks) |
utils/position_embedding_utils.py |
RoPE offset computation for multi-shot | L10 (comment), L66 (helper) |
Summary
- Scene-cut prefixes act as hard boundaries in LongLive, triggering KV cache pinning across shots.
- Configure the prefix via
inference.scene_cut_prefix(default:"[SCENE_CUT]") in your YAML. - Enable multi-shot RoPE offsets by setting
inference.multi_shot.enable: trueto handle positional encoding per shot. - The dataset automatically prepends prefixes at shot boundaries via
utils/dataset.py. - Use
pin_kv_on_scene_cutin the pipeline to maintain the attention sink across scene transitions. - Control shot length with
max_shot_chunksto force automatic scene cuts.
Frequently Asked Questions
What is the default scene-cut prefix in LongLive?
The default prefix is "[SCENE_CUT]", defined in pipeline/self_forcing_training.py at line 27. You can override this in your inference configuration file using the scene_cut_prefix key.
How does KV cache pinning work across scene cuts?
When the pipeline encounters a scene-cut token, it calls pin_kv_on_scene_cut (found at line 598 in pipeline/self_forcing_training.py and line 547 in pipeline/causal_diffusion_inference.py). This preserves the cached key-value pairs as an attention sink for the subsequent shot, allowing the model to maintain context without recomputing attention over the entire previous sequence.
Can I use a custom scene-cut prefix during training?
Yes. When initializing TextOnlyDataset from utils/dataset.py, pass your custom prefix string to the scene_cut_prefix parameter. This overrides the default and ensures the dataset prepends your custom token at shot boundaries during training.
What happens when max_shot_chunks is reached?
When a shot reaches the max_shot_chunks limit configured in your YAML, the dataset logic in utils/dataset.py (line 714) automatically inserts a scene cut. This forces a new shot boundary and triggers KV cache pinning, preventing memory overflow and maintaining generation coherence.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →