How to Configure Multi-Shot Video Generation with Scene Cut Prefixes in LongLive

To enable multi-shot video generation in LongLive, set scene_cut_prefix in your inference configuration, activate multi-shot RoPE offsets via inference.multi_shot.enable: true, and ensure the dataset prepends the prefix to new shots, which allows the model to pin the KV cache across scene boundaries.

Generating coherent long-form videos requires efficient handling of scene transitions without recomputing attention over the entire sequence. In NVlabs/LongLive, multi-shot video generation with scene cut prefixes enables the model to produce videos as sequences of shots sharing a common KV cache, using special tokens as hard boundaries that trigger cache pinning.

How Scene-Cut Prefixes Work

LongLive treats the scene-cut prefix as a hard boundary in the token stream. When the model encounters this token, it pins the KV cache for the following shot, creating an attention sink that enables context sharing across shots. This mechanism avoids the overhead of re-encoding the full prefix for each new shot while maintaining temporal coherence.

Step-by-Step Configuration

Step 1: Enable Scene-Cut Prefixes in the Config

Set scene_cut_prefix in your inference or training configuration. The default value is "[SCENE_CUT]", defined in pipeline/self_forcing_training.py at line 27. Override this globally via config.inference.scene_cut_prefix as referenced in pipeline/causal_diffusion_inference.py at line 71.

Step 2: Configure Dataset Prefix Injection

The dataset must prepend the prefix whenever a new shot begins. In utils/dataset.py, lines 109-125 and 251-252 handle this logic: when scene_cut_prefix is non-empty, each shot's first caption is concatenated with the prefix before tokenization.

Step 3: Activate Multi-Shot RoPE Offsets

Multi-shot attention requires adjusted positional encodings per shot. Set inference.multi_shot.enable: true in your YAML to activate the RoPE offset computation. The offset logic resides in utils/position_embedding_utils.py (see the header comment on line 10 and the helper function on line 66), which adds shot-specific offsets to the rotary positional encodings.

Step 4: Pin the KV Cache on Scene Cuts

During inference, the pipeline calls pin_kv_on_scene_cut to preserve the cache across shots. This function appears in pipeline/self_forcing_training.py at line 598 and pipeline/causal_diffusion_inference.py at line 547. The pinned region serves as the attention sink for the next shot, enabling efficient cross-shot generation.

Step 5: (Optional) Control Shot Length

Limit the maximum number of shot chunks via max_shot_chunks in your config. In utils/dataset.py at line 714, this parameter forces a scene cut when the limit is reached, ensuring manageable memory usage and automatic shot boundary creation.

Configuration YAML and Code Examples

Minimal Inference Configuration

Create or modify your configs/inference.yaml:


# configs/inference.yaml

inference:
  scene_cut_prefix: "[SCENE_CUT]"
  multi_shot:
    enable: true
    offset: 0
    max_shot_chunks: 64

Running Multi-Shot Inference

import torch
from omegaconf import OmegaConf
from pipeline import CausalDiffusionInferencePipeline
from utils.config import normalize_config
from utils.inference_utils import (
    load_generator_checkpoint,
    place_vae_for_streaming,
    prepare_single_prompt_inputs,
    save_video,
)

# Load configuration

config = normalize_config(OmegaConf.load("configs/inference.yaml"))
device = torch.device("cuda")
torch.set_grad_enabled(False)

# Initialize pipeline

pipe = CausalDiffusionInferencePipeline(config, device=device)
load_generator_checkpoint(pipe.generator, "LongLive-2.0-5B/model_bf16.pt")
pipe = pipe.to(device=device, dtype=torch.bfloat16)

# Optional: enable streaming VAE

place_vae_for_streaming(pipe, config)

# Prepare initial shot prompt

noise, prompts = prepare_single_prompt_inputs(
    config, 
    "A sunrise over a mountain range.", 
    device
)

# Generate video with automatic scene-cut handling

video = pipe.inference(noise=noise, text_prompts=prompts)
save_video(video[0], "output/multi_shot_demo.mp4", fps=24)

Custom Prefix in Training

from utils.dataset import TextOnlyDataset

# Use custom scene-cut prefix

dataset = TextOnlyDataset(
    data_path="path/to/video_captions.json",
    scene_cut_prefix="[NEW_SCENE]",
    # additional arguments...

)

Manual Scene Cut in Prompts

prompt = "[SCENE_CUT] A bustling city at night."

# The pipeline treats this as a new shot boundary regardless of max_shot_chunks

Key Source Files for Multi-Shot Generation

File Role Key Lines
pipeline/self_forcing_training.py Training-time pinning logic and default prefix L27 (default), L598 (pin_kv_on_scene_cut)
pipeline/causal_diffusion_inference.py Inference-time scene-cut handling and config override L71 (config), L547 (pin_kv_on_scene_cut)
utils/dataset.py Prefix injection and shot chunk limiting L109-125, L251-252 (prefix), L714 (max_shot_chunks)
utils/position_embedding_utils.py RoPE offset computation for multi-shot L10 (comment), L66 (helper)

Summary

  • Scene-cut prefixes act as hard boundaries in LongLive, triggering KV cache pinning across shots.
  • Configure the prefix via inference.scene_cut_prefix (default: "[SCENE_CUT]") in your YAML.
  • Enable multi-shot RoPE offsets by setting inference.multi_shot.enable: true to handle positional encoding per shot.
  • The dataset automatically prepends prefixes at shot boundaries via utils/dataset.py.
  • Use pin_kv_on_scene_cut in the pipeline to maintain the attention sink across scene transitions.
  • Control shot length with max_shot_chunks to force automatic scene cuts.

Frequently Asked Questions

What is the default scene-cut prefix in LongLive?

The default prefix is "[SCENE_CUT]", defined in pipeline/self_forcing_training.py at line 27. You can override this in your inference configuration file using the scene_cut_prefix key.

How does KV cache pinning work across scene cuts?

When the pipeline encounters a scene-cut token, it calls pin_kv_on_scene_cut (found at line 598 in pipeline/self_forcing_training.py and line 547 in pipeline/causal_diffusion_inference.py). This preserves the cached key-value pairs as an attention sink for the subsequent shot, allowing the model to maintain context without recomputing attention over the entire previous sequence.

Can I use a custom scene-cut prefix during training?

Yes. When initializing TextOnlyDataset from utils/dataset.py, pass your custom prefix string to the scene_cut_prefix parameter. This overrides the default and ensures the dataset prepends your custom token at shot boundaries during training.

What happens when max_shot_chunks is reached?

When a shot reaches the max_shot_chunks limit configured in your YAML, the dataset logic in utils/dataset.py (line 714) automatically inserts a scene cut. This forces a new shot boundary and triggers KV cache pinning, preventing memory overflow and maintaining generation coherence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →