How RoPE Position Embedding Offset Works for Multi-Shot Sequences in LongLive
LongLive applies a cumulative phase offset to rotary position embeddings (RoPE) at each shot boundary, allowing the transformer model to distinguish temporal positions across consecutive video generation shots without resetting the KV cache.
The NVlabs/LongLive repository implements autoregressive video generation through sequences of prompt-driven shots. To maintain temporal coherence while distinguishing between different shots, the codebase employs a RoPE position embedding offset mechanism that shifts the rotary phase at each shot boundary.
Understanding RoPE Offset in Multi-Shot Generation
The Challenge of Shot Boundaries
When generating long videos as multiple consecutive shots, the model must maintain temporal continuity while recognizing that each shot represents a distinct temporal segment. Without proper handling, the transformer sees overlapping position indices across shots, causing attention mechanisms to conflate tokens from different shots.
The Phase Offset Solution
LongLive solves this by adding a scalar phase offset phi to the RoPE calculation for each subsequent shot. As implemented in utils/position_embedding_utils.py, the compute_temporal_freqs function shifts positions by the temporal offset before computing angles, effectively injecting a constant phase shift proportional to the shot index.
Core Implementation Components
Position Embedding Utilities (utils/position_embedding_utils.py)
This file contains two critical functions that handle the offset logic:
-
select_temporal_offset_for_sample: Chooses the appropriate offset for a given sample, handling scalar values, tensors, or lists. When the offset is a scalar (the typical multi-shot case), it returns that value directly for all frames in the sample. -
compute_temporal_freqs: Builds the actual complex-valued RoPE frequencies. Whentemporal_offsetis non-zero, it adds the offset to frame positions before computing angles according to the formula:θ_t = (p + temporal_offset) * base_angle, wherepis the original frame index.
Pipeline Integration (pipeline/causal_diffusion_inference.py)
The inference pipeline tracks shot boundaries using the _is_shot_boundary(raw_prompts, chunk_index) method. When a new shot starts, it updates the model's offset attribute:
current_shot_index += 1
self._dit_model.rope_temporal_offset = current_shot_index * phi
Here, phi represents the multi-shot RoPE offset specified via the --multi_shot_rope_offset command-line argument.
Step-by-Step Execution Flow
-
Configuration: The user specifies
--multi_shot_rope_offset 0.2(stored asphi). InCausalDiffusionInferencePipeline.__init__(lines 75-80), this value initializesself.multi_shot_rope_offset. -
Boundary Detection: During inference, the pipeline iterates over frame chunks. The
_is_shot_boundarymethod identifies when a prompt change indicates a new shot. -
Offset Update: At shot boundaries (lines 62-66), the code increments
current_shot_indexand setsself._dit_model.rope_temporal_offset = current_shot_index * phi, resulting in cumulative offsets of0, phi, 2*phi, etc. -
Frequency Computation: Inside
WanModel(defined inwan_5b/modules/model.py), therope_applymethod callsselect_temporal_offset_for_sampleto retrieve the scalar offset, then passes it tocompute_temporal_freqs. Here, the offset shifts the position indices:positions = arange(f) + start_frame + temporal_offset. -
Embedding Application: The resulting complex frequencies multiply with token representations (
x_i * freqs_i), giving each shot a distinct phase signature while maintaining relative temporal ordering within each shot.
Practical Code Examples
Command-Line Configuration
python inference.py \
--model_name Wan2.2-TI2V-5B \
--multi_shot_rope_offset 0.25 \
--prompt "A sunrise over mountains" \
--prompt "A storm rolls in" \
--prompt "The storm clears"
Programmatic Implementation
from pipeline.causal_diffusion_inference import CausalDiffusionInferencePipeline
from types import SimpleNamespace
# Configure with a 0.3 radian offset per shot
args = SimpleNamespace(
model_kwargs=SimpleNamespace(model_name="Wan2.2-TI2V-5B"),
multi_shot_rope_offset=0.3,
# ... other required fields
)
pipeline = CausalDiffusionInferencePipeline(
args=args,
device="cuda",
generator=generator,
text_encoder=text_encoder,
vae=vae,
)
# The offset accumulates automatically at shot boundaries
video = pipeline.inference(
noise=torch.randn(1, 64, 4, 64, 64, device="cuda"),
text_prompts=["Scene 1", "Scene 2", "Scene 3"]
)
Debugging the Offset Value
# After processing the second shot (phi=0.3)
print(pipeline._dit_model.rope_temporal_offset) # Output: 0.6
Key Files Reference
utils/position_embedding_utils.py: Core offset logic (select_temporal_offset_for_sample,compute_temporal_freqs).wan_5b/modules/model.py: Definesrope_applythat wires the offset into frequency computation.pipeline/causal_diffusion_inference.py: Shot boundary detection andrope_temporal_offsetpropagation.pipeline/self_forcing_training.py: Training-time implementation mirroring inference behavior.
Summary
- LongLive uses a cumulative RoPE phase offset to distinguish between video shots in multi-shot generation.
- The offset
phiis multiplied by the shot index and applied viarope_temporal_offseton the DIT model. utils/position_embedding_utils.pyhandles the mathematical application of the offset to position indices before frequency calculation.- This approach requires no additional memory (just a scalar per shot) and works identically for both inference and training pipelines.
Frequently Asked Questions
What happens if I set multi_shot_rope_offset to 0.0?
When multi_shot_rope_offset is 0.0, the phase offset remains zero for all shots. According to the source code in causal_diffusion_inference.py, the pipeline only updates the offset when phi != 0.0, effectively disabling the mechanism and treating the entire sequence as a single continuous temporal block with no phase shifts between shots.
Does this mechanism work during training as well as inference?
Yes. The same rope_temporal_offset attribute is set in pipeline/self_forcing_training.py using identical logic to the inference pipeline. Both implementations track shot boundaries and update the model's offset attribute, ensuring consistent temporal encoding across training batches and generation.
How does the phase offset prevent attention confusion across shots?
By adding current_shot_index * phi to the position indices before computing rotary frequencies, each shot receives a unique phase rotation in the complex plane. Tokens from shot 1 have embeddings rotated by phi, shot 2 by 2*phi, etc., making their dot-product attention scores mathematically distinct from those of other shots while preserving intra-shot relative positions and temporal coherence.
Can I use different offset values for different shots in the same sequence?
The current implementation in CausalDiffusionInferencePipeline uses a single scalar multi_shot_rope_offset value multiplied by the shot index. However, the underlying select_temporal_offset_for_sample utility supports per-sample arrays and lists, so advanced users could modify the pipeline to pass shot-specific offsets by overriding how rope_temporal_offset is updated at each boundary.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →