# How RoPE Position Embedding Offset Works for Multi-Shot Sequences in LongLive

> Discover how LongLive's RoPE position embedding offset revolutionizes multi-shot video generation by maintaining temporal context without KV cache resets.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: deep-dive
- Published: 2026-05-24

---

**LongLive applies a cumulative phase offset to rotary position embeddings (RoPE) at each shot boundary, allowing the transformer model to distinguish temporal positions across consecutive video generation shots without resetting the KV cache.**

The NVlabs/LongLive repository implements autoregressive video generation through sequences of prompt-driven shots. To maintain temporal coherence while distinguishing between different shots, the codebase employs a **RoPE position embedding offset** mechanism that shifts the rotary phase at each shot boundary.

## Understanding RoPE Offset in Multi-Shot Generation

### The Challenge of Shot Boundaries

When generating long videos as multiple consecutive shots, the model must maintain temporal continuity while recognizing that each shot represents a distinct temporal segment. Without proper handling, the transformer sees overlapping position indices across shots, causing attention mechanisms to conflate tokens from different shots.

### The Phase Offset Solution

LongLive solves this by adding a scalar phase offset `phi` to the RoPE calculation for each subsequent shot. As implemented in [`utils/position_embedding_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/position_embedding_utils.py), the `compute_temporal_freqs` function shifts positions by the temporal offset before computing angles, effectively injecting a constant phase shift proportional to the shot index.

## Core Implementation Components

### Position Embedding Utilities ([`utils/position_embedding_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/position_embedding_utils.py))

This file contains two critical functions that handle the offset logic:

- **`select_temporal_offset_for_sample`**: Chooses the appropriate offset for a given sample, handling scalar values, tensors, or lists. When the offset is a scalar (the typical multi-shot case), it returns that value directly for all frames in the sample.

- **`compute_temporal_freqs`**: Builds the actual complex-valued RoPE frequencies. When `temporal_offset` is non-zero, it adds the offset to frame positions before computing angles according to the formula: `θ_t = (p + temporal_offset) * base_angle`, where `p` is the original frame index.

### Pipeline Integration ([`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py))

The inference pipeline tracks shot boundaries using the `_is_shot_boundary(raw_prompts, chunk_index)` method. When a new shot starts, it updates the model's offset attribute:

```python
current_shot_index += 1
self._dit_model.rope_temporal_offset = current_shot_index * phi

```

Here, `phi` represents the **multi-shot RoPE offset** specified via the `--multi_shot_rope_offset` command-line argument.

## Step-by-Step Execution Flow

1. **Configuration**: The user specifies `--multi_shot_rope_offset 0.2` (stored as `phi`). In `CausalDiffusionInferencePipeline.__init__` (lines 75-80), this value initializes `self.multi_shot_rope_offset`.

2. **Boundary Detection**: During inference, the pipeline iterates over frame chunks. The `_is_shot_boundary` method identifies when a prompt change indicates a new shot.

3. **Offset Update**: At shot boundaries (lines 62-66), the code increments `current_shot_index` and sets `self._dit_model.rope_temporal_offset = current_shot_index * phi`, resulting in cumulative offsets of `0, phi, 2*phi`, etc.

4. **Frequency Computation**: Inside `WanModel` (defined in [`wan_5b/modules/model.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/model.py)), the `rope_apply` method calls `select_temporal_offset_for_sample` to retrieve the scalar offset, then passes it to `compute_temporal_freqs`. Here, the offset shifts the position indices: `positions = arange(f) + start_frame + temporal_offset`.

5. **Embedding Application**: The resulting complex frequencies multiply with token representations (`x_i * freqs_i`), giving each shot a distinct phase signature while maintaining relative temporal ordering within each shot.

## Practical Code Examples

### Command-Line Configuration

```bash
python inference.py \
  --model_name Wan2.2-TI2V-5B \
  --multi_shot_rope_offset 0.25 \
  --prompt "A sunrise over mountains" \
  --prompt "A storm rolls in" \
  --prompt "The storm clears"

```

### Programmatic Implementation

```python
from pipeline.causal_diffusion_inference import CausalDiffusionInferencePipeline
from types import SimpleNamespace

# Configure with a 0.3 radian offset per shot

args = SimpleNamespace(
    model_kwargs=SimpleNamespace(model_name="Wan2.2-TI2V-5B"),
    multi_shot_rope_offset=0.3,
    # ... other required fields

)

pipeline = CausalDiffusionInferencePipeline(
    args=args,
    device="cuda",
    generator=generator,
    text_encoder=text_encoder,
    vae=vae,
)

# The offset accumulates automatically at shot boundaries

video = pipeline.inference(
    noise=torch.randn(1, 64, 4, 64, 64, device="cuda"),
    text_prompts=["Scene 1", "Scene 2", "Scene 3"]
)

```

### Debugging the Offset Value

```python

# After processing the second shot (phi=0.3)

print(pipeline._dit_model.rope_temporal_offset)  # Output: 0.6

```

## Key Files Reference

- **[`utils/position_embedding_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/position_embedding_utils.py)**: Core offset logic (`select_temporal_offset_for_sample`, `compute_temporal_freqs`).
- **[`wan_5b/modules/model.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/model.py)**: Defines `rope_apply` that wires the offset into frequency computation.
- **[`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py)**: Shot boundary detection and `rope_temporal_offset` propagation.
- **[`pipeline/self_forcing_training.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/self_forcing_training.py)**: Training-time implementation mirroring inference behavior.

## Summary

- LongLive uses a **cumulative RoPE phase offset** to distinguish between video shots in multi-shot generation.
- The offset `phi` is multiplied by the shot index and applied via `rope_temporal_offset` on the DIT model.
- [`utils/position_embedding_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/position_embedding_utils.py) handles the mathematical application of the offset to position indices before frequency calculation.
- This approach requires no additional memory (just a scalar per shot) and works identically for both inference and training pipelines.

## Frequently Asked Questions

### What happens if I set multi_shot_rope_offset to 0.0?

When `multi_shot_rope_offset` is `0.0`, the phase offset remains zero for all shots. According to the source code in [`causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/causal_diffusion_inference.py), the pipeline only updates the offset when `phi != 0.0`, effectively disabling the mechanism and treating the entire sequence as a single continuous temporal block with no phase shifts between shots.

### Does this mechanism work during training as well as inference?

Yes. The same `rope_temporal_offset` attribute is set in [`pipeline/self_forcing_training.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/self_forcing_training.py) using identical logic to the inference pipeline. Both implementations track shot boundaries and update the model's offset attribute, ensuring consistent temporal encoding across training batches and generation.

### How does the phase offset prevent attention confusion across shots?

By adding `current_shot_index * phi` to the position indices before computing rotary frequencies, each shot receives a unique phase rotation in the complex plane. Tokens from shot 1 have embeddings rotated by `phi`, shot 2 by `2*phi`, etc., making their dot-product attention scores mathematically distinct from those of other shots while preserving intra-shot relative positions and temporal coherence.

### Can I use different offset values for different shots in the same sequence?

The current implementation in `CausalDiffusionInferencePipeline` uses a single scalar `multi_shot_rope_offset` value multiplied by the shot index. However, the underlying `select_temporal_offset_for_sample` utility supports per-sample arrays and lists, so advanced users could modify the pipeline to pass shot-specific offsets by overriding how `rope_temporal_offset` is updated at each boundary.