# Video RoPE Positional Encoding: 3D Rotary Embeddings for Long-Sequence Video in Ling-Bot-Map

> Discover Video RoPE positional encoding, a 3D rotary embedding for transformers. Enhance long video sequence modeling with linear complexity and NTK-aware extrapolation. Explore the Robbyant/lingbot-map repository.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-29

---

**Video RoPE is a three-dimensional rotary position encoding that independently encodes temporal, vertical, and horizontal coordinates, enabling transformers to model long video sequences with linear complexity and NTK-aware extrapolation.**

Video RoPE positional encoding extends classic transformer rotary embeddings into three dimensions to capture spatio-temporal relationships in video data. In the `Robbyant/lingbot-map` repository, this mechanism is implemented through the `WanRotaryPosEmbed` class in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py), providing a computationally efficient alternative to absolute positional encodings for high-resolution video generation.

## What is Video RoPE Positional Encoding?

Video RoPE (Rotary Position Embedding) generalizes the standard 2D rotary embedding to handle the three-dimensional grid of video data: frames × height × width. Unlike spatial-only encodings, Video RoPE treats the temporal dimension as a continuous rotary coordinate, preserving relative positional information across frame sequences while maintaining the translation-invariant properties of rotary embeddings.

### The Three Rotary Dimensions

The encoding decomposes video coordinates into three independent 1D rotary embeddings that are concatenated into a single complex-valued frequency tensor of shape `[max_seq_len, head_dim//2]`:

- **t (temporal)** – Encodes the frame index using a dedicated 1D RoPE on the time head-dimension, capturing motion between frames
- **h (height)** – Encodes the patch row within a frame via vertical spatial rotation  
- **w (width)** – Encodes the patch column via horizontal spatial rotation

In [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) (lines 202-276), the `get_1d_rotary_pos_embed` function builds these three frequency tables independently, which `WanRotaryPosEmbed` then concatenates using `torch.cat` to form the final embedding. During inference, the `forward` method (lines 336-393) slices this tensor according to the current frame range (`f_start`, `f_end`) and broadcasts to shape `[1, 1, N, head_dim//2]`.

## Implementation in the Ling-Bot-Map Codebase

The integration of Video RoPE spans several key modules, with conditional logic that falls back to 2D spatial encoding when `enable_3d_rope` is disabled.

### Core RoPE Module

The `WanRotaryPosEmbed` class in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) manages the pre-computation and caching of frequency tables. The class caches tables per `(dim, seq_len, device, dtype)` tuple (lines 102-123), avoiding expensive host-to-device synchronizations that would break CUDA graph capture.

### Attention Integration

Inside [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) (lines 154-157 and 216-220), the model conditionally applies Video RoPE:

```python

# attention.py – Q/K preparation with conditional 3D RoPE

if self.rope is not None and not enable_3d_rope:
    q = self.rope(q, pos)  # 2-D spatial RoPE

    k = self.rope(k, pos)
elif enable_3d_rope and pos is not None:
    q = self.rope(q, pos)  # 3-D spatio-temporal RoPE

    k = self.rope(k, pos)

```

When enabled, the same `WanRotaryPosEmbed` instance receives a 3D positional tensor containing temporal and spatial coordinates.

## How Video RoPE Affects Long-Sequence Performance

Video RoPE enables efficient processing of video sequences far longer than training data through four key mechanisms that maintain computational linearity and numerical stability.

### NTK-Aware Frequency Scaling

The `get_1d_rotary_pos_embed` function accepts an `ntk_factor` argument (lines 42-44) that scales the base frequency `theta` (line 242: `theta = theta * ntk_factor`). This Neural-Tangents-Kernel scaling stretches the sinusoidal spectrum, allowing the embedding to extrapolate to sequence lengths never seen during training without degrading attention quality.

### Cache-Friendly Pre-computation

Frequency tables are cached per `(dim, seq_len, device, dtype)` tuple (lines 102-123 of [`rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/rope.py)). This eliminates the `positions.max()` host-to-device synchronization that would otherwise break CUDA graph capture, ensuring the embedding remains inexpensive even for streaming video processing.

### Linear-Time Positional Bias

The rotation applies as element-wise multiplication via `apply_rotary_emb` (lines 35-61), yielding **O(N·D)** complexity where *N* is the token count and *D* is the head dimension. Unlike absolute or learned positional encodings, Video RoPE requires no quadratic positional-bias matrix, keeping memory consumption linear with sequence length.

### Temporal Consistency in KV-Cache Streaming

For streaming inference, [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) (lines 24-30) and `CameraCausalHead` (lines 25-38) use `global_frame_idx` to generate the correct temporal slice. This ensures cached keys and values maintain alignment across frames, enabling causal inference on arbitrarily long videos without recomputing previous frame embeddings.

## Practical Implementation Examples

### Enabling 3D RoPE for Video Heads

To instantiate a causal head with Video RoPE support:

```python

# camera_head.py – constructing with 3-D RoPE

self.rope3d = WanRotaryPosEmbed(
    attention_head_dim=head_dim,          # dim_in // num_heads

    patch_size=(max_frame_num, 1, 1),    # (frames, height, width)

    theta=rope_theta,                    # typically 10_000.0

    fhw_dim=[40, 44, 44]                # explicit (t, h, w) dim split

)

```

The head receives positional tensors of shape `[1, 1, S, head_dim//2]` via the `pos` argument in `CameraCausalHead.trunk_fn`.

### Streaming Pipeline Integration

For processing video chunks in a streaming context:

```python

# stream.py – generating RoPE for each new chunk

if self.enable_3d_rope and self.rope3d is not None:
    pos3d = self.rope3d(
        ppf=S, pph=H, ppw=W,
        patch_start_idx=0,
        device=x.device,
        f_start=self.global_frame_idx,
        f_end=self.global_frame_idx + S
    )

```

The `global_frame_idx` tracker ensures proper temporal offsets for the KV cache across arbitrarily long sequences.

## Summary

- **Video RoPE** extends rotary embeddings to three dimensions (temporal, height, width) via independent 1D frequency tables concatenated in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py)
- **NTK-aware scaling** (`ntk_factor`) allows extrapolation to sequences longer than training data without retraining by stretching base frequencies
- **Linear complexity** is maintained through element-wise rotary multiplication (`apply_rotary_emb`) avoiding quadratic memory costs
- **Streaming compatibility** is achieved through frame-aware slicing in [`stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/stream.py) and [`camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/camera_head.py), enabling causal processing of unlimited video length
- The implementation falls back to 2D spatial RoPE when `enable_3d_rope=False`, controlled via CLI flags in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py)

## Frequently Asked Questions

### What makes Video RoPE different from standard 2D rotary embeddings?

Video RoPE adds a third temporal dimension to the traditional height and width spatial coordinates. While 2D RoPE encodes patch positions within static images, Video RoPE in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) constructs separate frequency tables for temporal frames, enabling the attention mechanism to understand relative motion and temporal ordering across video sequences.

### How does Video RoPE handle video sequences longer than the training context?

The implementation uses NTK-aware frequency scaling via the `ntk_factor` parameter in `get_1d_rotary_pos_embed` (line 242). By stretching the base frequency `theta`, the sinusoidal encodings maintain proper phase relationships for token positions beyond the training sequence length, allowing the model to generalize to arbitrarily long videos without positional aliasing.

### Why is Video RoPE more efficient than absolute positional encodings for long videos?

Video RoPE applies positional information through complex number rotation (element-wise multiplication in `apply_rotary_emb`) rather than additive bias vectors or learned embeddings. This approach scales linearly **O(N·D)** with sequence length and requires no additional memory for positional matrices, whereas absolute encodings often require quadratic attention bias computation or fixed-length position indices that cannot extrapolate.

### Can Video RoPE be disabled or used with 2D-only data?

Yes, the codebase supports a fallback to 2D spatial RoPE via the `enable_3d_rope` boolean flag found in [`attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/attention.py) and [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py). When disabled, the model uses standard rotary embeddings on spatial coordinates only, making the architecture compatible with both video (3D) and image (2D) inputs without architectural changes.