Video RoPE Positional Encoding: 3D Rotary Embeddings for Long-Sequence Video in Ling-Bot-Map

Video RoPE is a three-dimensional rotary position encoding that independently encodes temporal, vertical, and horizontal coordinates, enabling transformers to model long video sequences with linear complexity and NTK-aware extrapolation.

Video RoPE positional encoding extends classic transformer rotary embeddings into three dimensions to capture spatio-temporal relationships in video data. In the Robbyant/lingbot-map repository, this mechanism is implemented through the WanRotaryPosEmbed class in lingbot_map/layers/rope.py, providing a computationally efficient alternative to absolute positional encodings for high-resolution video generation.

What is Video RoPE Positional Encoding?

Video RoPE (Rotary Position Embedding) generalizes the standard 2D rotary embedding to handle the three-dimensional grid of video data: frames × height × width. Unlike spatial-only encodings, Video RoPE treats the temporal dimension as a continuous rotary coordinate, preserving relative positional information across frame sequences while maintaining the translation-invariant properties of rotary embeddings.

The Three Rotary Dimensions

The encoding decomposes video coordinates into three independent 1D rotary embeddings that are concatenated into a single complex-valued frequency tensor of shape [max_seq_len, head_dim//2]:

  • t (temporal) – Encodes the frame index using a dedicated 1D RoPE on the time head-dimension, capturing motion between frames
  • h (height) – Encodes the patch row within a frame via vertical spatial rotation
  • w (width) – Encodes the patch column via horizontal spatial rotation

In lingbot_map/layers/rope.py (lines 202-276), the get_1d_rotary_pos_embed function builds these three frequency tables independently, which WanRotaryPosEmbed then concatenates using torch.cat to form the final embedding. During inference, the forward method (lines 336-393) slices this tensor according to the current frame range (f_start, f_end) and broadcasts to shape [1, 1, N, head_dim//2].

Implementation in the Ling-Bot-Map Codebase

The integration of Video RoPE spans several key modules, with conditional logic that falls back to 2D spatial encoding when enable_3d_rope is disabled.

Core RoPE Module

The WanRotaryPosEmbed class in lingbot_map/layers/rope.py manages the pre-computation and caching of frequency tables. The class caches tables per (dim, seq_len, device, dtype) tuple (lines 102-123), avoiding expensive host-to-device synchronizations that would break CUDA graph capture.

Attention Integration

Inside lingbot_map/layers/attention.py (lines 154-157 and 216-220), the model conditionally applies Video RoPE:


# attention.py – Q/K preparation with conditional 3D RoPE

if self.rope is not None and not enable_3d_rope:
    q = self.rope(q, pos)  # 2-D spatial RoPE

    k = self.rope(k, pos)
elif enable_3d_rope and pos is not None:
    q = self.rope(q, pos)  # 3-D spatio-temporal RoPE

    k = self.rope(k, pos)

When enabled, the same WanRotaryPosEmbed instance receives a 3D positional tensor containing temporal and spatial coordinates.

How Video RoPE Affects Long-Sequence Performance

Video RoPE enables efficient processing of video sequences far longer than training data through four key mechanisms that maintain computational linearity and numerical stability.

NTK-Aware Frequency Scaling

The get_1d_rotary_pos_embed function accepts an ntk_factor argument (lines 42-44) that scales the base frequency theta (line 242: theta = theta * ntk_factor). This Neural-Tangents-Kernel scaling stretches the sinusoidal spectrum, allowing the embedding to extrapolate to sequence lengths never seen during training without degrading attention quality.

Cache-Friendly Pre-computation

Frequency tables are cached per (dim, seq_len, device, dtype) tuple (lines 102-123 of rope.py). This eliminates the positions.max() host-to-device synchronization that would otherwise break CUDA graph capture, ensuring the embedding remains inexpensive even for streaming video processing.

Linear-Time Positional Bias

The rotation applies as element-wise multiplication via apply_rotary_emb (lines 35-61), yielding O(N·D) complexity where N is the token count and D is the head dimension. Unlike absolute or learned positional encodings, Video RoPE requires no quadratic positional-bias matrix, keeping memory consumption linear with sequence length.

Temporal Consistency in KV-Cache Streaming

For streaming inference, lingbot_map/aggregator/stream.py (lines 24-30) and CameraCausalHead (lines 25-38) use global_frame_idx to generate the correct temporal slice. This ensures cached keys and values maintain alignment across frames, enabling causal inference on arbitrarily long videos without recomputing previous frame embeddings.

Practical Implementation Examples

Enabling 3D RoPE for Video Heads

To instantiate a causal head with Video RoPE support:


# camera_head.py – constructing with 3-D RoPE

self.rope3d = WanRotaryPosEmbed(
    attention_head_dim=head_dim,          # dim_in // num_heads

    patch_size=(max_frame_num, 1, 1),    # (frames, height, width)

    theta=rope_theta,                    # typically 10_000.0

    fhw_dim=[40, 44, 44]                # explicit (t, h, w) dim split

)

The head receives positional tensors of shape [1, 1, S, head_dim//2] via the pos argument in CameraCausalHead.trunk_fn.

Streaming Pipeline Integration

For processing video chunks in a streaming context:


# stream.py – generating RoPE for each new chunk

if self.enable_3d_rope and self.rope3d is not None:
    pos3d = self.rope3d(
        ppf=S, pph=H, ppw=W,
        patch_start_idx=0,
        device=x.device,
        f_start=self.global_frame_idx,
        f_end=self.global_frame_idx + S
    )

The global_frame_idx tracker ensures proper temporal offsets for the KV cache across arbitrarily long sequences.

Summary

  • Video RoPE extends rotary embeddings to three dimensions (temporal, height, width) via independent 1D frequency tables concatenated in lingbot_map/layers/rope.py
  • NTK-aware scaling (ntk_factor) allows extrapolation to sequences longer than training data without retraining by stretching base frequencies
  • Linear complexity is maintained through element-wise rotary multiplication (apply_rotary_emb) avoiding quadratic memory costs
  • Streaming compatibility is achieved through frame-aware slicing in stream.py and camera_head.py, enabling causal processing of unlimited video length
  • The implementation falls back to 2D spatial RoPE when enable_3d_rope=False, controlled via CLI flags in demo.py

Frequently Asked Questions

What makes Video RoPE different from standard 2D rotary embeddings?

Video RoPE adds a third temporal dimension to the traditional height and width spatial coordinates. While 2D RoPE encodes patch positions within static images, Video RoPE in lingbot_map/layers/rope.py constructs separate frequency tables for temporal frames, enabling the attention mechanism to understand relative motion and temporal ordering across video sequences.

How does Video RoPE handle video sequences longer than the training context?

The implementation uses NTK-aware frequency scaling via the ntk_factor parameter in get_1d_rotary_pos_embed (line 242). By stretching the base frequency theta, the sinusoidal encodings maintain proper phase relationships for token positions beyond the training sequence length, allowing the model to generalize to arbitrarily long videos without positional aliasing.

Why is Video RoPE more efficient than absolute positional encodings for long videos?

Video RoPE applies positional information through complex number rotation (element-wise multiplication in apply_rotary_emb) rather than additive bias vectors or learned embeddings. This approach scales linearly O(N·D) with sequence length and requires no additional memory for positional matrices, whereas absolute encodings often require quadratic attention bias computation or fixed-length position indices that cannot extrapolate.

Can Video RoPE be disabled or used with 2D-only data?

Yes, the codebase supports a fallback to 2D spatial RoPE via the enable_3d_rope boolean flag found in attention.py and demo.py. When disabled, the model uses standard rotary embeddings on spatial coordinates only, making the architecture compatible with both video (3D) and image (2D) inputs without architectural changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →