3D RoPE in LingBot-Map for Temporal Consistency: Implementation and Usage Guide

3D RoPE in LingBot-Map extends rotary position embeddings to three dimensions (time, height, width) using global frame indices, ensuring that key-value cache entries remain temporally aligned during streaming video inference.

Streaming SLAM and map-building transformers require stable temporal representations across variable-length frame sequences. The LingBot-Map repository solves this by implementing 3D Rotary Position Embedding (RoPE), which encodes absolute temporal positions alongside spatial coordinates to prevent the representation drift that occurs with relative window indexing.

What Is 3D RoPE in LingBot-Map?

Extending 2D Rotary Embeddings to Three Dimensions

Standard 2D RoPE processes spatial dimensions by treating height and width independently, concatenating the rotated feature vectors for each axis. In lingbot_map/layers/rope.py, this approach handles patch grids within single frames.

3D RoPE augments this with a temporal dimension (frame index). The implementation resides in the WanRotaryPosEmbed class, which partitions the attention head dimension into three sub-components: time (t), height (h), and width (w).

The WanRotaryPosEmbed Architecture

The constructor receives the attention head dimension and patch size (max_frames, h, w), then splits the embedding space according to fhw_dim (default [20, 22, 22]) or an auto-computed split:


# lingbot_map/layers/rope.py

class WanRotaryPosEmbed(nn.Module):
    """3D rotary position embedding (t, h, w)."""
    def __init__(self, attention_head_dim, patch_size, max_seq_len=1024,
                 theta=10000.0, fhw_dim=[20, 22, 22]):
        # Split head_dim into t-, h-, w-sub-dimensions

        # Pre-compute separate 1-D RoPE tables for each axis (get_1d_rotary_pos_embed)

        # Store them in self.freqs (complex-valued tensor)

The constraint t_dim + h_dim + w_dim = head_dim ensures full utilization of the embedding space. For each axis, get_1d_rotary_pos_embed builds sinusoidal frequency tables, which are concatenated into self.freqs and indexed using frame-row-column triples.

How 3D RoPE Ensures Temporal Consistency

Temporal consistency requires that cached keys and queries align to the same global timeline, regardless of whether the model processes a single frame or a sliding window. The mechanism relies on absolute frame indexing rather than relative window positions.

Global Frame Index Propagation

A model-level boolean flag enable_3d_rope propagates through the architecture hierarchy:

  • GCTBase (base model class)
  • GCTStream* (streaming variants)
  • AggregatorStream (KV cache manager)
  • CameraCausalHead (position generation)

This flag originates in lingbot_map/models/gct_base.py (line 56) and passes through constructor arguments in each downstream component.

Initialization and Position Generation

When AggregatorStream initializes with enable_3d_rope=True, it invokes _init_3d_rope() (lines 101-104 of lingbot_map/aggregator/stream.py) to instantiate a shared WanRotaryPosEmbed module. The CameraCausalHead receives this module during its own initialization (lines 86-108 of lingbot_map/heads/camera_head.py).

During streaming inference, the camera head generates frame-wise positional tensors using absolute indices:


# lingbot_map/heads/camera_head.py (excerpt)

if self.rope3d is not None:
    pos3d = self.rope3d(
        ppf=S,               # Number of frames in current chunk

        pph=1, ppw=1,        # Camera token occupies 1×1 patch

        patch_start_idx=0,
        device=pose_tokens.device,
        f_start=self.frame_idx,         # Global start frame

        f_end=self.frame_idx + S,       # Global end frame

    )

KV Cache Alignment Mechanism

The WanRotaryPosEmbed.__call__ method splits pre-computed frequencies into t, h, and w components, selects the slice for the current frame range (frame_slice), and returns a complex-valued tensor of shape [1, 1, S, head_dim//2].

This tensor is applied to queries and keys before they enter the KV cache (via the pos=pos3d argument). Consequently, all cached keys carry the correct temporal rotation for their absolute frame index. When the cache is read in subsequent steps, the temporal positions remain consistent across different sliding windows because they reference the same global timeline.

Preventing Double Rotation

The attention mechanism in lingbot_map/layers/attention.py (lines 154-159) checks if enable_3d_rope before applying standard 2D RoPE. This conditional guarantees that temporal rotation is applied exactly once during the initial forward pass and never reapplied to cached keys during retrieval.

Implementation Details and Source Code

The following files contain the core 3D RoPE implementation:

  • lingbot_map/layers/rope.py – Contains WanRotaryPosEmbed with get_1d_rotary_pos_embed helper and frequency table concatenation logic.

  • lingbot_map/aggregator/stream.py – Implements _init_3d_rope() (lines 101-104) to initialize the shared 3D RoPE module for the streaming aggregator.

  • lingbot_map/heads/camera_head.py – Instantiates the 3D RoPE module in __init__ (lines 86-108) and generates position tensors during forward passes using self.frame_idx tracking.

  • lingbot_map/layers/attention.py – Contains the conditional logic (lines 154-159) that disables 2D RoPE when 3D RoPE is active to prevent double application.

  • lingbot_map/models/gct_base.py – Defines the enable_3d_rope constructor argument (line 56) that propagates through the model hierarchy.

Practical Usage Example

Enable 3D RoPE when constructing a streaming model to activate temporal consistency:

import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream
from lingbot_map.utils.geometry import setup_dummy_input

# 1. Enable 3-D RoPE when building the model

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    disable_global_rope=False,
    enable_3d_rope=True,               # Activate temporal consistency

    max_frame_num=1024,
)

# 2. Dummy streaming batch (B=1, 3 frames, 1 camera token per frame)

tokens, mask = setup_dummy_input(num_frames=3)

# 3. Forward pass with causal inference

with torch.no_grad():
    out = model(tokens, mask=mask, causal_inference=True)

print(out.shape)   # → (B, S, output_dim) with temporally consistent embeddings

The setup_dummy_input helper creates tensors shaped [B, S, C] where S represents the number of frames. In production, these tokens originate from the vision backbone and maintain global frame indices across sliding window boundaries.

Summary

  • 3D RoPE extends rotary embeddings to independent time, height, and width axes using WanRotaryPosEmbed in lingbot_map/layers/rope.py.
  • Global frame indices ensure that KV cache entries align temporally across different processing windows, eliminating drift.
  • Flag-controlled activation via enable_3d_rope propagates from GCTBase through AggregatorStream to CameraCausalHead.
  • Single application guarantee prevents double rotation through conditional checks in the attention layer.
  • Streaming stability is achieved by computing complex-valued position tensors before caching, using absolute frame_idx values.

Frequently Asked Questions

What is the difference between 2D and 3D RoPE in LingBot-Map?

2D RoPE processes only spatial dimensions (height and width) independently, suitable for single-frame image processing. 3D RoPE adds a temporal dimension that encodes absolute frame indices, enabling consistent position embeddings across video sequences. The implementation splits the attention head dimension into t_dim, h_dim, and w_dim components to accommodate the additional axis.

How does 3D RoPE maintain consistency across sliding windows?

The system uses global frame indices rather than relative window positions. When generating position tensors in CameraCausalHead, the f_start parameter references self.frame_idx—an absolute counter that persists across inference steps. Because the rotation is a deterministic function of this global index, the same physical frame receives identical embeddings regardless of which sliding window processes it, ensuring cached keys remain aligned.

Where is the 3D RoPE module instantiated in the architecture?

The module is instantiated in AggregatorStream._init_3d_rope() (lines 101-104 of lingbot_map/aggregator/stream.py) when enable_3d_rope=True. It is then passed to CameraCausalHead during initialization (lines 86-108 of lingbot_map/heads/camera_head.py). The base model class GCTBase defines the flag in its constructor (line 56 of lingbot_map/models/gct_base.py), controlling whether the 3D pathway is activated.

Can 3D RoPE be disabled for standard (non-streaming) inference?

Yes. Setting enable_3d_rope=False (the default) disables the 3D pathway, causing the model to use standard 2D RoPE instead. This is controlled via the constructor argument in GCTBase and propagated through all sub-modules. When disabled, the attention mechanism applies standard spatial rotary embeddings without temporal components, suitable for single-frame or non-causal batch processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →