3D RoPE in LingBot-Map for Temporal Consistency: Implementation and Usage Guide
3D RoPE in LingBot-Map extends rotary position embeddings to three dimensions (time, height, width) using global frame indices, ensuring that key-value cache entries remain temporally aligned during streaming video inference.
Streaming SLAM and map-building transformers require stable temporal representations across variable-length frame sequences. The LingBot-Map repository solves this by implementing 3D Rotary Position Embedding (RoPE), which encodes absolute temporal positions alongside spatial coordinates to prevent the representation drift that occurs with relative window indexing.
What Is 3D RoPE in LingBot-Map?
Extending 2D Rotary Embeddings to Three Dimensions
Standard 2D RoPE processes spatial dimensions by treating height and width independently, concatenating the rotated feature vectors for each axis. In lingbot_map/layers/rope.py, this approach handles patch grids within single frames.
3D RoPE augments this with a temporal dimension (frame index). The implementation resides in the WanRotaryPosEmbed class, which partitions the attention head dimension into three sub-components: time (t), height (h), and width (w).
The WanRotaryPosEmbed Architecture
The constructor receives the attention head dimension and patch size (max_frames, h, w), then splits the embedding space according to fhw_dim (default [20, 22, 22]) or an auto-computed split:
# lingbot_map/layers/rope.py
class WanRotaryPosEmbed(nn.Module):
"""3D rotary position embedding (t, h, w)."""
def __init__(self, attention_head_dim, patch_size, max_seq_len=1024,
theta=10000.0, fhw_dim=[20, 22, 22]):
# Split head_dim into t-, h-, w-sub-dimensions
# Pre-compute separate 1-D RoPE tables for each axis (get_1d_rotary_pos_embed)
# Store them in self.freqs (complex-valued tensor)
The constraint t_dim + h_dim + w_dim = head_dim ensures full utilization of the embedding space. For each axis, get_1d_rotary_pos_embed builds sinusoidal frequency tables, which are concatenated into self.freqs and indexed using frame-row-column triples.
How 3D RoPE Ensures Temporal Consistency
Temporal consistency requires that cached keys and queries align to the same global timeline, regardless of whether the model processes a single frame or a sliding window. The mechanism relies on absolute frame indexing rather than relative window positions.
Global Frame Index Propagation
A model-level boolean flag enable_3d_rope propagates through the architecture hierarchy:
GCTBase(base model class)GCTStream*(streaming variants)AggregatorStream(KV cache manager)CameraCausalHead(position generation)
This flag originates in lingbot_map/models/gct_base.py (line 56) and passes through constructor arguments in each downstream component.
Initialization and Position Generation
When AggregatorStream initializes with enable_3d_rope=True, it invokes _init_3d_rope() (lines 101-104 of lingbot_map/aggregator/stream.py) to instantiate a shared WanRotaryPosEmbed module. The CameraCausalHead receives this module during its own initialization (lines 86-108 of lingbot_map/heads/camera_head.py).
During streaming inference, the camera head generates frame-wise positional tensors using absolute indices:
# lingbot_map/heads/camera_head.py (excerpt)
if self.rope3d is not None:
pos3d = self.rope3d(
ppf=S, # Number of frames in current chunk
pph=1, ppw=1, # Camera token occupies 1×1 patch
patch_start_idx=0,
device=pose_tokens.device,
f_start=self.frame_idx, # Global start frame
f_end=self.frame_idx + S, # Global end frame
)
KV Cache Alignment Mechanism
The WanRotaryPosEmbed.__call__ method splits pre-computed frequencies into t, h, and w components, selects the slice for the current frame range (frame_slice), and returns a complex-valued tensor of shape [1, 1, S, head_dim//2].
This tensor is applied to queries and keys before they enter the KV cache (via the pos=pos3d argument). Consequently, all cached keys carry the correct temporal rotation for their absolute frame index. When the cache is read in subsequent steps, the temporal positions remain consistent across different sliding windows because they reference the same global timeline.
Preventing Double Rotation
The attention mechanism in lingbot_map/layers/attention.py (lines 154-159) checks if enable_3d_rope before applying standard 2D RoPE. This conditional guarantees that temporal rotation is applied exactly once during the initial forward pass and never reapplied to cached keys during retrieval.
Implementation Details and Source Code
The following files contain the core 3D RoPE implementation:
-
lingbot_map/layers/rope.py– ContainsWanRotaryPosEmbedwithget_1d_rotary_pos_embedhelper and frequency table concatenation logic. -
lingbot_map/aggregator/stream.py– Implements_init_3d_rope()(lines 101-104) to initialize the shared 3D RoPE module for the streaming aggregator. -
lingbot_map/heads/camera_head.py– Instantiates the 3D RoPE module in__init__(lines 86-108) and generates position tensors during forward passes usingself.frame_idxtracking. -
lingbot_map/layers/attention.py– Contains the conditional logic (lines 154-159) that disables 2D RoPE when 3D RoPE is active to prevent double application. -
lingbot_map/models/gct_base.py– Defines theenable_3d_ropeconstructor argument (line 56) that propagates through the model hierarchy.
Practical Usage Example
Enable 3D RoPE when constructing a streaming model to activate temporal consistency:
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream
from lingbot_map.utils.geometry import setup_dummy_input
# 1. Enable 3-D RoPE when building the model
model = GCTStream(
img_size=518,
patch_size=14,
embed_dim=1024,
disable_global_rope=False,
enable_3d_rope=True, # Activate temporal consistency
max_frame_num=1024,
)
# 2. Dummy streaming batch (B=1, 3 frames, 1 camera token per frame)
tokens, mask = setup_dummy_input(num_frames=3)
# 3. Forward pass with causal inference
with torch.no_grad():
out = model(tokens, mask=mask, causal_inference=True)
print(out.shape) # → (B, S, output_dim) with temporally consistent embeddings
The setup_dummy_input helper creates tensors shaped [B, S, C] where S represents the number of frames. In production, these tokens originate from the vision backbone and maintain global frame indices across sliding window boundaries.
Summary
- 3D RoPE extends rotary embeddings to independent time, height, and width axes using
WanRotaryPosEmbedinlingbot_map/layers/rope.py. - Global frame indices ensure that KV cache entries align temporally across different processing windows, eliminating drift.
- Flag-controlled activation via
enable_3d_ropepropagates fromGCTBasethroughAggregatorStreamtoCameraCausalHead. - Single application guarantee prevents double rotation through conditional checks in the attention layer.
- Streaming stability is achieved by computing complex-valued position tensors before caching, using absolute
frame_idxvalues.
Frequently Asked Questions
What is the difference between 2D and 3D RoPE in LingBot-Map?
2D RoPE processes only spatial dimensions (height and width) independently, suitable for single-frame image processing. 3D RoPE adds a temporal dimension that encodes absolute frame indices, enabling consistent position embeddings across video sequences. The implementation splits the attention head dimension into t_dim, h_dim, and w_dim components to accommodate the additional axis.
How does 3D RoPE maintain consistency across sliding windows?
The system uses global frame indices rather than relative window positions. When generating position tensors in CameraCausalHead, the f_start parameter references self.frame_idx—an absolute counter that persists across inference steps. Because the rotation is a deterministic function of this global index, the same physical frame receives identical embeddings regardless of which sliding window processes it, ensuring cached keys remain aligned.
Where is the 3D RoPE module instantiated in the architecture?
The module is instantiated in AggregatorStream._init_3d_rope() (lines 101-104 of lingbot_map/aggregator/stream.py) when enable_3d_rope=True. It is then passed to CameraCausalHead during initialization (lines 86-108 of lingbot_map/heads/camera_head.py). The base model class GCTBase defines the flag in its constructor (line 56 of lingbot_map/models/gct_base.py), controlling whether the 3D pathway is activated.
Can 3D RoPE be disabled for standard (non-streaming) inference?
Yes. Setting enable_3d_rope=False (the default) disables the 3D pathway, causing the model to use standard 2D RoPE instead. This is controlled via the constructor argument in GCTBase and propagated through all sub-modules. When disabled, the attention mechanism applies standard spatial rotary embeddings without temporal components, suitable for single-frame or non-causal batch processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →