# 3D RoPE in LingBot-Map for Temporal Consistency: Implementation and Usage Guide

> Learn how 3D RoPE in LingBot-Map ensures temporal consistency for streaming video inference by extending rotary position embeddings to three dimensions. Explore implementation and usage.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-27

---

**3D RoPE in LingBot-Map extends rotary position embeddings to three dimensions (time, height, width) using global frame indices, ensuring that key-value cache entries remain temporally aligned during streaming video inference.**

Streaming SLAM and map-building transformers require stable temporal representations across variable-length frame sequences. The **LingBot-Map** repository solves this by implementing **3D Rotary Position Embedding (RoPE)**, which encodes absolute temporal positions alongside spatial coordinates to prevent the representation drift that occurs with relative window indexing.

## What Is 3D RoPE in LingBot-Map?

### Extending 2D Rotary Embeddings to Three Dimensions

Standard **2D RoPE** processes spatial dimensions by treating **height** and **width** independently, concatenating the rotated feature vectors for each axis. In [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py), this approach handles patch grids within single frames.

**3D RoPE** augments this with a **temporal** dimension (frame index). The implementation resides in the `WanRotaryPosEmbed` class, which partitions the attention head dimension into three sub-components: time (`t`), height (`h`), and width (`w`).

### The WanRotaryPosEmbed Architecture

The constructor receives the **attention head dimension** and **patch size** `(max_frames, h, w)`, then splits the embedding space according to `fhw_dim` (default `[20, 22, 22]`) or an auto-computed split:

```python

# lingbot_map/layers/rope.py

class WanRotaryPosEmbed(nn.Module):
    """3D rotary position embedding (t, h, w)."""
    def __init__(self, attention_head_dim, patch_size, max_seq_len=1024,
                 theta=10000.0, fhw_dim=[20, 22, 22]):
        # Split head_dim into t-, h-, w-sub-dimensions

        # Pre-compute separate 1-D RoPE tables for each axis (get_1d_rotary_pos_embed)

        # Store them in self.freqs (complex-valued tensor)

```

The constraint `t_dim + h_dim + w_dim = head_dim` ensures full utilization of the embedding space. For each axis, `get_1d_rotary_pos_embed` builds sinusoidal frequency tables, which are concatenated into `self.freqs` and indexed using frame-row-column triples.

## How 3D RoPE Ensures Temporal Consistency

Temporal consistency requires that cached keys and queries align to the same **global timeline**, regardless of whether the model processes a single frame or a sliding window. The mechanism relies on absolute frame indexing rather than relative window positions.

### Global Frame Index Propagation

A model-level boolean flag `enable_3d_rope` propagates through the architecture hierarchy:

- `GCTBase` (base model class)
- `GCTStream*` (streaming variants)
- `AggregatorStream` (KV cache manager)
- `CameraCausalHead` (position generation)

This flag originates in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py) (line 56) and passes through constructor arguments in each downstream component.

### Initialization and Position Generation

When `AggregatorStream` initializes with `enable_3d_rope=True`, it invokes `_init_3d_rope()` (lines 101-104 of [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py)) to instantiate a shared `WanRotaryPosEmbed` module. The `CameraCausalHead` receives this module during its own initialization (lines 86-108 of [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py)).

During streaming inference, the camera head generates frame-wise positional tensors using absolute indices:

```python

# lingbot_map/heads/camera_head.py (excerpt)

if self.rope3d is not None:
    pos3d = self.rope3d(
        ppf=S,               # Number of frames in current chunk

        pph=1, ppw=1,        # Camera token occupies 1×1 patch

        patch_start_idx=0,
        device=pose_tokens.device,
        f_start=self.frame_idx,         # Global start frame

        f_end=self.frame_idx + S,       # Global end frame

    )

```

### KV Cache Alignment Mechanism

The `WanRotaryPosEmbed.__call__` method splits pre-computed frequencies into `t`, `h`, and `w` components, selects the slice for the current frame range (`frame_slice`), and returns a complex-valued tensor of shape `[1, 1, S, head_dim//2]`.

This tensor is applied to queries and keys **before** they enter the KV cache (via the `pos=pos3d` argument). Consequently, all cached keys carry the correct temporal rotation for their absolute frame index. When the cache is read in subsequent steps, the temporal positions remain consistent across different sliding windows because they reference the same global timeline.

### Preventing Double Rotation

The attention mechanism in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) (lines 154-159) checks `if enable_3d_rope` before applying standard 2D RoPE. This conditional guarantees that temporal rotation is applied **exactly once** during the initial forward pass and never reapplied to cached keys during retrieval.

## Implementation Details and Source Code

The following files contain the core 3D RoPE implementation:

- **[`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py)** – Contains `WanRotaryPosEmbed` with `get_1d_rotary_pos_embed` helper and frequency table concatenation logic.

- **[`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py)** – Implements `_init_3d_rope()` (lines 101-104) to initialize the shared 3D RoPE module for the streaming aggregator.

- **[`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py)** – Instantiates the 3D RoPE module in `__init__` (lines 86-108) and generates position tensors during forward passes using `self.frame_idx` tracking.

- **[`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py)** – Contains the conditional logic (lines 154-159) that disables 2D RoPE when 3D RoPE is active to prevent double application.

- **[`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py)** – Defines the `enable_3d_rope` constructor argument (line 56) that propagates through the model hierarchy.

## Practical Usage Example

Enable 3D RoPE when constructing a streaming model to activate temporal consistency:

```python
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream
from lingbot_map.utils.geometry import setup_dummy_input

# 1. Enable 3-D RoPE when building the model

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    disable_global_rope=False,
    enable_3d_rope=True,               # Activate temporal consistency

    max_frame_num=1024,
)

# 2. Dummy streaming batch (B=1, 3 frames, 1 camera token per frame)

tokens, mask = setup_dummy_input(num_frames=3)

# 3. Forward pass with causal inference

with torch.no_grad():
    out = model(tokens, mask=mask, causal_inference=True)

print(out.shape)   # → (B, S, output_dim) with temporally consistent embeddings

```

The `setup_dummy_input` helper creates tensors shaped `[B, S, C]` where `S` represents the number of frames. In production, these tokens originate from the vision backbone and maintain global frame indices across sliding window boundaries.

## Summary

- **3D RoPE** extends rotary embeddings to independent time, height, and width axes using `WanRotaryPosEmbed` in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py).
- **Global frame indices** ensure that KV cache entries align temporally across different processing windows, eliminating drift.
- **Flag-controlled activation** via `enable_3d_rope` propagates from `GCTBase` through `AggregatorStream` to `CameraCausalHead`.
- **Single application** guarantee prevents double rotation through conditional checks in the attention layer.
- **Streaming stability** is achieved by computing complex-valued position tensors before caching, using absolute `frame_idx` values.

## Frequently Asked Questions

### What is the difference between 2D and 3D RoPE in LingBot-Map?

**2D RoPE** processes only spatial dimensions (height and width) independently, suitable for single-frame image processing. **3D RoPE** adds a temporal dimension that encodes absolute frame indices, enabling consistent position embeddings across video sequences. The implementation splits the attention head dimension into `t_dim`, `h_dim`, and `w_dim` components to accommodate the additional axis.

### How does 3D RoPE maintain consistency across sliding windows?

The system uses **global frame indices** rather than relative window positions. When generating position tensors in `CameraCausalHead`, the `f_start` parameter references `self.frame_idx`—an absolute counter that persists across inference steps. Because the rotation is a deterministic function of this global index, the same physical frame receives identical embeddings regardless of which sliding window processes it, ensuring cached keys remain aligned.

### Where is the 3D RoPE module instantiated in the architecture?

The module is instantiated in `AggregatorStream._init_3d_rope()` (lines 101-104 of [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py)) when `enable_3d_rope=True`. It is then passed to `CameraCausalHead` during initialization (lines 86-108 of [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py)). The base model class `GCTBase` defines the flag in its constructor (line 56 of [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py)), controlling whether the 3D pathway is activated.

### Can 3D RoPE be disabled for standard (non-streaming) inference?

Yes. Setting `enable_3d_rope=False` (the default) disables the 3D pathway, causing the model to use standard 2D RoPE instead. This is controlled via the constructor argument in `GCTBase` and propagated through all sub-modules. When disabled, the attention mechanism applies standard spatial rotary embeddings without temporal components, suitable for single-frame or non-causal batch processing.