# How 3D RoPE Positional Encoding Works with the Paged KV Cache in LingBot-Map

> Discover how 3D RoPE positional encoding pre-computes rotary embeddings for paged KV cache. Learn how LingBot-Map achieves efficient streaming video inference by avoiding recomputation.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-31

---

**3D RoPE positional encoding pre-computes frame-height-width rotary embeddings that are baked into token features before entering the paged KV cache, enabling FlashInfer to retrieve position-aware keys and values without positional recomputation during streaming video inference.**

LingBot-Map implements a streaming 3D transformer architecture that processes video frames sequentially using a novel integration of 3D Rotary Position Embeddings (RoPE) and a dual-stream paged KV cache. This design, found in the `Robbyant/lingbot-map` repository, enables efficient inference on extremely long video sequences by computing positional encodings upfront and storing position-aware tensors directly in reusable memory pages.

## 3D Rotary Position Embedding Architecture

The **3D RoPE encoder** in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) generates rotary position embeddings across three spatial axes: frame (t), height (h), and width (w). This allows the model to understand temporal relationships in video streams rather than treating each frame as an independent image.

### Multi-Axis Frequency Generation

The `WanRotaryPosEmbed` class creates independent 1D RoPE frequency tables for each axis using `get_1d_rotary_pos_embed`, then concatenates these three tables to form a unified 3D positional representation. The resulting complex-valued tensor has shape:

```

[1, 1, ppf * (patch_start_idx + pph * ppw), head_dim//2]

```

Where **ppf** represents patches-per-frame, **pph** and **ppw** denote the patch grid dimensions, and **patch_start_idx** accounts for special tokens preceding the patch sequence.

### Complex-Valued Tensor Application

The `forward` method returns a complex-valued tensor that is applied to real-valued token features via `apply_rotary_emb`. This function performs complex multiplication using pure real arithmetic, ensuring compatibility with Torch-Compile and CUDA graphs while maintaining mathematical equivalence to standard complex operations.

## Paged KV Cache Design

The **paged KV cache** in [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py) implements a two-stream memory layout that separates patch tokens from special tokens to optimize memory reuse and context preservation across thousands of frames.

### Dual-Stream Memory Layout

The `FlashInferKVCacheManager` maintains two distinct streams:

- **Patch stream**: Stores `patches_per_frame` tokens per page using a recyclable sliding window. The first `scale_frames` pages remain permanent, while subsequent pages form a `sliding_window` that evicts oldest frames back to a free list when full.
- **Special stream**: An append-only stream holding six per-frame special tokens (camera, register, and scale tokens). These pages are never evicted, ensuring global context remains available for cross-frame attention.

### Frame Appending and Position Encoding Integration

When `append_frame` receives new frame data, it splits the incoming tensor into special components (`sp_k`, `sp_v`) and patch components (`patch_k`, `patch_v`). Because the 3D RoPE encoding is applied to token features **before** this write operation (as implemented in [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) and [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py)), the cached values already carry their positional information.

The `_write_patch_page` method stores patch tokens in the recyclable stream, while `_write_special_tokens` appends special tokens to their dedicated stream. This separation allows the system to evict visual patch history while preserving critical metadata tokens.

## Attention Computation with Cached RoPE

During `compute_attention`, FlashInfer's `BatchPrefillWithPagedKVCacheWrapper` receives a visible page table constructed as:

```

scale_patch_pages → live_window_patch_pages → all_special_pages

```

The wrapper respects `compute_last_page_len` to handle partially-filled special pages correctly. Since RoPE is already encoded in the cached K/V tensors, the attention mechanism retrieves position-aware values directly without recomputing embeddings, regardless of how many times the patch pages are recycled.

## Implementation Example

The following code demonstrates the complete workflow from 3D RoPE generation to paged cache utilization:

```python
from lingbot_map.layers.rope import WanRotaryPosEmbed, apply_rotary_emb
from lingbot_map.layers.flashinfer_cache import FlashInferKVCacheManager
import torch

# Initialize 3D RoPE encoder

rope3d = WanRotaryPosEmbed(
    attention_head_dim=64,
    patch_size=(1, 14, 14),
    max_seq_len=1024,
    theta=10000.0,
    fhw_dim=[20, 22, 22]  # Split head dim across t/h/w axes

)

# Generate frequency embeddings for 30 frames with 14x14 patches

freqs = rope3d(
    ppf=30,
    pph=14, ppw=14,
    patch_start_idx=6,  # 6 special tokens per frame

    device=torch.device('cuda')
)

# Shape: (1, 1, 30 * (6 + 14*14), 32)

# Apply to token features before caching

x = torch.randn(1, 16, 262, 64)  # [B, H, N, D]

x_with_rope = apply_rotary_emb(x, freqs)

# Initialize paged KV cache manager

kv_mgr = FlashInferKVCacheManager(
    num_blocks=12,
    max_num_frames=200,
    tokens_per_frame=262,  # 256 patches + 6 specials

    num_heads=16,
    head_dim=64,
    dtype=torch.bfloat16,
    device=torch.device('cuda')
)

# Append encoded frame (splits automatically into patch/special streams)

kv_mgr.append_frame(block_idx=0, k=x_with_rope, v=x_with_rope)

# Evict old patch pages while keeping special tokens permanent

kv_mgr.evict_frames(block_idx=0, scale_frames=8, sliding_window=64)

# Compute attention with cached position-aware K/V

q = torch.randn(262, 16, 64, device='cuda', dtype=torch.bfloat16)
out = kv_mgr.compute_attention(block_idx=0, q=q)

```

## Summary

- **3D RoPE encoding** in `WanRotaryPosEmbed` generates frame-height-width positional embeddings by concatenating three 1D frequency tables and expanding them to full token sequences including special tokens.
- **Pre-caching application** occurs in the camera head and stream aggregator before tokens enter the KV cache, ensuring all stored values carry positional information.
- **Dual-stream architecture** separates recyclable patch pages from permanent special token pages, enabling long-context video processing with bounded memory.
- **FlashInfer integration** retrieves position-aware tensors directly from the paged cache without recomputation, supporting efficient attention on sequences exceeding 10,000 frames.

## Frequently Asked Questions

### How does 3D RoPE differ from standard 2D rotary embeddings in LingBot-Map?

Standard 2D RoPE processes height and width dimensions only, while the 3D implementation in `WanRotaryPosEmbed` adds temporal (frame) encoding through the `fhw_dim` parameter split across three axes. This allows the model to understand temporal relationships between video frames rather than treating each frame as an independent image.

### Why are special tokens stored in a separate append-only stream?

Special tokens contain critical global context including camera parameters and register embeddings that must remain accessible for all future frames. The special stream in `FlashInferKVCacheManager` never evicts these pages, whereas the patch stream recycles older visual features through a sliding window to manage memory consumption on long videos.

### Can the paged KV cache handle variable numbers of patches per frame?

Yes, the cache handles variable lengths through `compute_last_page_len`, which tracks valid tokens in partially-filled pages. While the configuration specifies `tokens_per_frame` for allocation purposes, the attention wrapper respects actual sequence lengths when reading from both patch and special streams.

### Where does the RoPE application occur relative to the KV cache write?

According to the source code in [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) and [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), RoPE is applied to query and key tensors immediately after linear projection but before calling `append_frame`. This ensures that cached keys and values in `FlashInferKVCacheManager` already incorporate their 3D positional encodings, eliminating the need for position recomputation during attention retrieval.