How 3D RoPE Positional Encoding Works with the Paged KV Cache in LingBot-Map
3D RoPE positional encoding pre-computes frame-height-width rotary embeddings that are baked into token features before entering the paged KV cache, enabling FlashInfer to retrieve position-aware keys and values without positional recomputation during streaming video inference.
LingBot-Map implements a streaming 3D transformer architecture that processes video frames sequentially using a novel integration of 3D Rotary Position Embeddings (RoPE) and a dual-stream paged KV cache. This design, found in the Robbyant/lingbot-map repository, enables efficient inference on extremely long video sequences by computing positional encodings upfront and storing position-aware tensors directly in reusable memory pages.
3D Rotary Position Embedding Architecture
The 3D RoPE encoder in lingbot_map/layers/rope.py generates rotary position embeddings across three spatial axes: frame (t), height (h), and width (w). This allows the model to understand temporal relationships in video streams rather than treating each frame as an independent image.
Multi-Axis Frequency Generation
The WanRotaryPosEmbed class creates independent 1D RoPE frequency tables for each axis using get_1d_rotary_pos_embed, then concatenates these three tables to form a unified 3D positional representation. The resulting complex-valued tensor has shape:
[1, 1, ppf * (patch_start_idx + pph * ppw), head_dim//2]
Where ppf represents patches-per-frame, pph and ppw denote the patch grid dimensions, and patch_start_idx accounts for special tokens preceding the patch sequence.
Complex-Valued Tensor Application
The forward method returns a complex-valued tensor that is applied to real-valued token features via apply_rotary_emb. This function performs complex multiplication using pure real arithmetic, ensuring compatibility with Torch-Compile and CUDA graphs while maintaining mathematical equivalence to standard complex operations.
Paged KV Cache Design
The paged KV cache in lingbot_map/layers/flashinfer_cache.py implements a two-stream memory layout that separates patch tokens from special tokens to optimize memory reuse and context preservation across thousands of frames.
Dual-Stream Memory Layout
The FlashInferKVCacheManager maintains two distinct streams:
- Patch stream: Stores
patches_per_frametokens per page using a recyclable sliding window. The firstscale_framespages remain permanent, while subsequent pages form asliding_windowthat evicts oldest frames back to a free list when full. - Special stream: An append-only stream holding six per-frame special tokens (camera, register, and scale tokens). These pages are never evicted, ensuring global context remains available for cross-frame attention.
Frame Appending and Position Encoding Integration
When append_frame receives new frame data, it splits the incoming tensor into special components (sp_k, sp_v) and patch components (patch_k, patch_v). Because the 3D RoPE encoding is applied to token features before this write operation (as implemented in lingbot_map/heads/camera_head.py and lingbot_map/aggregator/stream.py), the cached values already carry their positional information.
The _write_patch_page method stores patch tokens in the recyclable stream, while _write_special_tokens appends special tokens to their dedicated stream. This separation allows the system to evict visual patch history while preserving critical metadata tokens.
Attention Computation with Cached RoPE
During compute_attention, FlashInfer's BatchPrefillWithPagedKVCacheWrapper receives a visible page table constructed as:
scale_patch_pages → live_window_patch_pages → all_special_pages
The wrapper respects compute_last_page_len to handle partially-filled special pages correctly. Since RoPE is already encoded in the cached K/V tensors, the attention mechanism retrieves position-aware values directly without recomputing embeddings, regardless of how many times the patch pages are recycled.
Implementation Example
The following code demonstrates the complete workflow from 3D RoPE generation to paged cache utilization:
from lingbot_map.layers.rope import WanRotaryPosEmbed, apply_rotary_emb
from lingbot_map.layers.flashinfer_cache import FlashInferKVCacheManager
import torch
# Initialize 3D RoPE encoder
rope3d = WanRotaryPosEmbed(
attention_head_dim=64,
patch_size=(1, 14, 14),
max_seq_len=1024,
theta=10000.0,
fhw_dim=[20, 22, 22] # Split head dim across t/h/w axes
)
# Generate frequency embeddings for 30 frames with 14x14 patches
freqs = rope3d(
ppf=30,
pph=14, ppw=14,
patch_start_idx=6, # 6 special tokens per frame
device=torch.device('cuda')
)
# Shape: (1, 1, 30 * (6 + 14*14), 32)
# Apply to token features before caching
x = torch.randn(1, 16, 262, 64) # [B, H, N, D]
x_with_rope = apply_rotary_emb(x, freqs)
# Initialize paged KV cache manager
kv_mgr = FlashInferKVCacheManager(
num_blocks=12,
max_num_frames=200,
tokens_per_frame=262, # 256 patches + 6 specials
num_heads=16,
head_dim=64,
dtype=torch.bfloat16,
device=torch.device('cuda')
)
# Append encoded frame (splits automatically into patch/special streams)
kv_mgr.append_frame(block_idx=0, k=x_with_rope, v=x_with_rope)
# Evict old patch pages while keeping special tokens permanent
kv_mgr.evict_frames(block_idx=0, scale_frames=8, sliding_window=64)
# Compute attention with cached position-aware K/V
q = torch.randn(262, 16, 64, device='cuda', dtype=torch.bfloat16)
out = kv_mgr.compute_attention(block_idx=0, q=q)
Summary
- 3D RoPE encoding in
WanRotaryPosEmbedgenerates frame-height-width positional embeddings by concatenating three 1D frequency tables and expanding them to full token sequences including special tokens. - Pre-caching application occurs in the camera head and stream aggregator before tokens enter the KV cache, ensuring all stored values carry positional information.
- Dual-stream architecture separates recyclable patch pages from permanent special token pages, enabling long-context video processing with bounded memory.
- FlashInfer integration retrieves position-aware tensors directly from the paged cache without recomputation, supporting efficient attention on sequences exceeding 10,000 frames.
Frequently Asked Questions
How does 3D RoPE differ from standard 2D rotary embeddings in LingBot-Map?
Standard 2D RoPE processes height and width dimensions only, while the 3D implementation in WanRotaryPosEmbed adds temporal (frame) encoding through the fhw_dim parameter split across three axes. This allows the model to understand temporal relationships between video frames rather than treating each frame as an independent image.
Why are special tokens stored in a separate append-only stream?
Special tokens contain critical global context including camera parameters and register embeddings that must remain accessible for all future frames. The special stream in FlashInferKVCacheManager never evicts these pages, whereas the patch stream recycles older visual features through a sliding window to manage memory consumption on long videos.
Can the paged KV cache handle variable numbers of patches per frame?
Yes, the cache handles variable lengths through compute_last_page_len, which tracks valid tokens in partially-filled pages. While the configuration specifies tokens_per_frame for allocation purposes, the attention wrapper respects actual sequence lengths when reading from both patch and special streams.
Where does the RoPE application occur relative to the KV cache write?
According to the source code in lingbot_map/heads/camera_head.py and lingbot_map/aggregator/stream.py, RoPE is applied to query and key tensors immediately after linear projection but before calling append_frame. This ensures that cached keys and values in FlashInferKVCacheManager already incorporate their 3D positional encodings, eliminating the need for position recomputation during attention retrieval.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →