How LingBot-Map Achieves Real-Time 3D Reconstruction: Streaming Transformers and KV Caching
LingBot-Map achieves real-time 3D reconstruction by processing video frames through a streaming Vision Transformer that maintains a paged key-value (KV) cache and applies 3-D rotary position embeddings, enabling constant-memory causal attention across arbitrarily long sequences.
The Robbyant/lingbot-map repository implements a novel approach to real-time 3D reconstruction that eliminates the memory bottlenecks of traditional batch-processing methods. By treating video streams as temporal volumes rather than independent images, the system continuously updates depth maps and point clouds at video frame rates. This architecture relies on three core technical innovations: the streaming GCT-Stream transformer, FlashInfer-based KV caching, and 3-D rotary position embeddings.
Core Architecture Components
GCT-Stream: A Streaming-Aware Vision Transformer
The foundation of the system is the GCTStream class defined in lingbot_map/models/gct_stream.py. Unlike standard Vision Transformers that process entire sequences in parallel, this model inherits from GCTBase and overrides _build_aggregator to instantiate a streaming aggregator.
The class accepts critical streaming parameters including kv_cache_sliding_window, enable_3d_rope, and use_sdpa, configuring the model for frame-by-frame processing rather than batch inference. This design allows the model to ingest frames continuously without requiring the entire video to reside in memory simultaneously.
AggregatorStream with FlashInfer Paged KV Cache
At the heart of the memory-efficient processing lies the AggregatorStream class in lingbot_map/aggregator/stream.py. This component implements a paged KV cache using FlashInfer, allowing the model to maintain attention states from previous frames without storing the entire video history in GPU memory.
The implementation includes several key mechanisms:
_init_kv_cache: Initializes the FlashInfer manager (or a dictionary-based fallback) to handle cache read/write operations._build_blocks: Constructs frame blocks and global blocks, utilizingFlashInferBlock(orSDPABlockwhenuse_sdpa=True) where the paged KV cache resides.- Sliding-window eviction: The
kv_cache_sliding_windowparameter bounds memory usage by evicting older frames when the cache exceeds the specified limit, whilekv_cache_scale_framesandkv_cache_cross_frame_specialcontrol how scale and special tokens interact with the temporal cache.
3-D Rotary Position Embeddings for Temporal Consistency
To maintain geometric consistency across time, LingBot-Map extends standard 2-D RoPE into three dimensions. The WanRotaryPosEmbed class in lingbot_map/layers/rope.py computes frequency components for the temporal dimension alongside spatial coordinates.
When the enable_3d_rope flag is activated (passed to both AggregatorStream and CameraCausalHead), the model applies rotation based on absolute frame indices, ensuring that positional encodings respect the temporal structure of the 3-D volume being reconstructed.
The Real-Time Inference Pipeline
The system executes real-time 3D reconstruction through a carefully orchestrated sequence of operations that minimize latency while preserving geometric accuracy.
Model Initialization and KV Cache Warm-Up
Inference begins when demo.py or a custom script instantiates GCTStream with streaming-specific arguments. The constructor chains through __init__ → super().__init__ → _build_aggregator to create the AggregatorStream with the configured KV cache settings.
Before processing the first frame, the aggregator lazily initializes the FlashInfer manager via _get_flashinfer_manager, and the initial forward pass populates the cache with baseline key/value tensors.
Per-Frame Processing and Causal Attention
For each new video frame, the system performs the following operations:
- Feature aggregation: The current frame tensor (shape
[B, 1, 3, H, W]) passes throughmodel._aggregate_features(), which invokesAggregatorStream.__call__. - Cached attention retrieval: The aggregator retrieves KV tensors from previous frames through the paged cache.
- Temporal causal attention: The current frame's tokens attend only to cached past tokens and special tokens (camera or scale tokens), preventing future information leakage.
- 3-D RoPE application: If enabled, the system uses the frame index to rotate query/key embeddings via the temporal RoPE implementation, maintaining spatial-temporal coherence.
Cache Eviction and Output Generation
After processing each frame, AggregatorStream increments total_frames_processed and evaluates the sliding-window policy. If the number of stored frames exceeds kv_cache_sliding_window, the oldest blocks are evicted from the FlashInfer cache, maintaining constant memory consumption regardless of video length.
Simultaneously, the CameraCausalHead (also using the KV cache and optional 3-D RoPE) refines camera poses frame-by-frame. Finally, the point_head decodes aggregated tokens into depth maps and point clouds, visualized in real time via lingbot_map/vis/point_cloud_viewer.py.
Code Implementation Examples
Python Streaming Inference
The following implementation demonstrates how to configure and run the streaming model for real-time 3D reconstruction:
# ----------------------------------------------------------------------
# Streaming inference example (equivalent to the demo script)
# ----------------------------------------------------------------------
import torch
from lingbot_map.models.gct_stream import GCTStream
from lingbot_map.utils.load_fn import load_and_preprocess_images
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# ---- Load a checkpoint -------------------------------------------------
model = GCTStream(
img_size=518,
patch_size=14,
enable_3d_rope=True, # activate temporal RoPE
max_frame_num=1024,
kv_cache_sliding_window=64, # keep 64 recent frames in KV cache
kv_cache_scale_frames=8,
kv_cache_cross_frame_special=True,
kv_cache_include_scale_frames=True,
use_sdpa=False, # use FlashInfer (default)
camera_num_iterations=4,
).to(device).eval()
# ---- Prepare a short video sequence ------------------------------------
image_paths = ["frame000.jpg", "frame001.jpg", "frame002.jpg"]
images = load_and_preprocess_images(
image_paths,
mode="crop",
image_size=518,
patch_size=14,
).to(device) # shape: [B, S, 3, H, W]
# ---- Run streaming inference -------------------------------------------
# The model will automatically manage the KV cache across frames.
with torch.no_grad():
for i in range(images.shape[1]): # iterate over time dimension
cur_img = images[:, i:i+1] # shape: [B, 1, 3, H, W]
# Aggregator returns tokens; camera head returns pose refinements, etc.
tokens, _ = model._aggregate_features(cur_img)
depth, point_cloud = model.point_head(tokens) # example decode step
# … visualize or save depth / point cloud here …
Command-Line Interface
For direct execution from the repository root:
# ----------------------------------------------------------------------
# Command‑line demo (run from the repository root)
# ----------------------------------------------------------------------
python demo.py \
--model_path checkpoints/gct_stream.pt \
--image_folder ./my_video_frames \
--mode streaming \
--enable_3d_rope \
--kv_cache_sliding_window 64
The demo script handles KV cache warm-up, per-frame forwarding, and real-time visualization automatically.
Summary
- Streaming Architecture: The
GCTStreamclass inlingbot_map/models/gct_stream.pyenables frame-by-frame processing by overriding the standard aggregator with a streaming implementation. - Memory Efficiency:
AggregatorStreamutilizes FlashInfer's paged KV cache with configurable sliding-window eviction (kv_cache_sliding_window) to maintain constant GPU memory usage during arbitrarily long video sequences. - Temporal Coherence: 3-D rotary position embeddings (
WanRotaryPosEmbedinlingbot_map/layers/rope.py) encode absolute frame indices into the attention mechanism, preserving geometric consistency across the reconstructed volume. - Real-Time Performance: By combining causal attention, cached state management, and optimized pose refinement through
CameraCausalHead, the system achieves video-rate reconstruction (approximately 10 fps) while continuously updating 3-D geometry.
Frequently Asked Questions
What is the KV cache in LingBot-Map and why is it necessary for real-time 3D reconstruction?
The KV cache is a memory buffer that stores key and value tensors from previous video frames, allowing the transformer to attend to past information without recomputing attention for the entire sequence. According to the source code in lingbot_map/aggregator/stream.py, the paged KV cache implementation using FlashInfer enables the model to process streaming video with constant memory usage rather than linear growth, which is essential for real-time operation on limited GPU resources.
How does 3-D RoPE differ from standard 2-D positional encoding in Vision Transformers?
Standard 2-D RoPE encodes spatial positions within individual images, while 3-D RoPE extends this to include the temporal dimension. As implemented in lingbot_map/layers/rope.py, the WanRotaryPosEmbed class computes rotation frequencies for frame indices alongside spatial coordinates, allowing the model to treat the video as a 3-D volume and maintain geometric consistency across time steps during reconstruction.
Can LingBot-Map process arbitrarily long videos without running out of memory?
Yes, the sliding-window KV cache mechanism ensures bounded memory consumption regardless of video length. The kv_cache_sliding_window parameter in AggregatorStream configures how many recent frames remain in the FlashInfer cache, automatically evicting older frames while preserving the temporal context necessary for coherent 3-D reconstruction.
What is the difference between GCT-Stream and the windowed variant?
The standard GCTStream class processes video as a continuous stream using a single KV cache, while the windowed variant in lingbot_map/models/gct_stream_window.py splits very long sequences into overlapping windows to manage extreme-length videos. Both implementations inherit from GCTBase and use the same underlying AggregatorStream mechanism, but the windowed version provides additional segmentation for sequences exceeding the memory capacity of even the sliding-window cache.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →