Architecture of LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction
LingBot-Map implements a Geometric Context Transformer (GCT) that combines a KV-cached Feature Aggregator with a causal Camera Head to enable real-time, memory-bounded streaming 3D reconstruction from video sequences.
LingBot-Map is an open-source framework for streaming 3D reconstruction built around a Geometric Context Transformer (GCT). The architecture fuses multi-scale visual features with 3-D positional encoding through a causal attention mechanism, enabling frame-by-frame processing without unbounded memory growth. At its core, the system uses a paged KV-cache strategy with configurable keyframe intervals to handle arbitrarily long video sequences on consumer GPUs.
Core Architectural Components
The architecture consists of three tightly coupled modules implemented in lingbot_map/models/gct_stream.py:
Feature Aggregator (AggregatorStream)
The AggregatorStream extracts multi-scale visual tokens from each input frame and manages temporal context through a paged KV-cache. This component supports 3-D RoPE (Rotary Positional Encoding) for temporal consistency across frames.
Key implementation details in _build_aggregator() include:
- Configuring
kv_cache_sliding_windowfor maximum cache size - Setting
kv_cache_scale_frames(typically 8) for initial scale estimation - Storing visual tokens using FlashInfer-backed paging when
use_sdpa=False
Camera Head (CameraCausalHead)
The CameraCausalHead refines per-frame camera poses using accumulated tokens from the aggregator. Implemented in _build_camera_head(), this module supports:
- Iterative refinement via
camera_num_iterations - Causal processing that only attends to previous frames
- Optional 3-D RoPE specific to the camera pose branch
Streaming Engine (GCTStream)
The GCTStream class orchestrates the aggregator and camera head, managing the KV-cache lifecycle through inference_streaming(). The engine provides:
- Skip-append logic for non-keyframes to limit cache growth
- Two-phase processing: bidirectional attention on scale frames, followed by causal streaming
- Automatic device management for offloading predictions to CPU
How the Streaming Pipeline Works
The architecture processes video through a defined lifecycle that balances reconstruction quality with memory constraints.
Scale Frame Initialization
The first N frames (default 8, configured via num_scale_frames) undergo bidirectional attention processing to establish scene scale. This "scale token" phase yields initial depth and pose estimates that anchor the coordinate system for subsequent causal processing.
KV-Cache Management and Keyframe Logic
After initialization, the system switches to causal processing with configurable KV-cache eviction:
keyframe_interval = 1: Every frame's keys and values are cached (high quality, high memory)keyframe_interval > 1: Only every k-th frame triggers_set_skip_append(False); non-keyframes use_set_skip_append(True)to discard KV tensors after processing
This mechanism keeps GPU memory bounded even for 10,000+ frame sequences.
Windowed Inference for Long Sequences
For sequences exceeding the training context window (320 frames), the inference_windowed() method (available in both gct_stream.py and gct_stream_window.py) implements sliding-window processing:
window_size: Maximum KV slots per window (e.g., 128)overlap_size: Shared keyframes between adjacent windows (e.g., 8)- Automatic cache reset between windows to prevent position encoding overflow
Key Implementation Files
| File Path | Architectural Role |
|---|---|
lingbot_map/models/gct_stream.py |
Core GCT implementation with GCTStream, AggregatorStream, and CameraCausalHead classes |
lingbot_map/models/gct_stream_window.py |
Windowed inference variant with sliding-window KV eviction |
benchmark/methods/lingbot_map.py |
Benchmark adapter providing LingbotMapMethod for evaluation pipelines |
demo_render/batch_demo.py |
End-to-end rendering pipeline for long video sequences |
Practical Usage Examples
Running the Interactive Streaming Demo
Execute the streaming mode with configurable keyframe intervals to balance speed and accuracy:
python demo.py \
--model_path /path/to/lingbot-map-long.pt \
--image_folder example/oxford \
--mask_sky \
--keyframe_interval 2
This command processes frames through LingbotMapMethod → GCTStream.inference_streaming(), caching only every second frame's KV tensors.
Programmatic Inference with GCTStream
Build and run the model directly for custom pipelines:
import torch
from lingbot_map.models.gct_stream import GCTStream
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri
# Initialize streaming model
model = GCTStream(
img_size=518,
patch_size=14,
enable_3d_rope=True,
kv_cache_sliding_window=64,
kv_cache_scale_frames=8,
use_sdpa=False, # Enables FlashInfer backend
)
# Load checkpoint
ckpt = torch.load("lingbot-map-long.pt", map_location="cpu")
model.load_state_dict(ckpt["model"], strict=False)
model.eval().cuda()
# Prepare RGB frames [S, 3, H, W]
preds = model.inference_streaming(
imgs,
num_scale_frames=8,
keyframe_interval=4,
output_device=torch.device("cpu")
)
# Decode to extrinsic/intrinsic matrices
extr, intr = pose_encoding_to_extri_intri(preds["pose_enc"], imgs.shape[-2:])
Windowed Processing for Long Videos
Handle sequences beyond the RoPE training range using windowed mode:
from lingbot_map.models.gct_stream_window import GCTStream
model = GCTStream(...)
# Load checkpoint...
preds = model.inference_windowed(
imgs,
window_size=128,
overlap_size=8,
keyframe_interval=13,
flow_threshold=0.0,
max_non_keyframe_gap=30,
)
This pattern automatically manages KV-cache resets between windows while maintaining temporal continuity through overlapping keyframes.
Benchmark Integration
Use the provided adapter for standardized evaluation:
from benchmark.methods.lingbot_map import LingbotMapMethod
method = LingbotMapMethod(
checkpoint="lingbot-map-long.pt",
device="cuda",
mode="streaming",
keyframe_interval="auto", # Auto-select based on sequence length
)
results = method.process_scene(gt_artifact)
The process_scene() method handles image preparation via _prepare_images(), inference via _run_inference(), and output formatting via _process_outputs().
Summary
- LingBot-Map implements a Geometric Context Transformer (GCT) combining a KV-cached Feature Aggregator with a causal Camera Head for streaming 3D reconstruction.
- The architecture uses 3-D RoPE and paged KV-caches with configurable
keyframe_intervalsettings to maintain bounded GPU memory during causal processing. - Two-phase processing handles scale initialization (bidirectional) followed by causal streaming, while windowed inference in
inference_windowed()supports arbitrary sequence lengths. - Core implementation resides in
lingbot_map/models/gct_stream.py, with benchmark adapters inbenchmark/methods/lingbot_map.pyand rendering tools indemo_render/batch_demo.py.
Frequently Asked Questions
What makes LingBot-Map suitable for real-time streaming applications?
The causal attention mechanism and KV-cache management enable frame-by-frame processing without revisiting previous frames. By configuring keyframe_interval and using _set_skip_append() for non-keyframes, the system maintains constant GPU memory regardless of video length. The CameraCausalHead processes each frame using only accumulated history from the AggregatorStream's cache, eliminating the quadratic memory growth typical of standard transformers.
How does the keyframe interval affect reconstruction quality and performance?
Setting keyframe_interval=1 stores every frame's KV tensors, maximizing temporal coherence but consuming significant VRAM. Increasing the interval (e.g., to 4 or 13) reduces memory usage linearly while the CameraCausalHead continues to output poses for every frame. The _set_skip_append(True) mechanism ensures non-keyframes compute forward passes without polluting the cache, making the trade-off between memory and drift customizable per hardware constraints.
What is the difference between streaming mode and windowed mode?
Streaming mode (inference_streaming()) processes sequences continuously with a monotonic KV-cache, suitable for online scenarios up to the RoPE training limit (320 frames). Windowed mode (inference_windowed()) divides long sequences into overlapping chunks (controlled by window_size and overlap_size), resetting the KV-cache between windows to handle 10,000+ frame videos without position encoding overflow. Windowed mode is implemented in gct_stream_window.py and used in demo_render/batch_demo.py for offline batch processing.
Where is the camera pose decoding implemented?
Raw pose encodings from the model are converted to standard extrinsic and intrinsic matrices via pose_encoding_to_extri_intri() in lingbot_map/utils/pose_enc.py. This function accepts the model's pose_enc output tensor and image dimensions, returning camera matrices compatible with standard computer vision benchmarks. The benchmark adapter in benchmark/methods/lingbot_map.py automatically handles this conversion in _process_outputs().
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →