Core Components of the LingBot-Map Architecture: A Deep Dive into the Geometric Context Transformer
The LingBot-Map architecture centers on a Geometric Context Transformer (GCT) that combines a Feature Aggregator with KV-caching, a causal Camera Head, and a Streaming Engine to enable real-time 3D reconstruction from long video sequences.
LingBot-Map is an open-source streaming 3D reconstruction system that processes visual sequences through a transformer-based causal attention mechanism. The architecture integrates 3-D Rotary Positional Encoding (RoPE) with a paged KV-cache to handle arbitrarily long videos without linear memory growth. All core modules reside in the Robbyant/lingbot-map repository, specifically within the lingbot_map/models/ directory.
The Three Pillars of the LingBot-Map Architecture
The system decomposes into three tightly coupled components that operate within the GCTStream class defined in lingbot_map/models/gct_stream.py.
Feature Aggregator (AggregatorStream)
The Feature Aggregator extracts multi-scale visual tokens from each input frame and manages temporal consistency through a sophisticated caching mechanism. Implemented via _build_aggregator() in lingbot_map/models/gct_stream.py, this component returns an AggregatorStream object configured with KV-cache parameters including kv_cache_sliding_window and kv_cache_scale_frames.
The aggregator optionally applies 3-D RoPE (Rotary Positional Encoding) to maintain geometric consistency across frames. It stores visual tokens in a paged KV-cache, which enables efficient memory management during streaming inference.
Camera Head (CameraCausalHead)
The Camera Head refines per-frame camera poses using a causal attention mechanism that only attends to previous frames. Created by _build_camera_head() in the same file, the CameraCausalHead receives the same KV-cache settings as the aggregator and supports iterative refinement through the camera_num_iterations parameter.
This component can apply 3-D RoPE specific to the camera branch and operates on "keyframe" logic to limit computational overhead. The causal design ensures that pose estimation for frame $t$ depends only on frames $0$ through $t-1$, maintaining temporal coherence without future context.
Streaming Engine (GCTStream)
The Streaming Engine orchestrates the aggregator and camera head through the GCTStream class. This engine manages the forward pass frame-by-frame, handles KV-cache lifecycle operations, and provides both streaming and windowed inference modes.
Key methods include inference_streaming(), which implements a two-phase pipeline processing scale frames bidirectionally followed by per-frame causal processing. The engine also manages the keyframe-interval mechanism that limits cache growth by invoking _set_skip_append(True) for non-keyframe frames.
The Streaming Inference Pipeline
Understanding how these components interact requires examining the five-stage pipeline implemented in the LingBot-Map architecture.
Input Preprocessing and Scale Frames
Images enter the system as tensors with shape [S, 3, H, W] through LingbotMapMethod._prepare_images in benchmark/methods/lingbot_map.py. The first N frames (default 8) serve as scale frames, processed together using bidirectional attention via a scale token to establish initial scene scale and depth estimates.
KV-Cache Streaming with Keyframe Selection
After processing scale frames, the system switches to causal streaming mode. When keyframe_interval=1, every frame's key-value tensors are cached. With keyframe_interval > 1, only every k-th frame is stored; non-keyframes trigger _set_skip_append(True), discarding their KV tensors immediately after the forward pass to bound GPU memory usage.
Windowed Mode for Long Sequences
For sequences exceeding the 320-frame RoPE training range, the architecture supports sliding-window eviction through inference_windowed() in benchmark/methods/lingbot_map.py. This mode uses window_size and overlap_size parameters to reset the KV cache periodically while maintaining continuity between windows.
Post-Processing and Output Decoding
Raw pose encodings (pose_enc) are decoded into extrinsic and intrinsic matrices using pose_encoding_to_extri_intri(). Depth maps are reshaped to match input dimensions, and optional confidence maps are attached to the output dictionary.
Key Implementation Files
The LingBot-Map architecture spans several critical files in the repository:
lingbot_map/models/gct_stream.py: Contains the coreGCTStreamclass,_build_aggregator(),_build_camera_head(), and theinference_streaming()method.lingbot_map/models/gct_stream_window.py: Implements the windowed inference variant with sliding-window KV eviction.benchmark/methods/lingbot_map.py: Provides the benchmark adapter that handles checkpoint loading, image preparation, and result formatting.demo_render/batch_demo.py: Showcases the streaming engine on long video sequences with offline rendering capabilities.
Practical Usage Examples
Running Interactive Streaming Inference
Execute the streaming mode with controlled KV-cache growth using the command-line interface:
python demo.py \
--model_path /path/to/lingbot-map-long.pt \
--image_folder example/oxford \
--mask_sky \
--keyframe_interval 2
Internally, this invokes LingbotMapMethod which loads the checkpoint and calls inference_streaming() on the GCTStream model.
Programmatic Model Initialization
Instantiate the streaming architecture directly in Python:
import torch
from lingbot_map.models.gct_stream import GCTStream
model = GCTStream(
img_size=518,
patch_size=14,
enable_3d_rope=True,
kv_cache_sliding_window=64,
kv_cache_scale_frames=8,
use_sdpa=False, # Uses FlashInfer backend for fast KV paging
)
ckpt = torch.load("lingbot-map-long.pt", map_location="cpu")
model.load_state_dict(ckpt["model"], strict=False)
model.eval().cuda()
Streaming Inference with Keyframe Control
Process video frames with explicit keyframe intervals to manage memory:
rgb_list = [...] # List of HxWx3 uint8 frames
imgs = torch.stack([torchvision.transforms.ToTensor()(im) for im in rgb_list]).cuda()
preds = model.inference_streaming(
imgs,
num_scale_frames=8,
keyframe_interval=4,
output_device=torch.device("cpu")
)
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri
extr, intr = pose_encoding_to_extri_intri(preds["pose_enc"], imgs.shape[-2:])
Windowed Inference for Extended Videos
Handle sequences exceeding 10,000 frames using sliding windows:
from lingbot_map.models.gct_stream_window import GCTStream
model = GCTStream(
img_size=518,
patch_size=14,
enable_3d_rope=True,
kv_cache_sliding_window=64,
use_sdpa=False,
)
preds = model.inference_windowed(
imgs,
window_size=128,
overlap_size=8,
keyframe_interval=13,
flow_threshold=0.0,
max_non_keyframe_gap=30,
)
Benchmark Adapter Integration
Use the standardized interface for evaluation pipelines:
from benchmark.methods.lingbot_map import LingbotMapMethod
method = LingbotMapMethod(
checkpoint="lingbot-map-long.pt",
device="cuda",
mode="streaming",
keyframe_interval="auto",
)
results = method.process_scene(gt_artifact)
The process_scene() method internally calls _prepare_images(), _run_inference(), and _process_outputs() to return benchmark-compatible results.
Summary
The LingBot-Map architecture delivers streaming 3D reconstruction through three integrated components:
- Feature Aggregator (AggregatorStream): Extracts multi-scale visual tokens and manages the paged KV-cache with configurable sliding windows.
- Camera Head (CameraCausalHead): Performs causal pose refinement with iterative optimization and 3-D RoPE support.
- Streaming Engine (GCTStream): Orchestrates frame-by-frame processing with keyframe-based memory management and dual inference modes.
The system achieves unbounded sequence processing via _set_skip_append() logic for non-keyframes and sliding-window eviction in inference_windowed(), keeping GPU memory constant regardless of video length.
Frequently Asked Questions
What is the Geometric Context Transformer in LingBot-Map?
The Geometric Context Transformer (GCT) is the foundational neural architecture that fuses visual features, 3-D positional encoding, and causal attention mechanisms. It enables the system to perform streaming 3D reconstruction by processing frames sequentially while maintaining geometric consistency through 3-D RoPE and bounded memory usage through KV-cache management.
How does the KV-cache manage memory in LingBot-Map?
The KV-cache implements a paged storage system with two primary constraints: kv_cache_sliding_window defines the maximum cache size, while keyframe_interval controls how frequently frames are stored. When processing non-keyframes, the system invokes _set_skip_append(True) to discard KV tensors immediately after the forward pass, preventing linear memory growth during long sequences.
What is the difference between streaming and windowed inference modes?
Streaming mode (inference_streaming()) processes videos continuously with a monotonically growing KV-cache that respects keyframe_interval constraints. Windowed mode (inference_windowed()) periodically resets the cache using window_size and overlap_size parameters, enabling processing of sequences that exceed the 320-frame RoPE training limit by treating the video as overlapping chunks.
Which files contain the core LingBot-Map implementation?
The primary implementation resides in lingbot_map/models/gct_stream.py, which contains the GCTStream class, _build_aggregator(), and _build_camera_head() methods. Windowed inference variants are in lingbot_map/models/gct_stream_window.py, while the benchmark integration layer is implemented in benchmark/methods/lingbot_map.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →