LingBot-Map Architecture Explained: Core Components of the Geometric Context Transformer
LingBot-Map architecture centers on a Geometric Context Transformer (GCT) that fuses multi-scale visual features with 3-D positional encoding and causal attention to enable real-time, memory-efficient streaming 3D reconstruction.
The LingBot-Map architecture powers the Robbyant/lingbot-map repository's approach to dense 3D reconstruction from monocular video sequences. Built around a streaming transformer design, this architecture processes frames causally while maintaining temporal geometric consistency through a novel KV-caching mechanism and 3-D Rotary Positional Encoding (RoPE).
Core Components of the LingBot-Map Architecture
The LingBot-Map architecture comprises three tightly coupled modules that operate within the GCTStream class defined in lingbot_map/models/gct_stream.py.
Feature Aggregator (AggregatorStream)
The Feature Aggregator extracts multi-scale visual tokens from each input frame and manages the temporal memory through a paged KV-cache. Implemented via _build_aggregator() in gct_stream.py, this component returns an AggregatorStream instance configured with cache parameters including kv_cache_sliding_window and kv_cache_scale_frames. The aggregator optionally applies 3-D RoPE (Rotary Positional Encoding) to maintain temporal consistency across the video sequence.
Camera Head (CameraCausalHead)
The Camera Head refines per-frame camera poses using accumulated tokens from the aggregator. Created by _build_camera_head() in the same source file, the CameraCausalHead processes geometry in a strictly causal manner, supporting iterative refinement through the camera_num_iterations parameter. This module receives identical KV-cache settings to the aggregator and can be driven by keyframe logic to optimize computation.
Streaming Engine (GCTStream)
The Streaming Engine orchestrates the aggregator and camera head modules while managing the KV-cache lifecycle. The GCTStream class provides inference_streaming() for frame-by-frame processing and inference_windowed() for long sequences. The engine implements cache cleaning, "skip-append" logic for non-keyframes via _set_skip_append(True), and maintains statistics to prevent unbounded memory growth.
Architectural Pipeline and Data Flow
The LingBot-Map architecture processes video through a five-phase pipeline that balances geometric accuracy with computational efficiency.
-
Input preprocessing: Images are converted to tensor format
[S, 3, H, W]and optionally resized viaLingbotMapMethod._prepare_imagesinbenchmark/methods/lingbot_map.py. -
Scale frame initialization: The first N frames (default 8) are processed together using bidirectional attention through a dedicated scale token. This establishes the initial scene scale and provides baseline depth and pose estimates.
-
KV-cache streaming: Subsequent frames are processed causally. When
keyframe_interval=1, every frame's key-value pairs are cached. Withkeyframe_interval > 1, only every k-th frame is stored; non-keyframes trigger_set_skip_append(True)to discard their KV tensors immediately after the forward pass. -
Windowed eviction: For sequences exceeding memory limits, the architecture supports sliding-window cache eviction via
window_sizeandoverlap_sizeparameters. Theinference_windowed()method inbenchmark/methods/lingbot_map.pywraps this functionality and handles flow-based keyframe selection. -
Post-processing: Raw pose encodings (
pose_enc) are decoded into extrinsic and intrinsic matrices usingpose_encoding_to_extri_intri, while depth maps are reshaped and optional confidence maps are attached to the output.
Key Implementation Files
Understanding the LingBot-Map architecture requires familiarity with these critical source files:
-
lingbot_map/models/gct_stream.py: Contains the coreGCTStreamimplementation, including the aggregator builder, camera head builder, and inference loops. -
lingbot_map/models/gct_stream_window.py: Implements the windowed inference variant with sliding-window KV eviction for sequences exceeding the standard 320-frame RoPE training range. -
benchmark/methods/lingbot_map.py: Provides the benchmark adapter that handles checkpoint loading, image preparation, and output formatting for evaluation pipelines. -
demo_render/batch_demo.py: Demonstrates the end-to-end offline rendering pipeline for long video sequences.
Implementation Examples
Running Streaming Inference
The following bash command executes the interactive demo in streaming mode with keyframe-based memory management:
python demo.py \
--model_path /path/to/lingbot-map-long.pt \
--image_folder example/oxford \
--mask_sky \
--keyframe_interval 2
This internally instantiates LingbotMapMethod, loads the checkpoint, and calls inference_streaming() on the GCTStream model. The KV-cache grows only on keyframes, maintaining bounded GPU memory usage.
Programmatic Model Usage
For custom applications, instantiate the streaming engine directly:
import torch
from lingbot_map.models.gct_stream import GCTStream
# Initialize model with KV-cache parameters
model = GCTStream(
img_size=518,
patch_size=14,
enable_3d_rope=True,
kv_cache_sliding_window=64,
kv_cache_scale_frames=8,
use_sdpa=False, # FlashInfer backend
)
# Load checkpoint
ckpt = torch.load("lingbot-map-long.pt", map_location="cpu")
model.load_state_dict(ckpt["model"], strict=False)
model.eval().cuda()
# Prepare images and run inference
preds = model.inference_streaming(
imgs,
num_scale_frames=8,
keyframe_interval=4,
output_device=torch.device("cpu")
)
# Decode poses
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri
extr, intr = pose_encoding_to_extri_intri(preds["pose_enc"], imgs.shape[-2:])
Windowed Inference for Long Sequences
For videos exceeding 10,000 frames, use the windowed implementation:
from lingbot_map.models.gct_stream_window import GCTStream
model = GCTStream(
img_size=518,
patch_size=14,
enable_3d_rope=True,
kv_cache_sliding_window=64,
kv_cache_scale_frames=8,
use_sdpa=False,
)
# Process with sliding windows
preds = model.inference_windowed(
imgs,
window_size=128,
overlap_size=8,
keyframe_interval=13,
flow_threshold=0.0,
max_non_keyframe_gap=30,
)
This mode automatically resets the KV cache after each window, enabling stable inference beyond the standard training range.
Benchmark Integration
Integrate with evaluation frameworks using the provided adapter:
from benchmark.methods.lingbot_map import LingbotMapMethod
method = LingbotMapMethod(
checkpoint="lingbot-map-long.pt",
device="cuda",
mode="streaming",
keyframe_interval="auto",
)
results = method.process_scene(gt_artifact)
The process_scene() method internally orchestrates _prepare_images(), _run_inference(), and _process_outputs() to return benchmark-compatible results.
Offline Rendering Pipeline
For processing very long video sequences end-to-end:
python demo_render/batch_demo.py \
--video_path /data/demo_videos/indoor_travel.MP4 \
--output_folder /data/outputs/indoor_travel/ \
--model_path /path/to/lingbot-map.pt \
--config demo_render/config/indoor.yaml \
--mode windowed --window_size 128 \
--keyframe_interval 13 --overlap_keyframes 8 \
--mask_sky \
--save_predictions
This script calls GCTStream.inference_windowed() under the hood to generate MP4 fly-throughs with optional sky-mask visualizations.
Summary
The LingBot-Map architecture delivers efficient streaming 3D reconstruction through several key innovations:
- Geometric Context Transformer (GCT) design combining visual feature aggregation with causal camera pose estimation
- Paged KV-cache system with configurable
kv_cache_sliding_windowand keyframe intervals to bound memory usage - 3-D RoPE integration for maintaining temporal geometric consistency across video sequences
- Dual inference modes: streaming for real-time applications and windowed for arbitrarily long sequences
- Modular implementation separating concerns between
AggregatorStream,CameraCausalHead, and theGCTStreamorchestration layer
Frequently Asked Questions
What is the Geometric Context Transformer in LingBot-Map?
The Geometric Context Transformer (GCT) is the core neural architecture that fuses multi-scale visual features with 3-D positional information through causal attention mechanisms. It consists of the AggregatorStream for feature extraction and the CameraCausalHead for pose refinement, both implemented in lingbot_map/models/gct_stream.py.
How does LingBot-Map handle memory constraints during streaming?
The architecture implements a keyframe-based KV-cache that stores attention keys and values only at specified intervals. By setting keyframe_interval > 1 and invoking _set_skip_append(True) for non-keyframes, the model discards intermediate tensors immediately after processing, keeping GPU memory bounded regardless of sequence length.
What is the difference between streaming and windowed inference modes?
Streaming mode (inference_streaming()) processes frames sequentially with a monotonically growing KV-cache, suitable for real-time applications. Windowed mode (inference_windowed()) processes sequences in overlapping chunks with periodic cache resets, enabling processing of videos exceeding the 320-frame RoPE training limit while maintaining geometric consistency through overlap_size shared frames.
Which files contain the core LingBot-Map architecture implementation?
The primary implementation resides in lingbot_map/models/gct_stream.py, which defines the GCTStream class, _build_aggregator(), and _build_camera_head(). Windowed variants are in lingbot_map/models/gct_stream_window.py, while the benchmark wrapper is located at benchmark/methods/lingbot_map.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →