Architecture of LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction

LingBot-Map implements a Geometric Context Transformer (GCT) that combines a KV-cached Feature Aggregator with a causal Camera Head to enable real-time, memory-bounded streaming 3D reconstruction from video sequences.

LingBot-Map is an open-source framework for streaming 3D reconstruction built around a Geometric Context Transformer (GCT). The architecture fuses multi-scale visual features with 3-D positional encoding through a causal attention mechanism, enabling frame-by-frame processing without unbounded memory growth. At its core, the system uses a paged KV-cache strategy with configurable keyframe intervals to handle arbitrarily long video sequences on consumer GPUs.

Core Architectural Components

The architecture consists of three tightly coupled modules implemented in lingbot_map/models/gct_stream.py:

Feature Aggregator (AggregatorStream)

The AggregatorStream extracts multi-scale visual tokens from each input frame and manages temporal context through a paged KV-cache. This component supports 3-D RoPE (Rotary Positional Encoding) for temporal consistency across frames.

Key implementation details in _build_aggregator() include:

  • Configuring kv_cache_sliding_window for maximum cache size
  • Setting kv_cache_scale_frames (typically 8) for initial scale estimation
  • Storing visual tokens using FlashInfer-backed paging when use_sdpa=False

Camera Head (CameraCausalHead)

The CameraCausalHead refines per-frame camera poses using accumulated tokens from the aggregator. Implemented in _build_camera_head(), this module supports:

  • Iterative refinement via camera_num_iterations
  • Causal processing that only attends to previous frames
  • Optional 3-D RoPE specific to the camera pose branch

Streaming Engine (GCTStream)

The GCTStream class orchestrates the aggregator and camera head, managing the KV-cache lifecycle through inference_streaming(). The engine provides:

  • Skip-append logic for non-keyframes to limit cache growth
  • Two-phase processing: bidirectional attention on scale frames, followed by causal streaming
  • Automatic device management for offloading predictions to CPU

How the Streaming Pipeline Works

The architecture processes video through a defined lifecycle that balances reconstruction quality with memory constraints.

Scale Frame Initialization

The first N frames (default 8, configured via num_scale_frames) undergo bidirectional attention processing to establish scene scale. This "scale token" phase yields initial depth and pose estimates that anchor the coordinate system for subsequent causal processing.

KV-Cache Management and Keyframe Logic

After initialization, the system switches to causal processing with configurable KV-cache eviction:

  • keyframe_interval = 1: Every frame's keys and values are cached (high quality, high memory)
  • keyframe_interval > 1: Only every k-th frame triggers _set_skip_append(False); non-keyframes use _set_skip_append(True) to discard KV tensors after processing

This mechanism keeps GPU memory bounded even for 10,000+ frame sequences.

Windowed Inference for Long Sequences

For sequences exceeding the training context window (320 frames), the inference_windowed() method (available in both gct_stream.py and gct_stream_window.py) implements sliding-window processing:

  • window_size: Maximum KV slots per window (e.g., 128)
  • overlap_size: Shared keyframes between adjacent windows (e.g., 8)
  • Automatic cache reset between windows to prevent position encoding overflow

Key Implementation Files

File Path Architectural Role
lingbot_map/models/gct_stream.py Core GCT implementation with GCTStream, AggregatorStream, and CameraCausalHead classes
lingbot_map/models/gct_stream_window.py Windowed inference variant with sliding-window KV eviction
benchmark/methods/lingbot_map.py Benchmark adapter providing LingbotMapMethod for evaluation pipelines
demo_render/batch_demo.py End-to-end rendering pipeline for long video sequences

Practical Usage Examples

Running the Interactive Streaming Demo

Execute the streaming mode with configurable keyframe intervals to balance speed and accuracy:

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/oxford \
    --mask_sky \
    --keyframe_interval 2

This command processes frames through LingbotMapMethod → GCTStream.inference_streaming(), caching only every second frame's KV tensors.

Programmatic Inference with GCTStream

Build and run the model directly for custom pipelines:

import torch
from lingbot_map.models.gct_stream import GCTStream
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri

# Initialize streaming model

model = GCTStream(
    img_size=518,
    patch_size=14,
    enable_3d_rope=True,
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    use_sdpa=False,  # Enables FlashInfer backend

)

# Load checkpoint

ckpt = torch.load("lingbot-map-long.pt", map_location="cpu")
model.load_state_dict(ckpt["model"], strict=False)
model.eval().cuda()

# Prepare RGB frames [S, 3, H, W]

preds = model.inference_streaming(
    imgs,
    num_scale_frames=8,
    keyframe_interval=4,
    output_device=torch.device("cpu")
)

# Decode to extrinsic/intrinsic matrices

extr, intr = pose_encoding_to_extri_intri(preds["pose_enc"], imgs.shape[-2:])

Windowed Processing for Long Videos

Handle sequences beyond the RoPE training range using windowed mode:

from lingbot_map.models.gct_stream_window import GCTStream

model = GCTStream(...)

# Load checkpoint...

preds = model.inference_windowed(
    imgs,
    window_size=128,
    overlap_size=8,
    keyframe_interval=13,
    flow_threshold=0.0,
    max_non_keyframe_gap=30,
)

This pattern automatically manages KV-cache resets between windows while maintaining temporal continuity through overlapping keyframes.

Benchmark Integration

Use the provided adapter for standardized evaluation:

from benchmark.methods.lingbot_map import LingbotMapMethod

method = LingbotMapMethod(
    checkpoint="lingbot-map-long.pt",
    device="cuda",
    mode="streaming",
    keyframe_interval="auto",  # Auto-select based on sequence length

)

results = method.process_scene(gt_artifact)

The process_scene() method handles image preparation via _prepare_images(), inference via _run_inference(), and output formatting via _process_outputs().

Summary

  • LingBot-Map implements a Geometric Context Transformer (GCT) combining a KV-cached Feature Aggregator with a causal Camera Head for streaming 3D reconstruction.
  • The architecture uses 3-D RoPE and paged KV-caches with configurable keyframe_interval settings to maintain bounded GPU memory during causal processing.
  • Two-phase processing handles scale initialization (bidirectional) followed by causal streaming, while windowed inference in inference_windowed() supports arbitrary sequence lengths.
  • Core implementation resides in lingbot_map/models/gct_stream.py, with benchmark adapters in benchmark/methods/lingbot_map.py and rendering tools in demo_render/batch_demo.py.

Frequently Asked Questions

What makes LingBot-Map suitable for real-time streaming applications?

The causal attention mechanism and KV-cache management enable frame-by-frame processing without revisiting previous frames. By configuring keyframe_interval and using _set_skip_append() for non-keyframes, the system maintains constant GPU memory regardless of video length. The CameraCausalHead processes each frame using only accumulated history from the AggregatorStream's cache, eliminating the quadratic memory growth typical of standard transformers.

How does the keyframe interval affect reconstruction quality and performance?

Setting keyframe_interval=1 stores every frame's KV tensors, maximizing temporal coherence but consuming significant VRAM. Increasing the interval (e.g., to 4 or 13) reduces memory usage linearly while the CameraCausalHead continues to output poses for every frame. The _set_skip_append(True) mechanism ensures non-keyframes compute forward passes without polluting the cache, making the trade-off between memory and drift customizable per hardware constraints.

What is the difference between streaming mode and windowed mode?

Streaming mode (inference_streaming()) processes sequences continuously with a monotonic KV-cache, suitable for online scenarios up to the RoPE training limit (320 frames). Windowed mode (inference_windowed()) divides long sequences into overlapping chunks (controlled by window_size and overlap_size), resetting the KV-cache between windows to handle 10,000+ frame videos without position encoding overflow. Windowed mode is implemented in gct_stream_window.py and used in demo_render/batch_demo.py for offline batch processing.

Where is the camera pose decoding implemented?

Raw pose encodings from the model are converted to standard extrinsic and intrinsic matrices via pose_encoding_to_extri_intri() in lingbot_map/utils/pose_enc.py. This function accepts the model's pose_enc output tensor and image dimensions, returning camera matrices compatible with standard computer vision benchmarks. The benchmark adapter in benchmark/methods/lingbot_map.py automatically handles this conversion in _process_outputs().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →