What Is LingBot-Map? A Feed-Forward 3D Foundation Model for Streaming 3D Reconstruction

LingBot-Map is a feed-forward 3D foundation model that performs real-time streaming 3D reconstruction from monocular video, solving the problem of long-range drift and high computational cost inherent in traditional SLAM and NeRF pipelines.

LingBot-Map is an open-source streaming 3D reconstruction system developed by Robbyant/lingbot-map that unifies geometric context understanding with high-efficiency inference. Unlike iterative SLAM or NeRF methods that require per-frame optimization, this model processes video sequences in a single forward pass while maintaining geometric consistency across thousands of frames.

Core Architecture: The Geometric Context Transformer

The foundation of LingBot-Map is the Geometric Context Transformer, implemented in lingbot_map/layers/vision_transformer.py. This architecture extends the Vision Transformer (ViT) design with custom blocks for geometric token handling, memory-efficient attention, and optional sky-masking capabilities.

The model employs three key mechanisms to correct long-range drift:

  • Anchor context – provides global reference points to stabilize absolute pose
  • Pose-reference window – maintains local geometric consistency across recent frames
  • Trajectory memory – accumulates path information to prevent error accumulation

Real-Time Streaming Performance

LingBot-Map achieves real-time inference through paged KV-cache attention using FlashInfer, maintaining approximately 20 FPS on 518×378 resolution images even for sequences exceeding 10,000 frames. This efficiency eliminates the per-frame optimization bottlenecks typical of neural rendering pipelines, enabling processing of long-video sequences without exponential memory growth.

Running LingBot-Map: Interactive and Batch Modes

The repository provides two primary execution modes through demo.py and demo_render/batch_demo.py, with visualization support via lingbot_map/vis/viser_wrapper.py.

Interactive Visualization

Launch the browser-based Viser viewer for real-time reconstruction:

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky

This starts an interactive 3D viewer at http://localhost:8080.

Streaming with Keyframe Intervals

Reduce memory usage while preserving drift correction by processing only select frames:

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/loop \
    --keyframe_interval 5 \
    --mask_sky

Windowed Inference for Extended Sequences

Enable sliding-window attention for videos exceeding 3000 frames:

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/long_video \
    --window_size 3000 \
    --mask_sky

Offline Batch Rendering

Process very long sequences (e.g., 25,000-frame indoor walkthroughs) without interactive latency using demo_render/batch_demo.py:

python demo_render/batch_demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/indoor_long \
    --out_dir renders/indoor_long \
    --window_size 4000 \
    --mask_sky

This outputs video or GLB files for archival or downstream processing.

Benchmark Results and Evaluation

According to the evaluation scripts in benchmark/run.py, LingBot-Map outperforms both streaming baselines and classic iterative SLAM/NeRF pipelines on standard datasets including KITTI, Oxford Spires, and indoor long-video demonstrations. The feed-forward architecture eliminates the accumulation of optimization errors that plague iterative methods while maintaining geometric fidelity comparable to offline batch reconstruction systems.

Summary

  • LingBot-Map is a feed-forward 3D foundation model specifically designed for streaming 3D reconstruction from monocular RGB video.
  • The Geometric Context Transformer in lingbot_map/layers/vision_transformer.py unifies coordinate grounding, dense geometric cues, and drift correction through anchor contexts and trajectory memory.
  • Paged KV-cache attention enables 20 FPS inference on 518×378 images for sequences longer than 10,000 frames.
  • Users can run interactive reconstructions via demo.py or offline batch processing via demo_render/batch_demo.py for very long sequences.
  • The system achieves state-of-the-art reconstruction quality on robotics benchmarks including KITTI and Oxford Spires.

Frequently Asked Questions

What is LingBot-Map used for?

LingBot-Map solves the problem of real-time 3D scene reconstruction from monocular video streams, making it suitable for robotics navigation, AR/VR applications, and large-scale virtual mapping. Unlike traditional SLAM systems that require iterative bundle adjustment, it processes video in a single forward pass while correcting for long-range drift through geometric context transformers.

How does LingBot-Map handle long video sequences?

The model implements sliding-window attention mechanisms and paged KV-cache (FlashInfer) to maintain constant memory usage regardless of sequence length. For videos exceeding 3000 frames, users can specify --window_size parameters in demo.py, while the batch_demo.py script handles sequences of 25,000+ frames through chunked processing and trajectory memory management.

What hardware is required to run LingBot-Map?

The implementation uses FlashInfer for optimized attention computation and achieves 20 FPS inference on 518×378 images according to the source benchmarks. This performance profile indicates a CUDA-capable GPU is required for real-time streaming, though the offline batch mode in demo_render/batch_demo.py can accommodate longer processing times on less powerful hardware.

How does LingBot-Map differ from traditional SLAM systems?

Traditional SLAM systems rely on iterative optimization that accumulates latency and drift over time. LingBot-Map replaces these iterative components with a feed-forward transformer that simultaneously grounds image coordinates, integrates depth cues, and corrects drift through anchor contexts and pose-reference windows, enabling true streaming reconstruction without per-frame optimization loops.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →