# How LingBot-Map Performs 3D Reconstruction from Video Streams: Architecture and Implementation

> Discover how LingBot-Map achieves efficient 3D reconstruction from video streams with its geometric-context transformer and paged KV-cache for constant GPU memory and real-time inference. Learn more about this innovative archit...

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: architecture
- Published: 2026-07-23

---

**LingBot-Map performs 3D reconstruction from video streams using a feed-forward geometric-context transformer with a paged KV-cache that stores only keyframe embeddings, enabling constant GPU memory usage and near real-time inference regardless of video length.**

LingBot-Map is an open-source framework developed by Robbyant that enables incremental 3D reconstruction from video streams through a novel streaming transformer architecture. Unlike traditional structure-from-motion pipelines that require batch processing or expensive global bundle adjustment, this system processes frames sequentially while maintaining bounded memory consumption. The implementation leverages FlashInfer's paged attention mechanisms to cache geometric context efficiently, making it suitable for arbitrarily long video sequences.

## Core Streaming Architecture: The GCTStream Model

The heart of the system is the **GCTStream** model, implemented in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py). This **feed-forward geometric-context transformer** processes each incoming frame independently while re-using cached key/value pairs from previous frames.

Unlike recurrent neural networks that maintain hidden states, `GCTStream` operates with no persistent state beyond the KV-cache. Each frame's forward pass attends to the cached context of previous keyframes, allowing the model to build a coherent 3D scene representation without gradient propagation through time. This architectural choice enables **O(1)** computational complexity per frame with respect to sequence length, once the cache is populated.

## Token Aggregation and Scene Anchoring

Before frames enter the transformer, the **AggregatorStream** class in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) constructs the input token sequence. This component performs three critical functions:

- **Patch Embedding**: Converts incoming RGB-D or RGB frames into visual tokens.
- **Anchor Token Insertion**: Prepends a global **anchor token** that provides persistent scene context across the entire video.
- **Pose-Reference Routing**: Injects **pose-reference tokens** that supply geometric constraints to the camera head.

The aggregator maintains the interface between raw video frames and the transformer blocks, ensuring that spatial relationships are properly encoded before cache storage.

## Memory Management with FlashInfer KV-Cache

To prevent GPU memory exhaustion during long videos, LingBot-Map implements the **FlashInferKVCacheManager** in [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py). This wrapper around FlashInfer's paged attention kernels stores only **keyframes**—specifically every N-th frame as controlled by the `--keyframe_interval` flag.

The cache manager employs **paged attention** to organize keyframe embeddings into fixed-size memory pages. When the cache reaches capacity, it evicts older slots using a least-recently-used policy, keeping memory usage roughly constant even for videos exceeding thousands of frames. Non-keyframes still generate 3D predictions during processing but are not cached, significantly reducing memory pressure.

## Handling Long Videos: Windowed Inference Mode

When video sequences exceed the training RoPE (Rotary Position Embedding) range of approximately 320 frames, LingBot-Map automatically splits the stream into **overlapping windows**. This windowed inference mode is triggered via the `--mode windowed` flag in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py).

The windowing mechanism operates as follows:

1. When the current window reaches the `--window_size` limit (measured in KV-cache slots), the system creates a new window.
2. The KV-cache resets, and anchor tokens re-initialize to prevent drift.
3. The system copies overlapping keyframe KV entries specified by `--overlap_keyframes` to ensure pose continuity between windows.
4. Processing continues seamlessly without dropping frames.

This approach prevents **RoPE overflow** while maintaining geometric consistency across window boundaries through the overlap mechanism.

## Camera Pose Refinement Without Bundle Adjustment

Accurate camera tracking is achieved through the **camera head** in [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py). Rather than performing expensive global bundle adjustment, this component refines each frame's camera pose by attending to a **sliding window of recent keyframes**.

The pose-reference tokens generated by the AggregatorStream feed into the camera head, which iteratively refines extrinsic parameters. This local attention mechanism provides accurate pose estimation with computational costs linear in the window size rather than quadratic in the total frame count.

## Implementing the Streaming Pipeline

The repository provides multiple interfaces for running 3D reconstruction from video streams.

### Basic Streaming with Keyframe Intervals

For standard video processing with keyframe caching every 2 frames:

```bash
python demo.py \
    --image_folder example/university \
    --model_path /path/to/lingbot-map-long.pt \
    --keyframe_interval 2 \
    --mask_sky

```

### Windowed Inference for Long Videos

For sequences exceeding 3000 frames, use windowed mode to manage memory:

```bash
python demo.py \
    --video_path /data/long_walk.mp4 \
    --fps 10 \
    --model_path /path/to/lingbot-map-long.pt \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8 \
    --keyframe_interval 13 \
    --mask_sky

```

### Direct API Integration

Embed the streaming pipeline in custom Python code:

```python
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream
from lingbot_map.aggregator.stream import AggregatorStream

model = GCTStream.load_from_checkpoint("lingbot-map-long.pt")
model.eval().cuda()

# Assume `frames` is a list of (RGB, depth) tensors

for i, (rgb, depth) in enumerate(frames):
    # Build token stream for the current frame

    stream = AggregatorStream()
    stream.add_frame(rgb, depth)
    
    # Forward pass returns dense depth + camera pose

    out = model(stream.tokens, stream.kv_cache)
    # `out` contains reconstruction predictions for frame i

```

## Summary

- **LingBot-Map** processes video streams incrementally using `GCTStream`, a feed-forward transformer that reuses cached key/value pairs from previous frames.
- The **FlashInferKVCacheManager** in [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py) maintains constant memory usage by storing only keyframe embeddings in a paged cache structure.
- **AggregatorStream** handles token preparation, inserting anchor tokens for global context and pose-reference tokens for camera tracking.
- **Windowed inference** mode prevents RoPE overflow in long videos by splitting sequences into overlapping windows with configurable `--window_size` and `--overlap_keyframes` parameters.
- Camera poses refine locally through the **camera head** without global bundle adjustment, attending to recent keyframes via pose-reference tokens.

## Frequently Asked Questions

### What is the KV-cache in LingBot-Map?

The **KV-cache** (key-value cache) is a memory buffer managed by `FlashInferKVCacheManager` that stores embeddings of past keyframes. It enables the transformer to attend to previous video frames without reprocessing them, reducing computation from quadratic to constant time per frame. The cache uses FlashInfer's paged attention to organize memory into pages that can be efficiently accessed and evicted.

### How does LingBot-Map handle videos longer than 3000 frames?

For videos exceeding the model's RoPE limit of approximately 320 frames, LingBot-Map activates **windowed inference** via the `--mode windowed` flag. The system splits the video into overlapping segments defined by `--window_size`, resets the KV-cache at each window boundary, and copies `--overlap_keyframes` entries between windows to maintain pose continuity. This technique caps memory usage regardless of video length.

### What is the difference between keyframes and non-keyframes?

**Keyframes** are frames whose embeddings get stored in the KV-cache, controlled by the `--keyframe_interval` parameter (e.g., every 13th frame). **Non-keyframes** are processed through the transformer to generate 3D predictions but are immediately discarded from memory. This distinction allows the system to maintain high temporal resolution in output while keeping GPU memory consumption bounded.

### Why does LingBot-Map use windowed inference?

Windowed inference prevents **RoPE (Rotary Position Embedding) overflow**, which occurs when position indices exceed the training distribution (~320 frames). By splitting long sequences into overlapping windows and reinitializing anchor tokens at each split, the system maintains geometric accuracy without positional encoding degradation. The overlapping keyframes (`--overlap_keyframes`) ensure smooth transitions between windows, preventing reconstruction drift.