# Architecture of LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction

> Explore the LingBot-Map architecture featuring a Geometric Context Transformer for efficient, real-time 3D reconstruction from video streams. Discover its KV-cached Feature Aggregator and causal Camera Head.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: architecture
- Published: 2026-07-27

---

**LingBot-Map implements a Geometric Context Transformer (GCT) that combines a KV-cached Feature Aggregator with a causal Camera Head to enable real-time, memory-bounded streaming 3D reconstruction from video sequences.**

LingBot-Map is an open-source framework for streaming 3D reconstruction built around a **Geometric Context Transformer (GCT)**. The architecture fuses multi-scale visual features with 3-D positional encoding through a causal attention mechanism, enabling frame-by-frame processing without unbounded memory growth. At its core, the system uses a paged KV-cache strategy with configurable keyframe intervals to handle arbitrarily long video sequences on consumer GPUs.

## Core Architectural Components

The architecture consists of three tightly coupled modules implemented in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py):

### Feature Aggregator (AggregatorStream)

The **AggregatorStream** extracts multi-scale visual tokens from each input frame and manages temporal context through a **paged KV-cache**. This component supports **3-D RoPE (Rotary Positional Encoding)** for temporal consistency across frames.

Key implementation details in `_build_aggregator()` include:
- Configuring `kv_cache_sliding_window` for maximum cache size
- Setting `kv_cache_scale_frames` (typically 8) for initial scale estimation
- Storing visual tokens using FlashInfer-backed paging when `use_sdpa=False`

### Camera Head (CameraCausalHead)

The **CameraCausalHead** refines per-frame camera poses using accumulated tokens from the aggregator. Implemented in `_build_camera_head()`, this module supports:
- **Iterative refinement** via `camera_num_iterations`
- **Causal processing** that only attends to previous frames
- Optional 3-D RoPE specific to the camera pose branch

### Streaming Engine (GCTStream)

The **GCTStream** class orchestrates the aggregator and camera head, managing the KV-cache lifecycle through `inference_streaming()`. The engine provides:
- **Skip-append logic** for non-keyframes to limit cache growth
- **Two-phase processing**: bidirectional attention on scale frames, followed by causal streaming
- **Automatic device management** for offloading predictions to CPU

## How the Streaming Pipeline Works

The architecture processes video through a defined lifecycle that balances reconstruction quality with memory constraints.

### Scale Frame Initialization

The first *N* frames (default 8, configured via `num_scale_frames`) undergo bidirectional attention processing to establish scene scale. This "scale token" phase yields initial depth and pose estimates that anchor the coordinate system for subsequent causal processing.

### KV-Cache Management and Keyframe Logic

After initialization, the system switches to causal processing with configurable KV-cache eviction:
- **`keyframe_interval = 1`**: Every frame's keys and values are cached (high quality, high memory)
- **`keyframe_interval > 1`**: Only every *k*-th frame triggers `_set_skip_append(False)`; non-keyframes use `_set_skip_append(True)` to discard KV tensors after processing

This mechanism keeps GPU memory bounded even for 10,000+ frame sequences.

### Windowed Inference for Long Sequences

For sequences exceeding the training context window (320 frames), the `inference_windowed()` method (available in both [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py) and [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py)) implements sliding-window processing:
- **`window_size`**: Maximum KV slots per window (e.g., 128)
- **`overlap_size`**: Shared keyframes between adjacent windows (e.g., 8)
- **Automatic cache reset** between windows to prevent position encoding overflow

## Key Implementation Files

| File Path | Architectural Role |
|-----------|-------------------|
| [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) | Core GCT implementation with `GCTStream`, `AggregatorStream`, and `CameraCausalHead` classes |
| [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) | Windowed inference variant with sliding-window KV eviction |
| [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py) | Benchmark adapter providing `LingbotMapMethod` for evaluation pipelines |
| [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py) | End-to-end rendering pipeline for long video sequences |

## Practical Usage Examples

### Running the Interactive Streaming Demo

Execute the streaming mode with configurable keyframe intervals to balance speed and accuracy:

```bash
python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/oxford \
    --mask_sky \
    --keyframe_interval 2

```

This command processes frames through `LingbotMapMethod` → `GCTStream.inference_streaming()`, caching only every second frame's KV tensors.

### Programmatic Inference with GCTStream

Build and run the model directly for custom pipelines:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri

# Initialize streaming model

model = GCTStream(
    img_size=518,
    patch_size=14,
    enable_3d_rope=True,
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    use_sdpa=False,  # Enables FlashInfer backend

)

# Load checkpoint

ckpt = torch.load("lingbot-map-long.pt", map_location="cpu")
model.load_state_dict(ckpt["model"], strict=False)
model.eval().cuda()

# Prepare RGB frames [S, 3, H, W]

preds = model.inference_streaming(
    imgs,
    num_scale_frames=8,
    keyframe_interval=4,
    output_device=torch.device("cpu")
)

# Decode to extrinsic/intrinsic matrices

extr, intr = pose_encoding_to_extri_intri(preds["pose_enc"], imgs.shape[-2:])

```

### Windowed Processing for Long Videos

Handle sequences beyond the RoPE training range using windowed mode:

```python
from lingbot_map.models.gct_stream_window import GCTStream

model = GCTStream(...)

# Load checkpoint...

preds = model.inference_windowed(
    imgs,
    window_size=128,
    overlap_size=8,
    keyframe_interval=13,
    flow_threshold=0.0,
    max_non_keyframe_gap=30,
)

```

This pattern automatically manages KV-cache resets between windows while maintaining temporal continuity through overlapping keyframes.

### Benchmark Integration

Use the provided adapter for standardized evaluation:

```python
from benchmark.methods.lingbot_map import LingbotMapMethod

method = LingbotMapMethod(
    checkpoint="lingbot-map-long.pt",
    device="cuda",
    mode="streaming",
    keyframe_interval="auto",  # Auto-select based on sequence length

)

results = method.process_scene(gt_artifact)

```

The `process_scene()` method handles image preparation via `_prepare_images()`, inference via `_run_inference()`, and output formatting via `_process_outputs()`.

## Summary

- **LingBot-Map** implements a **Geometric Context Transformer (GCT)** combining a KV-cached Feature Aggregator with a causal Camera Head for streaming 3D reconstruction.
- The architecture uses **3-D RoPE** and **paged KV-caches** with configurable `keyframe_interval` settings to maintain bounded GPU memory during causal processing.
- **Two-phase processing** handles scale initialization (bidirectional) followed by causal streaming, while **windowed inference** in `inference_windowed()` supports arbitrary sequence lengths.
- Core implementation resides in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), with benchmark adapters in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py) and rendering tools in [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py).

## Frequently Asked Questions

### What makes LingBot-Map suitable for real-time streaming applications?

The **causal attention mechanism** and **KV-cache management** enable frame-by-frame processing without revisiting previous frames. By configuring `keyframe_interval` and using `_set_skip_append()` for non-keyframes, the system maintains constant GPU memory regardless of video length. The `CameraCausalHead` processes each frame using only accumulated history from the `AggregatorStream`'s cache, eliminating the quadratic memory growth typical of standard transformers.

### How does the keyframe interval affect reconstruction quality and performance?

Setting `keyframe_interval=1` stores every frame's KV tensors, maximizing temporal coherence but consuming significant VRAM. Increasing the interval (e.g., to 4 or 13) reduces memory usage linearly while the **CameraCausalHead** continues to output poses for every frame. The `_set_skip_append(True)` mechanism ensures non-keyframes compute forward passes without polluting the cache, making the trade-off between memory and drift customizable per hardware constraints.

### What is the difference between streaming mode and windowed mode?

**Streaming mode** (`inference_streaming()`) processes sequences continuously with a monotonic KV-cache, suitable for online scenarios up to the RoPE training limit (320 frames). **Windowed mode** (`inference_windowed()`) divides long sequences into overlapping chunks (controlled by `window_size` and `overlap_size`), resetting the KV-cache between windows to handle 10,000+ frame videos without position encoding overflow. Windowed mode is implemented in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py) and used in [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py) for offline batch processing.

### Where is the camera pose decoding implemented?

Raw pose encodings from the model are converted to standard extrinsic and intrinsic matrices via `pose_encoding_to_extri_intri()` in [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py). This function accepts the model's `pose_enc` output tensor and image dimensions, returning camera matrices compatible with standard computer vision benchmarks. The benchmark adapter in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py) automatically handles this conversion in `_process_outputs()`.