# LingBot-Map Architecture Explained: Core Components of the Geometric Context Transformer

> Explore the LingBot-Map architecture and its Geometric Context Transformer. Learn how it fuses multi-scale vision, 3D position, and causal attention for real-time 3D reconstruction.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: architecture
- Published: 2026-07-28

---

**LingBot-Map architecture centers on a Geometric Context Transformer (GCT) that fuses multi-scale visual features with 3-D positional encoding and causal attention to enable real-time, memory-efficient streaming 3D reconstruction.**

The LingBot-Map architecture powers the `Robbyant/lingbot-map` repository's approach to dense 3D reconstruction from monocular video sequences. Built around a streaming transformer design, this architecture processes frames causally while maintaining temporal geometric consistency through a novel KV-caching mechanism and 3-D Rotary Positional Encoding (RoPE).

## Core Components of the LingBot-Map Architecture

The LingBot-Map architecture comprises three tightly coupled modules that operate within the `GCTStream` class defined in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py).

### Feature Aggregator (AggregatorStream)

The **Feature Aggregator** extracts multi-scale visual tokens from each input frame and manages the temporal memory through a **paged KV-cache**. Implemented via `_build_aggregator()` in [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py), this component returns an `AggregatorStream` instance configured with cache parameters including `kv_cache_sliding_window` and `kv_cache_scale_frames`. The aggregator optionally applies **3-D RoPE** (Rotary Positional Encoding) to maintain temporal consistency across the video sequence.

### Camera Head (CameraCausalHead)

The **Camera Head** refines per-frame camera poses using accumulated tokens from the aggregator. Created by `_build_camera_head()` in the same source file, the `CameraCausalHead` processes geometry in a strictly causal manner, supporting iterative refinement through the `camera_num_iterations` parameter. This module receives identical KV-cache settings to the aggregator and can be driven by keyframe logic to optimize computation.

### Streaming Engine (GCTStream)

The **Streaming Engine** orchestrates the aggregator and camera head modules while managing the KV-cache lifecycle. The `GCTStream` class provides `inference_streaming()` for frame-by-frame processing and `inference_windowed()` for long sequences. The engine implements cache cleaning, "skip-append" logic for non-keyframes via `_set_skip_append(True)`, and maintains statistics to prevent unbounded memory growth.

## Architectural Pipeline and Data Flow

The LingBot-Map architecture processes video through a five-phase pipeline that balances geometric accuracy with computational efficiency.

1. **Input preprocessing**: Images are converted to tensor format `[S, 3, H, W]` and optionally resized via `LingbotMapMethod._prepare_images` in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py).

2. **Scale frame initialization**: The first *N* frames (default 8) are processed together using bidirectional attention through a dedicated scale token. This establishes the initial scene scale and provides baseline depth and pose estimates.

3. **KV-cache streaming**: Subsequent frames are processed causally. When `keyframe_interval=1`, every frame's key-value pairs are cached. With `keyframe_interval > 1`, only every *k*-th frame is stored; non-keyframes trigger `_set_skip_append(True)` to discard their KV tensors immediately after the forward pass.

4. **Windowed eviction**: For sequences exceeding memory limits, the architecture supports sliding-window cache eviction via `window_size` and `overlap_size` parameters. The `inference_windowed()` method in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py) wraps this functionality and handles flow-based keyframe selection.

5. **Post-processing**: Raw pose encodings (`pose_enc`) are decoded into extrinsic and intrinsic matrices using `pose_encoding_to_extri_intri`, while depth maps are reshaped and optional confidence maps are attached to the output.

## Key Implementation Files

Understanding the LingBot-Map architecture requires familiarity with these critical source files:

- [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py): Contains the core `GCTStream` implementation, including the aggregator builder, camera head builder, and inference loops.

- [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py): Implements the windowed inference variant with sliding-window KV eviction for sequences exceeding the standard 320-frame RoPE training range.

- [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py): Provides the benchmark adapter that handles checkpoint loading, image preparation, and output formatting for evaluation pipelines.

- [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py): Demonstrates the end-to-end offline rendering pipeline for long video sequences.

## Implementation Examples

### Running Streaming Inference

The following bash command executes the interactive demo in streaming mode with keyframe-based memory management:

```bash
python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/oxford \
    --mask_sky \
    --keyframe_interval 2

```

This internally instantiates `LingbotMapMethod`, loads the checkpoint, and calls `inference_streaming()` on the `GCTStream` model. The KV-cache grows only on keyframes, maintaining bounded GPU memory usage.

### Programmatic Model Usage

For custom applications, instantiate the streaming engine directly:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

# Initialize model with KV-cache parameters

model = GCTStream(
    img_size=518,
    patch_size=14,
    enable_3d_rope=True,
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    use_sdpa=False,  # FlashInfer backend

)

# Load checkpoint

ckpt = torch.load("lingbot-map-long.pt", map_location="cpu")
model.load_state_dict(ckpt["model"], strict=False)
model.eval().cuda()

# Prepare images and run inference

preds = model.inference_streaming(
    imgs,
    num_scale_frames=8,
    keyframe_interval=4,
    output_device=torch.device("cpu")
)

# Decode poses

from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri
extr, intr = pose_encoding_to_extri_intri(preds["pose_enc"], imgs.shape[-2:])

```

### Windowed Inference for Long Sequences

For videos exceeding 10,000 frames, use the windowed implementation:

```python
from lingbot_map.models.gct_stream_window import GCTStream

model = GCTStream(
    img_size=518,
    patch_size=14,
    enable_3d_rope=True,
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    use_sdpa=False,
)

# Process with sliding windows

preds = model.inference_windowed(
    imgs,
    window_size=128,
    overlap_size=8,
    keyframe_interval=13,
    flow_threshold=0.0,
    max_non_keyframe_gap=30,
)

```

This mode automatically resets the KV cache after each window, enabling stable inference beyond the standard training range.

### Benchmark Integration

Integrate with evaluation frameworks using the provided adapter:

```python
from benchmark.methods.lingbot_map import LingbotMapMethod

method = LingbotMapMethod(
    checkpoint="lingbot-map-long.pt",
    device="cuda",
    mode="streaming",
    keyframe_interval="auto",
)

results = method.process_scene(gt_artifact)

```

The `process_scene()` method internally orchestrates `_prepare_images()`, `_run_inference()`, and `_process_outputs()` to return benchmark-compatible results.

### Offline Rendering Pipeline

For processing very long video sequences end-to-end:

```bash
python demo_render/batch_demo.py \
    --video_path /data/demo_videos/indoor_travel.MP4 \
    --output_folder /data/outputs/indoor_travel/ \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed --window_size 128 \
    --keyframe_interval 13 --overlap_keyframes 8 \
    --mask_sky \
    --save_predictions

```

This script calls `GCTStream.inference_windowed()` under the hood to generate MP4 fly-throughs with optional sky-mask visualizations.

## Summary

The LingBot-Map architecture delivers efficient streaming 3D reconstruction through several key innovations:

- **Geometric Context Transformer (GCT)** design combining visual feature aggregation with causal camera pose estimation
- **Paged KV-cache system** with configurable `kv_cache_sliding_window` and keyframe intervals to bound memory usage
- **3-D RoPE integration** for maintaining temporal geometric consistency across video sequences
- **Dual inference modes**: streaming for real-time applications and windowed for arbitrarily long sequences
- **Modular implementation** separating concerns between `AggregatorStream`, `CameraCausalHead`, and the `GCTStream` orchestration layer

## Frequently Asked Questions

### What is the Geometric Context Transformer in LingBot-Map?

The **Geometric Context Transformer (GCT)** is the core neural architecture that fuses multi-scale visual features with 3-D positional information through causal attention mechanisms. It consists of the `AggregatorStream` for feature extraction and the `CameraCausalHead` for pose refinement, both implemented in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py).

### How does LingBot-Map handle memory constraints during streaming?

The architecture implements a **keyframe-based KV-cache** that stores attention keys and values only at specified intervals. By setting `keyframe_interval > 1` and invoking `_set_skip_append(True)` for non-keyframes, the model discards intermediate tensors immediately after processing, keeping GPU memory bounded regardless of sequence length.

### What is the difference between streaming and windowed inference modes?

**Streaming mode** (`inference_streaming()`) processes frames sequentially with a monotonically growing KV-cache, suitable for real-time applications. **Windowed mode** (`inference_windowed()`) processes sequences in overlapping chunks with periodic cache resets, enabling processing of videos exceeding the 320-frame RoPE training limit while maintaining geometric consistency through `overlap_size` shared frames.

### Which files contain the core LingBot-Map architecture implementation?

The primary implementation resides in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), which defines the `GCTStream` class, `_build_aggregator()`, and `_build_camera_head()`. Windowed variants are in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py), while the benchmark wrapper is located at [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py).