# Core Components of the LingBot-Map Architecture: A Deep Dive into the Geometric Context Transformer

> Explore the LingBot-Map architecture and its core Geometric Context Transformer. Discover its Feature Aggregator, Camera Head, and Streaming Engine for real-time 3D reconstruction.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-23

---

**The LingBot-Map architecture centers on a Geometric Context Transformer (GCT) that combines a Feature Aggregator with KV-caching, a causal Camera Head, and a Streaming Engine to enable real-time 3D reconstruction from long video sequences.**

LingBot-Map is an open-source streaming 3D reconstruction system that processes visual sequences through a transformer-based causal attention mechanism. The architecture integrates **3-D Rotary Positional Encoding (RoPE)** with a paged KV-cache to handle arbitrarily long videos without linear memory growth. All core modules reside in the `Robbyant/lingbot-map` repository, specifically within the `lingbot_map/models/` directory.

## The Three Pillars of the LingBot-Map Architecture

The system decomposes into three tightly coupled components that operate within the `GCTStream` class defined in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py).

### Feature Aggregator (AggregatorStream)

The **Feature Aggregator** extracts multi-scale visual tokens from each input frame and manages temporal consistency through a sophisticated caching mechanism. Implemented via `_build_aggregator()` in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), this component returns an `AggregatorStream` object configured with KV-cache parameters including `kv_cache_sliding_window` and `kv_cache_scale_frames`.

The aggregator optionally applies **3-D RoPE** (Rotary Positional Encoding) to maintain geometric consistency across frames. It stores visual tokens in a **paged KV-cache**, which enables efficient memory management during streaming inference.

### Camera Head (CameraCausalHead)

The **Camera Head** refines per-frame camera poses using a causal attention mechanism that only attends to previous frames. Created by `_build_camera_head()` in the same file, the `CameraCausalHead` receives the same KV-cache settings as the aggregator and supports iterative refinement through the `camera_num_iterations` parameter.

This component can apply 3-D RoPE specific to the camera branch and operates on "keyframe" logic to limit computational overhead. The causal design ensures that pose estimation for frame $t$ depends only on frames $0$ through $t-1$, maintaining temporal coherence without future context.

### Streaming Engine (GCTStream)

The **Streaming Engine** orchestrates the aggregator and camera head through the `GCTStream` class. This engine manages the forward pass frame-by-frame, handles KV-cache lifecycle operations, and provides both **streaming** and **windowed** inference modes.

Key methods include `inference_streaming()`, which implements a two-phase pipeline processing scale frames bidirectionally followed by per-frame causal processing. The engine also manages the keyframe-interval mechanism that limits cache growth by invoking `_set_skip_append(True)` for non-keyframe frames.

## The Streaming Inference Pipeline

Understanding how these components interact requires examining the five-stage pipeline implemented in the LingBot-Map architecture.

### Input Preprocessing and Scale Frames

Images enter the system as tensors with shape `[S, 3, H, W]` through `LingbotMapMethod._prepare_images` in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py). The first *N* frames (default 8) serve as **scale frames**, processed together using bidirectional attention via a scale token to establish initial scene scale and depth estimates.

### KV-Cache Streaming with Keyframe Selection

After processing scale frames, the system switches to causal streaming mode. When `keyframe_interval=1`, every frame's key-value tensors are cached. With `keyframe_interval > 1`, only every *k*-th frame is stored; non-keyframes trigger `_set_skip_append(True)`, discarding their KV tensors immediately after the forward pass to bound GPU memory usage.

### Windowed Mode for Long Sequences

For sequences exceeding the 320-frame RoPE training range, the architecture supports sliding-window eviction through `inference_windowed()` in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py). This mode uses `window_size` and `overlap_size` parameters to reset the KV cache periodically while maintaining continuity between windows.

### Post-Processing and Output Decoding

Raw pose encodings (`pose_enc`) are decoded into extrinsic and intrinsic matrices using `pose_encoding_to_extri_intri()`. Depth maps are reshaped to match input dimensions, and optional confidence maps are attached to the output dictionary.

## Key Implementation Files

The LingBot-Map architecture spans several critical files in the repository:

- **[`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py)**: Contains the core `GCTStream` class, `_build_aggregator()`, `_build_camera_head()`, and the `inference_streaming()` method.
- **[`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py)**: Implements the windowed inference variant with sliding-window KV eviction.
- **[`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py)**: Provides the benchmark adapter that handles checkpoint loading, image preparation, and result formatting.
- **[`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py)**: Showcases the streaming engine on long video sequences with offline rendering capabilities.

## Practical Usage Examples

### Running Interactive Streaming Inference

Execute the streaming mode with controlled KV-cache growth using the command-line interface:

```bash
python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/oxford \
    --mask_sky \
    --keyframe_interval 2

```

Internally, this invokes `LingbotMapMethod` which loads the checkpoint and calls `inference_streaming()` on the `GCTStream` model.

### Programmatic Model Initialization

Instantiate the streaming architecture directly in Python:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    img_size=518,
    patch_size=14,
    enable_3d_rope=True,
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    use_sdpa=False,  # Uses FlashInfer backend for fast KV paging

)

ckpt = torch.load("lingbot-map-long.pt", map_location="cpu")
model.load_state_dict(ckpt["model"], strict=False)
model.eval().cuda()

```

### Streaming Inference with Keyframe Control

Process video frames with explicit keyframe intervals to manage memory:

```python
rgb_list = [...]  # List of HxWx3 uint8 frames

imgs = torch.stack([torchvision.transforms.ToTensor()(im) for im in rgb_list]).cuda()

preds = model.inference_streaming(
    imgs,
    num_scale_frames=8,
    keyframe_interval=4,
    output_device=torch.device("cpu")
)

from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri
extr, intr = pose_encoding_to_extri_intri(preds["pose_enc"], imgs.shape[-2:])

```

### Windowed Inference for Extended Videos

Handle sequences exceeding 10,000 frames using sliding windows:

```python
from lingbot_map.models.gct_stream_window import GCTStream

model = GCTStream(
    img_size=518,
    patch_size=14,
    enable_3d_rope=True,
    kv_cache_sliding_window=64,
    use_sdpa=False,
)

preds = model.inference_windowed(
    imgs,
    window_size=128,
    overlap_size=8,
    keyframe_interval=13,
    flow_threshold=0.0,
    max_non_keyframe_gap=30,
)

```

### Benchmark Adapter Integration

Use the standardized interface for evaluation pipelines:

```python
from benchmark.methods.lingbot_map import LingbotMapMethod

method = LingbotMapMethod(
    checkpoint="lingbot-map-long.pt",
    device="cuda",
    mode="streaming",
    keyframe_interval="auto",
)

results = method.process_scene(gt_artifact)

```

The `process_scene()` method internally calls `_prepare_images()`, `_run_inference()`, and `_process_outputs()` to return benchmark-compatible results.

## Summary

The LingBot-Map architecture delivers streaming 3D reconstruction through three integrated components:

- **Feature Aggregator (AggregatorStream)**: Extracts multi-scale visual tokens and manages the paged KV-cache with configurable sliding windows.
- **Camera Head (CameraCausalHead)**: Performs causal pose refinement with iterative optimization and 3-D RoPE support.
- **Streaming Engine (GCTStream)**: Orchestrates frame-by-frame processing with keyframe-based memory management and dual inference modes.

The system achieves unbounded sequence processing via `_set_skip_append()` logic for non-keyframes and sliding-window eviction in `inference_windowed()`, keeping GPU memory constant regardless of video length.

## Frequently Asked Questions

### What is the Geometric Context Transformer in LingBot-Map?

The **Geometric Context Transformer (GCT)** is the foundational neural architecture that fuses visual features, 3-D positional encoding, and causal attention mechanisms. It enables the system to perform streaming 3D reconstruction by processing frames sequentially while maintaining geometric consistency through 3-D RoPE and bounded memory usage through KV-cache management.

### How does the KV-cache manage memory in LingBot-Map?

The KV-cache implements a **paged storage system** with two primary constraints: `kv_cache_sliding_window` defines the maximum cache size, while `keyframe_interval` controls how frequently frames are stored. When processing non-keyframes, the system invokes `_set_skip_append(True)` to discard KV tensors immediately after the forward pass, preventing linear memory growth during long sequences.

### What is the difference between streaming and windowed inference modes?

**Streaming mode** (`inference_streaming()`) processes videos continuously with a monotonically growing KV-cache that respects `keyframe_interval` constraints. **Windowed mode** (`inference_windowed()`) periodically resets the cache using `window_size` and `overlap_size` parameters, enabling processing of sequences that exceed the 320-frame RoPE training limit by treating the video as overlapping chunks.

### Which files contain the core LingBot-Map implementation?

The primary implementation resides in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), which contains the `GCTStream` class, `_build_aggregator()`, and `_build_camera_head()` methods. Windowed inference variants are in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py), while the benchmark integration layer is implemented in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py).