# How LingBot-Map Achieves Real-Time 3D Reconstruction: Streaming Transformers and KV Caching

> Discover how LingBot-Map achieves real-time 3D reconstruction using streaming Transformers and KV caching for efficient, constant-memory causal attention. Explore the technical details.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-28

---

**LingBot-Map achieves real-time 3D reconstruction by processing video frames through a streaming Vision Transformer that maintains a paged key-value (KV) cache and applies 3-D rotary position embeddings, enabling constant-memory causal attention across arbitrarily long sequences.**

The `Robbyant/lingbot-map` repository implements a novel approach to real-time 3D reconstruction that eliminates the memory bottlenecks of traditional batch-processing methods. By treating video streams as temporal volumes rather than independent images, the system continuously updates depth maps and point clouds at video frame rates. This architecture relies on three core technical innovations: the streaming GCT-Stream transformer, FlashInfer-based KV caching, and 3-D rotary position embeddings.

## Core Architecture Components

### GCT-Stream: A Streaming-Aware Vision Transformer

The foundation of the system is the **`GCTStream`** class defined in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py). Unlike standard Vision Transformers that process entire sequences in parallel, this model inherits from `GCTBase` and overrides **`_build_aggregator`** to instantiate a streaming aggregator. 

The class accepts critical streaming parameters including `kv_cache_sliding_window`, `enable_3d_rope`, and `use_sdpa`, configuring the model for frame-by-frame processing rather than batch inference. This design allows the model to ingest frames continuously without requiring the entire video to reside in memory simultaneously.

### AggregatorStream with FlashInfer Paged KV Cache

At the heart of the memory-efficient processing lies the **`AggregatorStream`** class in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py). This component implements a **paged KV cache** using FlashInfer, allowing the model to maintain attention states from previous frames without storing the entire video history in GPU memory.

The implementation includes several key mechanisms:

- **`_init_kv_cache`**: Initializes the FlashInfer manager (or a dictionary-based fallback) to handle cache read/write operations.
- **`_build_blocks`**: Constructs frame blocks and global blocks, utilizing `FlashInferBlock` (or `SDPABlock` when `use_sdpa=True`) where the paged KV cache resides.
- **Sliding-window eviction**: The `kv_cache_sliding_window` parameter bounds memory usage by evicting older frames when the cache exceeds the specified limit, while `kv_cache_scale_frames` and `kv_cache_cross_frame_special` control how scale and special tokens interact with the temporal cache.

### 3-D Rotary Position Embeddings for Temporal Consistency

To maintain geometric consistency across time, LingBot-Map extends standard 2-D RoPE into three dimensions. The **`WanRotaryPosEmbed`** class in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) computes frequency components for the temporal dimension alongside spatial coordinates. 

When the `enable_3d_rope` flag is activated (passed to both `AggregatorStream` and `CameraCausalHead`), the model applies rotation based on absolute frame indices, ensuring that positional encodings respect the temporal structure of the 3-D volume being reconstructed.

## The Real-Time Inference Pipeline

The system executes real-time 3D reconstruction through a carefully orchestrated sequence of operations that minimize latency while preserving geometric accuracy.

### Model Initialization and KV Cache Warm-Up

Inference begins when [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) or a custom script instantiates `GCTStream` with streaming-specific arguments. The constructor chains through `__init__` → `super().__init__` → `_build_aggregator` to create the `AggregatorStream` with the configured KV cache settings. 

Before processing the first frame, the aggregator lazily initializes the FlashInfer manager via `_get_flashinfer_manager`, and the initial forward pass populates the cache with baseline key/value tensors.

### Per-Frame Processing and Causal Attention

For each new video frame, the system performs the following operations:

1. **Feature aggregation**: The current frame tensor (shape `[B, 1, 3, H, W]`) passes through `model._aggregate_features()`, which invokes `AggregatorStream.__call__`.
2. **Cached attention retrieval**: The aggregator retrieves KV tensors from previous frames through the paged cache.
3. **Temporal causal attention**: The current frame's tokens attend only to cached past tokens and special tokens (camera or scale tokens), preventing future information leakage.
4. **3-D RoPE application**: If enabled, the system uses the frame index to rotate query/key embeddings via the temporal RoPE implementation, maintaining spatial-temporal coherence.

### Cache Eviction and Output Generation

After processing each frame, `AggregatorStream` increments `total_frames_processed` and evaluates the sliding-window policy. If the number of stored frames exceeds `kv_cache_sliding_window`, the oldest blocks are evicted from the FlashInfer cache, maintaining constant memory consumption regardless of video length. 

Simultaneously, the **`CameraCausalHead`** (also using the KV cache and optional 3-D RoPE) refines camera poses frame-by-frame. Finally, the `point_head` decodes aggregated tokens into depth maps and point clouds, visualized in real time via [`lingbot_map/vis/point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/point_cloud_viewer.py).

## Code Implementation Examples

### Python Streaming Inference

The following implementation demonstrates how to configure and run the streaming model for real-time 3D reconstruction:

```python

# ----------------------------------------------------------------------

# Streaming inference example (equivalent to the demo script)

# ----------------------------------------------------------------------

import torch
from lingbot_map.models.gct_stream import GCTStream
from lingbot_map.utils.load_fn import load_and_preprocess_images

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

# ---- Load a checkpoint -------------------------------------------------

model = GCTStream(
    img_size=518,
    patch_size=14,
    enable_3d_rope=True,            # activate temporal RoPE

    max_frame_num=1024,
    kv_cache_sliding_window=64,     # keep 64 recent frames in KV cache

    kv_cache_scale_frames=8,
    kv_cache_cross_frame_special=True,
    kv_cache_include_scale_frames=True,
    use_sdpa=False,                  # use FlashInfer (default)

    camera_num_iterations=4,
).to(device).eval()

# ---- Prepare a short video sequence ------------------------------------

image_paths = ["frame000.jpg", "frame001.jpg", "frame002.jpg"]
images = load_and_preprocess_images(
    image_paths,
    mode="crop",
    image_size=518,
    patch_size=14,
).to(device)      # shape: [B, S, 3, H, W]

# ---- Run streaming inference -------------------------------------------

# The model will automatically manage the KV cache across frames.

with torch.no_grad():
    for i in range(images.shape[1]):            # iterate over time dimension

        cur_img = images[:, i:i+1]             # shape: [B, 1, 3, H, W]

        # Aggregator returns tokens; camera head returns pose refinements, etc.

        tokens, _ = model._aggregate_features(cur_img)
        depth, point_cloud = model.point_head(tokens)   # example decode step

        # … visualize or save depth / point cloud here …

```

### Command-Line Interface

For direct execution from the repository root:

```bash

# ----------------------------------------------------------------------

# Command‑line demo (run from the repository root)

# ----------------------------------------------------------------------

python demo.py \
    --model_path checkpoints/gct_stream.pt \
    --image_folder ./my_video_frames \
    --mode streaming \
    --enable_3d_rope \
    --kv_cache_sliding_window 64

```

The demo script handles KV cache warm-up, per-frame forwarding, and real-time visualization automatically.

## Summary

- **Streaming Architecture**: The `GCTStream` class in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) enables frame-by-frame processing by overriding the standard aggregator with a streaming implementation.
- **Memory Efficiency**: `AggregatorStream` utilizes FlashInfer's paged KV cache with configurable sliding-window eviction (`kv_cache_sliding_window`) to maintain constant GPU memory usage during arbitrarily long video sequences.
- **Temporal Coherence**: 3-D rotary position embeddings (`WanRotaryPosEmbed` in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py)) encode absolute frame indices into the attention mechanism, preserving geometric consistency across the reconstructed volume.
- **Real-Time Performance**: By combining causal attention, cached state management, and optimized pose refinement through `CameraCausalHead`, the system achieves video-rate reconstruction (approximately 10 fps) while continuously updating 3-D geometry.

## Frequently Asked Questions

### What is the KV cache in LingBot-Map and why is it necessary for real-time 3D reconstruction?

The KV cache is a memory buffer that stores key and value tensors from previous video frames, allowing the transformer to attend to past information without recomputing attention for the entire sequence. According to the source code in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), the paged KV cache implementation using FlashInfer enables the model to process streaming video with constant memory usage rather than linear growth, which is essential for real-time operation on limited GPU resources.

### How does 3-D RoPE differ from standard 2-D positional encoding in Vision Transformers?

Standard 2-D RoPE encodes spatial positions within individual images, while 3-D RoPE extends this to include the temporal dimension. As implemented in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py), the `WanRotaryPosEmbed` class computes rotation frequencies for frame indices alongside spatial coordinates, allowing the model to treat the video as a 3-D volume and maintain geometric consistency across time steps during reconstruction.

### Can LingBot-Map process arbitrarily long videos without running out of memory?

Yes, the sliding-window KV cache mechanism ensures bounded memory consumption regardless of video length. The `kv_cache_sliding_window` parameter in `AggregatorStream` configures how many recent frames remain in the FlashInfer cache, automatically evicting older frames while preserving the temporal context necessary for coherent 3-D reconstruction.

### What is the difference between GCT-Stream and the windowed variant?

The standard `GCTStream` class processes video as a continuous stream using a single KV cache, while the windowed variant in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) splits very long sequences into overlapping windows to manage extreme-length videos. Both implementations inherit from `GCTBase` and use the same underlying `AggregatorStream` mechanism, but the windowed version provides additional segmentation for sequences exceeding the memory capacity of even the sliding-window cache.