# How LingBot-Map Handles Loop Closure Trajectories: KV-Cache Streaming and Geometric Alignment

> Discover how LingBot-Map handles loop closure trajectories using KV-cache streaming and geometric alignment, ensuring global consistency without memory bloat.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: internals
- Published: 2026-07-26

---

**LingBot-Map processes loop closure trajectories through a three-tiered architecture combining streaming KV-cache sampling, sliding-window inference for sequences exceeding 320 frames, and anchor-context pose alignment to maintain global consistency without unbounded memory growth.**

Loop closure trajectories—where a camera revisits previously mapped areas—require specialized handling to prevent drift and maintain geometric consistency across thousands of frames. The Robbyant/lingbot-map repository implements these capabilities within its `GCTStream` model, leveraging a **Geometric Context Transformer (GCT)** architecture that balances memory efficiency with temporal coherence. By integrating paged key-value caching, configurable windowing strategies, and learned anchor tokens, the system enables robust 3D reconstruction during extended loop-closure sequences.

## Streaming KV-Cache with Keyframe Interval

The core model (`GCTStream`) operates in a **causal streaming mode** where a paged key-value (KV) cache stores attention context for selected frames only. This mechanism, implemented in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), prevents unbounded memory growth during long trajectories while preserving geometric consistency.

Every **keyframe**—controlled via the `--keyframe_interval` parameter—is retained in the KV cache, while non-keyframes still generate predictions but are **not** stored. For example, setting `--keyframe_interval 5` retains every fifth frame in cache memory, dramatically reducing GPU memory requirements for loop-closure sequences. You can inspect KV-cache statistics programmatically using `model.get_kv_cache_info()` or enable detailed logging by setting the environment variable `LINGBOT_DEBUG_KV=1`.

The `AggregatorStream` class in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) manages the integration of this cache into the transformer architecture, ensuring that only relevant geometric contexts persist across the trajectory.

## Sliding-Window Inference for Extended Sequences

When a loop closure trajectory exceeds the training-time RoPE range of approximately **320 frames**, the model automatically switches to **windowed mode** (`--mode windowed`). This strategy enables processing of ultra-long sequences—such as the repository's 25,000-frame indoor walkthrough example—without exhausting memory resources.

In windowed mode, the system:

1. Resets the KV cache at the beginning of each window
2. Processes blocks of frames defined by `--window_size` (default: 128)
3. Overlaps keyframes (`--overlap_keyframes`) across window boundaries to preserve temporal continuity

This overlapping approach ensures smooth transitions between windows, allowing the model to maintain pose consistency even when processing sequences of 30,000+ frames. The windowed inference logic resides alongside the streaming implementation in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py).

## Geometric Context Alignment via Anchors

Beyond caching strategies, the GCT architecture handles loop closure through **multi-modal context fusion** defined in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py). The model integrates three distinct context types to recognize revisited areas and align poses:

- **Anchor context**: Learnable tokens that ground scene geometry and provide stable reference points across the trajectory
- **Pose-reference window**: A sliding window of recent camera poses that guides temporal consistency between frames
- **Trajectory memory**: The KV cache itself, storing past visual features for comparison against current views

When the camera revisits a previously mapped area, the anchor context mechanism reconciles the current view with historical features stored in the trajectory memory. This alignment process closes the loop without accumulating drift, producing globally consistent 3D reconstructions even in complex indoor environments.

## Practical Implementation Examples

Run loop closure trajectories using the CLI entry point in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) or integrate directly via the Python API.

### Basic Loop Closure Demo

Process the built-in `example/loop` scene with default settings (keyframe interval of 1):

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --image_folder example/loop

```

### Memory-Optimized Long Trajectory

Reduce KV-cache memory for extended sequences by keeping every fifth frame:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --image_folder /data/long_loop_seq \
    --keyframe_interval 5

```

### Ultra-Long Windowed Inference

Enable windowed mode for sequences exceeding 30,000 frames:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path /data/loop_video.mp4 \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8 \
    --keyframe_interval 10

```

### Direct Python API Usage

Instantiate `GCTStream` with custom cache parameters for fine-grained control:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    enable_camera=True,
    enable_depth=True,
    sliding_window_size=-1,          # full causal KV cache

    kv_cache_sliding_window=64,      # eviction policy

    kv_cache_cross_frame_special=True,
)

# Load checkpoint

state = torch.load("lingbot-map.pt", map_location="cpu")
model.load_state_dict(state)

# Streaming inference with custom keyframe sampling

imgs = torch.randn(1, 1200, 3, 518, 378)  # [B, S, C, H, W]

preds = model.inference_streaming(
    imgs,
    keyframe_interval=4,               # cache every 4th frame

    output_device=torch.device("cpu")  # offload to CPU

)

```

## Summary

- **Streaming KV-cache** with configurable keyframe intervals prevents memory overflow by storing only selected frames (controlled via `--keyframe_interval`), while `LINGBOT_DEBUG_KV` enables detailed cache inspection.
- **Sliding-window inference** (`--mode windowed`) processes sequences beyond 320 frames by resetting the cache every `--window_size` frames with `--overlap_keyframes` ensuring continuity.
- **Anchor-context alignment** in the GCT architecture fuses learnable scene anchors, pose-reference windows, and trajectory memory to recognize revisited areas and close loops without drift.
- The `GCTStream` class in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) and base architecture in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py) provide the complete implementation for handling loop closure trajectories at scale.

## Frequently Asked Questions

### What is the maximum sequence length LingBot-Map can handle for loop closure?

According to the source code, LingBot-Map handles sequences of **25,000+ frames** using windowed inference mode. The training-time RoPE range limits continuous processing to approximately 320 frames, but the sliding-window mechanism with overlapping keyframes enables arbitrary sequence lengths by processing in 128-frame blocks (configurable via `--window_size`).

### How does the keyframe interval affect memory usage and accuracy?

The `--keyframe_interval` parameter directly trades memory for temporal resolution. Setting `keyframe_interval=5` reduces KV-cache memory by 80% compared to `interval=1`, but increases reliance on the pose-reference window for intermediate frames. For loop closure scenes with high visual overlap, intervals of 4-10 maintain accuracy while enabling processing of hour-long trajectories on consumer GPUs.

### What is the difference between streaming mode and windowed mode?

**Streaming mode** (`--mode stream` or default) maintains a persistent KV-cache across the entire sequence up to the RoPE limit (~320 frames), providing maximum temporal consistency. **Windowed mode** (`--mode windowed`) resets the KV-cache every `--window_size` frames (default 128), using overlapping keyframes to maintain continuity. Use streaming for short, dense loops under 320 frames; use windowed mode for extended trajectories like building-scale walkthroughs.

### How can I verify the KV-cache is functioning correctly during loop closure?

Set the environment variable `LINGBOT_DEBUG_KV=1` before running to enable detailed cache logging. Programmatically, call `model.get_kv_cache_info()` after `inference_streaming()` to retrieve statistics about cached frames, memory allocation, and eviction events. This is particularly useful when tuning `--keyframe_interval` for memory-constrained deployments.