# How LingBot-Map Handles Long Sequences and Drift Correction

> LingBot-Map processes long sequences and corrects drift using KV cache blocks and optical flow. Learn how this approach ensures accurate mapping for extended video streams.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-28

---

**LingBot-Map processes arbitrarily long video streams by splitting the sequence into a sliding window of KV cache blocks and periodically correcting drift with a lightweight optical-flow estimator.**

Robbyant/lingbot-map is a streaming visual localization system designed to operate on arbitrarily long video sequences without unbounded memory growth. The architecture combines a constant-memory attention mechanism with optical-flow-based drift detection to maintain accurate camera pose estimates over extended durations. This article examines the specific implementation details of how LingBot-Map handles long sequences and drift correction in the `GCTStream` class.

## Sliding-Window KV Cache for Long Sequences

The core streaming pipeline resides in `GCTStream`, implemented in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) (lines 78-87). When a new frame arrives, the model feeds it to `AggregatorStream`, which maintains a KV cache storing attention keys and values for recent frames.

The cache size is strictly bounded by the `kv_cache_sliding_window` parameter, which defaults to 64 frames (lines 122-132). When the cache reaches this limit, the oldest blocks are automatically evicted, ensuring memory usage remains constant regardless of video length. The `clean_kv_cache` method (lines 88-99) provides explicit clearing when a new sequence starts, preventing stale keys from leaking across independent videos.

### Constant-Memory Attention Mechanism

The attention computation is backed by FlashInfer with a fallback to PyTorch SDPA. Because the KV cache is restricted to the sliding window, attention complexity remains **O(window size)** rather than **O(total frames)**, enabling real-time inference on arbitrarily long streams.

## Drift Detection and Correction

### Optical-Flow-Based Drift Detection

LingBot-Map measures camera displacement between the current frame and the last keyframe using `_compute_flow_magnitude` (lines 92-110 in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py)). This method projects the current depth map into the previous camera pose, computes pixel-wise displacement, and returns the **mean flow magnitude** in pixels. Large flow values indicate significant deviation from the reference pose, triggering corrective action.

### Camera-Head Iterative Refinement

The `CameraCausalHead`, constructed in `_build_camera_head` (lines 142-151), receives the current feature tokens and estimated pose. It executes `camera_num_iterations` (default 4) of gradient-descent refinement to optimize the pose estimate. This head utilizes the flow magnitude as an additional signal to re-anchor the pose when drift exceeds acceptable thresholds.

### Automatic Cache Reset

When the detected flow magnitude exceeds a hard-coded or user-defined limit, the higher-level inference loop (typically in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py)) invokes `model.clean_kv_cache()`. This discards stale context and initializes a fresh window, preventing error accumulation over very long runs.

## Practical Implementation Example

The following example demonstrates initializing a streaming model with drift monitoring:

```python
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

# Initialize streaming model with KV cache and camera head

model = GCTStream(
    sliding_window_size=64,          # Retain last 64 frames

    kv_cache_sliding_window=64,      # Explicit cache window size

    camera_num_iterations=4,         # Iterative pose refinement steps

)

# Process video frame-by-frame

for frame_idx, (rgb, depth) in enumerate(video_frames):
    # Forward pass: rgb [B,3,H,W], depth [B,1,H,W,1]

    outputs = model(rgb.unsqueeze(0), depth.unsqueeze(0))
    
    # Detect drift via mean optical flow magnitude

    flow_mag = model._compute_flow_magnitude(
        cur_pose_enc=outputs['pose_encoding'],
        kf_pose_enc=outputs['keyframe_pose_encoding'],
        cur_depth=depth.unsqueeze(0),
        image_size_hw=rgb.shape[-2:],
    )
    
    # Reset cache if drift exceeds threshold (e.g., 30 pixels)

    if flow_mag > 30.0:
        model.clean_kv_cache()
        print(f"[Drift] Reset KV cache at frame {frame_idx}")

```

## Summary

- **Bounded Memory**: The `kv_cache_sliding_window` parameter (default 64) in `AggregatorStream` ensures constant memory usage by evicting old attention blocks, keeping complexity at O(window size).
- **Drift Detection**: The `_compute_flow_magnitude` method in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py) quantifies camera displacement by projecting depth maps and calculating mean optical flow.
- **Pose Refinement**: `CameraCausalHead` performs iterative gradient descent (`camera_num_iterations`=4) to refine pose estimates when drift is detected.
- **Cache Hygiene**: The `clean_kv_cache` method (lines 88-99) allows explicit resetting of the KV cache between sequences or when drift thresholds are breached, preventing error propagation.

## Frequently Asked Questions

### What is the default sliding window size for the KV cache?

The default `kv_cache_sliding_window` is **64 frames**, defined in the `GCTStream` initialization in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) (lines 122-132). This value can be adjusted based on available GPU memory and the temporal coherence requirements of your specific application.

### How does LingBot-Map detect camera drift?

Drift detection relies on `_compute_flow_magnitude` (lines 92-110), which computes the mean optical flow in pixels between the current frame and the last keyframe. By projecting the current depth map into the previous camera coordinate system and measuring pixel displacement, the system quantifies how far the estimated pose has diverged from the reference.

### What happens when drift exceeds the threshold?

When the flow magnitude exceeds a defined threshold (e.g., 30 pixels), the inference loop typically calls `model.clean_kv_cache()` to discard all stale attention blocks. Simultaneously, the `CameraCausalHead` performs iterative refinement to re-anchor the pose estimate, effectively resetting the localization context to prevent catastrophic drift accumulation.

### Can the KV cache be manually cleared between sequences?

Yes. The `clean_kv_cache` method (lines 88-99 in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py)) is specifically designed for this purpose. It should be invoked when switching between independent video files or when the system detects a loop closure or scene cut, ensuring no attention keys leak across unrelated sequences.