How LingBot-Map Handles Long Sequences and Drift Correction

LingBot-Map processes arbitrarily long video streams by splitting the sequence into a sliding window of KV cache blocks and periodically correcting drift with a lightweight optical-flow estimator.

Robbyant/lingbot-map is a streaming visual localization system designed to operate on arbitrarily long video sequences without unbounded memory growth. The architecture combines a constant-memory attention mechanism with optical-flow-based drift detection to maintain accurate camera pose estimates over extended durations. This article examines the specific implementation details of how LingBot-Map handles long sequences and drift correction in the GCTStream class.

Sliding-Window KV Cache for Long Sequences

The core streaming pipeline resides in GCTStream, implemented in lingbot_map/models/gct_stream_window_v2.py (lines 78-87). When a new frame arrives, the model feeds it to AggregatorStream, which maintains a KV cache storing attention keys and values for recent frames.

The cache size is strictly bounded by the kv_cache_sliding_window parameter, which defaults to 64 frames (lines 122-132). When the cache reaches this limit, the oldest blocks are automatically evicted, ensuring memory usage remains constant regardless of video length. The clean_kv_cache method (lines 88-99) provides explicit clearing when a new sequence starts, preventing stale keys from leaking across independent videos.

Constant-Memory Attention Mechanism

The attention computation is backed by FlashInfer with a fallback to PyTorch SDPA. Because the KV cache is restricted to the sliding window, attention complexity remains O(window size) rather than O(total frames), enabling real-time inference on arbitrarily long streams.

Drift Detection and Correction

Optical-Flow-Based Drift Detection

LingBot-Map measures camera displacement between the current frame and the last keyframe using _compute_flow_magnitude (lines 92-110 in gct_stream_window_v2.py). This method projects the current depth map into the previous camera pose, computes pixel-wise displacement, and returns the mean flow magnitude in pixels. Large flow values indicate significant deviation from the reference pose, triggering corrective action.

Camera-Head Iterative Refinement

The CameraCausalHead, constructed in _build_camera_head (lines 142-151), receives the current feature tokens and estimated pose. It executes camera_num_iterations (default 4) of gradient-descent refinement to optimize the pose estimate. This head utilizes the flow magnitude as an additional signal to re-anchor the pose when drift exceeds acceptable thresholds.

Automatic Cache Reset

When the detected flow magnitude exceeds a hard-coded or user-defined limit, the higher-level inference loop (typically in benchmark/methods/lingbot_map.py) invokes model.clean_kv_cache(). This discards stale context and initializes a fresh window, preventing error accumulation over very long runs.

Practical Implementation Example

The following example demonstrates initializing a streaming model with drift monitoring:

import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

# Initialize streaming model with KV cache and camera head

model = GCTStream(
    sliding_window_size=64,          # Retain last 64 frames

    kv_cache_sliding_window=64,      # Explicit cache window size

    camera_num_iterations=4,         # Iterative pose refinement steps

)

# Process video frame-by-frame

for frame_idx, (rgb, depth) in enumerate(video_frames):
    # Forward pass: rgb [B,3,H,W], depth [B,1,H,W,1]

    outputs = model(rgb.unsqueeze(0), depth.unsqueeze(0))
    
    # Detect drift via mean optical flow magnitude

    flow_mag = model._compute_flow_magnitude(
        cur_pose_enc=outputs['pose_encoding'],
        kf_pose_enc=outputs['keyframe_pose_encoding'],
        cur_depth=depth.unsqueeze(0),
        image_size_hw=rgb.shape[-2:],
    )
    
    # Reset cache if drift exceeds threshold (e.g., 30 pixels)

    if flow_mag > 30.0:
        model.clean_kv_cache()
        print(f"[Drift] Reset KV cache at frame {frame_idx}")

Summary

  • Bounded Memory: The kv_cache_sliding_window parameter (default 64) in AggregatorStream ensures constant memory usage by evicting old attention blocks, keeping complexity at O(window size).
  • Drift Detection: The _compute_flow_magnitude method in gct_stream_window_v2.py quantifies camera displacement by projecting depth maps and calculating mean optical flow.
  • Pose Refinement: CameraCausalHead performs iterative gradient descent (camera_num_iterations=4) to refine pose estimates when drift is detected.
  • Cache Hygiene: The clean_kv_cache method (lines 88-99) allows explicit resetting of the KV cache between sequences or when drift thresholds are breached, preventing error propagation.

Frequently Asked Questions

What is the default sliding window size for the KV cache?

The default kv_cache_sliding_window is 64 frames, defined in the GCTStream initialization in lingbot_map/models/gct_stream_window_v2.py (lines 122-132). This value can be adjusted based on available GPU memory and the temporal coherence requirements of your specific application.

How does LingBot-Map detect camera drift?

Drift detection relies on _compute_flow_magnitude (lines 92-110), which computes the mean optical flow in pixels between the current frame and the last keyframe. By projecting the current depth map into the previous camera coordinate system and measuring pixel displacement, the system quantifies how far the estimated pose has diverged from the reference.

What happens when drift exceeds the threshold?

When the flow magnitude exceeds a defined threshold (e.g., 30 pixels), the inference loop typically calls model.clean_kv_cache() to discard all stale attention blocks. Simultaneously, the CameraCausalHead performs iterative refinement to re-anchor the pose estimate, effectively resetting the localization context to prevent catastrophic drift accumulation.

Can the KV cache be manually cleared between sequences?

Yes. The clean_kv_cache method (lines 88-99 in gct_stream_window_v2.py) is specifically designed for this purpose. It should be invoked when switching between independent video files or when the system detects a loop closure or scene cut, ensuring no attention keys leak across unrelated sequences.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →