How LingBot-Map's Paged KV Cache Handles Long Video Sequences

LingBot-Map implements a paged KV cache that combines sliding-window eviction, scale-frame retention, and cross-frame special tokens to maintain bounded GPU memory while preserving critical context across thousands of video frames.

LingBot-Map (available at Robbyant/lingbot-map) processes high-resolution 3-D mapping tasks from extended video sequences. To prevent GPU memory exhaustion during transformer inference, the project implements a sophisticated paged key-value cache mechanism primarily defined in lingbot_map/models/gct_stream_window_v2.py.

Core Architecture Components

Sliding-Window Eviction

The cache enforces a strict memory bound through the kv_cache_sliding_window parameter (defined at lines 223–227). This setting establishes a fixed-size window that retains only the most recent N frames. When new frames arrive, the manager automatically discards the oldest entries by slicing the KV tensors.

Scale-Frame Retention

Beyond the sliding window, the system preserves critical keyframes using kv_cache_scale_frames (lines 223–227). These scale frames provide global scene understanding and remain cached even if they fall outside the temporal window, provided kv_cache_include_scale_frames is enabled.

Cross-Frame Special Tokens

When standard frames are evicted, the kv_cache_cross_frame_special mechanism (lines 223–227) copies their special tokens into dedicated *_special slots rather than discarding them entirely. This allows later frames to attend to summarized information from distant past frames without retaining the full KV tensors.

Camera-Only Mode

For scenarios requiring minimal memory footprint, kv_cache_camera_only (lines 223–227) retains only camera-related tokens from evicted frames while discarding all other tensor data.

Lifecycle of the KV Cache

Construction and Initialization

When instantiating GCTStreamWindow or GCTStreamWindowV2, the system saves cache parameters as member variables including self.kv_cache_sliding_window and self.kv_cache_scale_frames around line 276. These values configure the low-level KVCacheManager that physically holds the tensors.

Appending New KV Pairs

During each forward pass, the model produces new k and v tensors. The aggregator's kv_cache_manager appends these to the cache unless the _skip_append flag is set (handled at lines 414–420). When skip=True, the attention layer reads from [cached_kv + current_kv] but does not persist the current KV, enabling read-only passes.

Eviction Logic

After appending, the manager checks cache size against the sliding-window limit. If exceeded, it removes the oldest frame(s) using the logic at lines 458–463:


# Slicing operation to drop the oldest time-step

kv[key] = kv[key][:, :, :-1]

# If the tensor's third dimension reaches size 1, clear it entirely

if kv[key].shape[2] == 1:
    kv[key] = None

Deferred Eviction

For batch processing scenarios, the _defer_eviction flag (lines 432–439) postpones eviction across multiple steps. This ensures that related frames are processed together before the cache trims older entries.

Special-Token Handling

During eviction, if kv_cache_cross_frame_special is active, the system moves evicted frames' special tokens to separate *_special buckets. This preserves long-range continuity without the memory cost of full tensor retention.

Inspection and Cleanup

The get_kv_cache_info() method (lines 493–501) reports cache statistics by iterating over the dictionary and counting entries whose keys start with k_ but do not end with _special, providing visibility into normal keys versus special tokens.

To free memory at episode boundaries or between benchmark runs, call clean_kv_cache(). This method appears throughout scripts/benchmark_gct_memory.py (lines 152–155, 210, 250, 318, 377, 412) and wipes all stored KV tensors from the aggregator and camera-head sub-modules.

Implementation Example

Configure and control the paged KV cache during inference:

from lingbot_map.models.gct_stream_window_v2 import GCTStreamWindowV2

# Initialize with a 64-frame window and 8 retained scale frames

model = GCTStreamWindowV2(
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    kv_cache_cross_frame_special=True,
    kv_cache_include_scale_frames=True,
    kv_cache_camera_only=False,
)

# Read from cache without appending new KV (dry-run mode)

model.set_skip_append(True)

# ... inference pass ...

model.set_skip_append(False)

# Defer eviction while processing a batch

model.set_defer_eviction(True)

# ... process multiple frames ...

model.set_defer_eviction(False)

# Clear all cached tensors

model.clean_kv_cache()

Summary

  • Memory Boundedness: The sliding-window mechanism ensures GPU memory usage remains constant regardless of total sequence length by evicting old frames via tensor slicing.
  • Long-Range Context: Cross-frame special tokens retain critical information from distant frames without storing complete KV tensors.
  • Flexible Control: Flags like _skip_append, _defer_eviction, and kv_cache_camera_only allow fine-tuning of the speed-memory-accuracy trade-off.
  • Explicit Management: Methods such as clean_kv_cache() and get_kv_cache_info() provide programmatic control over the cache lifecycle.

Frequently Asked Questions

How does the sliding-window eviction prevent GPU memory overflow?

The system enforces a hard limit via kv_cache_sliding_window, keeping only the most recent N frames in GPU memory. When new frames arrive, the manager slices the oldest tensors using kv[key][:, :, :-1] or sets them to None if they reach size 1 (lines 458–463), ensuring memory usage never grows linearly with sequence length.

What is the purpose of cross-frame special tokens in the paged KV cache?

Cross-frame special tokens act as a compressed summary of evicted frames. When kv_cache_cross_frame_special is enabled and a frame leaves the sliding window, its special tokens move to a *_special slot rather than being discarded. Later frames can attend to these tokens, maintaining long-range continuity without the memory cost of full KV tensors.

How can I temporarily disable KV cache writes during inference?

Call model.set_skip_append(True) to set the internal _skip_append flag (lines 414–420). In this mode, attention layers read from the concatenation of cached and current KV pairs but do not persist the new tensors to the cache. Use set_skip_append(False) to resume normal caching behavior.

Where is the eviction logic implemented in the source code?

The core eviction logic resides in lingbot_map/models/gct_stream_window_v2.py at lines 458–463. This section handles the tensor slicing operations that remove the oldest time-steps when the cache exceeds the configured sliding-window size.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →