# How LingBot-Map's Paged KV Cache Handles Long Video Sequences

> Discover how LingBot-Map's paged KV cache efficiently handles long video sequences preserving critical context within bounded GPU memory.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: internals
- Published: 2026-07-28

---

**LingBot-Map implements a paged KV cache that combines sliding-window eviction, scale-frame retention, and cross-frame special tokens to maintain bounded GPU memory while preserving critical context across thousands of video frames.**

LingBot-Map (available at Robbyant/lingbot-map) processes high-resolution 3-D mapping tasks from extended video sequences. To prevent GPU memory exhaustion during transformer inference, the project implements a sophisticated paged key-value cache mechanism primarily defined in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py).

## Core Architecture Components

### Sliding-Window Eviction

The cache enforces a strict memory bound through the **`kv_cache_sliding_window`** parameter (defined at lines 223–227). This setting establishes a fixed-size window that retains only the most recent *N* frames. When new frames arrive, the manager automatically discards the oldest entries by slicing the KV tensors.

### Scale-Frame Retention

Beyond the sliding window, the system preserves critical keyframes using **`kv_cache_scale_frames`** (lines 223–227). These scale frames provide global scene understanding and remain cached even if they fall outside the temporal window, provided **`kv_cache_include_scale_frames`** is enabled.

### Cross-Frame Special Tokens

When standard frames are evicted, the **`kv_cache_cross_frame_special`** mechanism (lines 223–227) copies their special tokens into dedicated `*_special` slots rather than discarding them entirely. This allows later frames to attend to summarized information from distant past frames without retaining the full KV tensors.

### Camera-Only Mode

For scenarios requiring minimal memory footprint, **`kv_cache_camera_only`** (lines 223–227) retains only camera-related tokens from evicted frames while discarding all other tensor data.

## Lifecycle of the KV Cache

### Construction and Initialization

When instantiating `GCTStreamWindow` or `GCTStreamWindowV2`, the system saves cache parameters as member variables including `self.kv_cache_sliding_window` and `self.kv_cache_scale_frames` around line 276. These values configure the low-level `KVCacheManager` that physically holds the tensors.

### Appending New KV Pairs

During each forward pass, the model produces new `k` and `v` tensors. The aggregator's `kv_cache_manager` appends these to the cache unless the **`_skip_append`** flag is set (handled at lines 414–420). When `skip=True`, the attention layer reads from `[cached_kv + current_kv]` but does not persist the current KV, enabling read-only passes.

### Eviction Logic

After appending, the manager checks cache size against the sliding-window limit. If exceeded, it removes the oldest frame(s) using the logic at lines 458–463:

```python

# Slicing operation to drop the oldest time-step

kv[key] = kv[key][:, :, :-1]

# If the tensor's third dimension reaches size 1, clear it entirely

if kv[key].shape[2] == 1:
    kv[key] = None

```

### Deferred Eviction

For batch processing scenarios, the **`_defer_eviction`** flag (lines 432–439) postpones eviction across multiple steps. This ensures that related frames are processed together before the cache trims older entries.

### Special-Token Handling

During eviction, if `kv_cache_cross_frame_special` is active, the system moves evicted frames' special tokens to separate `*_special` buckets. This preserves long-range continuity without the memory cost of full tensor retention.

### Inspection and Cleanup

The **`get_kv_cache_info()`** method (lines 493–501) reports cache statistics by iterating over the dictionary and counting entries whose keys start with `k_` but do not end with `_special`, providing visibility into normal keys versus special tokens.

To free memory at episode boundaries or between benchmark runs, call **`clean_kv_cache()`**. This method appears throughout [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py) (lines 152–155, 210, 250, 318, 377, 412) and wipes all stored KV tensors from the aggregator and camera-head sub-modules.

## Implementation Example

Configure and control the paged KV cache during inference:

```python
from lingbot_map.models.gct_stream_window_v2 import GCTStreamWindowV2

# Initialize with a 64-frame window and 8 retained scale frames

model = GCTStreamWindowV2(
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    kv_cache_cross_frame_special=True,
    kv_cache_include_scale_frames=True,
    kv_cache_camera_only=False,
)

# Read from cache without appending new KV (dry-run mode)

model.set_skip_append(True)

# ... inference pass ...

model.set_skip_append(False)

# Defer eviction while processing a batch

model.set_defer_eviction(True)

# ... process multiple frames ...

model.set_defer_eviction(False)

# Clear all cached tensors

model.clean_kv_cache()

```

## Summary

- **Memory Boundedness**: The sliding-window mechanism ensures GPU memory usage remains constant regardless of total sequence length by evicting old frames via tensor slicing.
- **Long-Range Context**: Cross-frame special tokens retain critical information from distant frames without storing complete KV tensors.
- **Flexible Control**: Flags like `_skip_append`, `_defer_eviction`, and `kv_cache_camera_only` allow fine-tuning of the speed-memory-accuracy trade-off.
- **Explicit Management**: Methods such as `clean_kv_cache()` and `get_kv_cache_info()` provide programmatic control over the cache lifecycle.

## Frequently Asked Questions

### How does the sliding-window eviction prevent GPU memory overflow?

The system enforces a hard limit via `kv_cache_sliding_window`, keeping only the most recent *N* frames in GPU memory. When new frames arrive, the manager slices the oldest tensors using `kv[key][:, :, :-1]` or sets them to `None` if they reach size 1 (lines 458–463), ensuring memory usage never grows linearly with sequence length.

### What is the purpose of cross-frame special tokens in the paged KV cache?

Cross-frame special tokens act as a compressed summary of evicted frames. When `kv_cache_cross_frame_special` is enabled and a frame leaves the sliding window, its special tokens move to a `*_special` slot rather than being discarded. Later frames can attend to these tokens, maintaining long-range continuity without the memory cost of full KV tensors.

### How can I temporarily disable KV cache writes during inference?

Call `model.set_skip_append(True)` to set the internal `_skip_append` flag (lines 414–420). In this mode, attention layers read from the concatenation of cached and current KV pairs but do not persist the new tensors to the cache. Use `set_skip_append(False)` to resume normal caching behavior.

### Where is the eviction logic implemented in the source code?

The core eviction logic resides in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) at lines 458–463. This section handles the tensor slicing operations that remove the oldest time-steps when the cache exceeds the configured sliding-window size.