How the Sliding Window Eviction Policy Works in Ling‑Bot’s KV Cache Manager

The sliding window eviction policy bounds memory usage during streaming inference by retaining only a configurable number of recent frames plus optional "scale" frames, automatically discarding older entries through tensor slicing in the _apply_kv_cache_eviction method.

Ling‑Bot caches key (K) and value (V) tensors from transformer blocks to avoid recomputing attention for previous tokens during long streaming sessions. Without eviction, this KV cache grows unbounded, eventually exhausting GPU memory. The repository implements a sliding window eviction policy that enforces a hard limit on cached frames while preserving critical temporal context.

Core Configuration Parameters

The eviction behavior is governed by hyperparameters passed to the attention layer during model construction:

  • kv_cache_sliding_window – The number of most recent frames to retain in the cache.
  • kv_cache_scale_frames – The number of oldest frames to preserve as "scale" frames for global context.
  • kv_cache_include_scale_frames – Boolean flag determining whether to concatenate scale frames with the recent window or discard them.
  • kv_cache_cross_frame_special – Enables preservation of a subset of evicted tokens in a separate special cache for cross-frame attention.

The _apply_kv_cache_eviction Method

The eviction logic lives in lingbot_map/layers/attention.py inside the private method _apply_kv_cache_eviction. This method inspects the cached tensor dimensions and conditionally truncates history when the total frame count exceeds the configured limits.

Step‑by‑Step Eviction Process

The method executes the following algorithmic steps:

  1. Determine Limits – The policy reads sliding_window_frames = self.kv_cache_sliding_window and scale_frames = self.kv_cache_scale_frames to establish retention boundaries.

  2. Check Eviction Necessity – Eviction proceeds only if the cached K tensor has more than one token (shape[3] > 1) and the total cached frame count exceeds sliding_window_frames + scale_frames.

  3. Compute Slice Bounds – The algorithm calculates the indices for the evicted region:

    evict_start = scale_frames
    evict_end = num_cached_frames - sliding_window_frames
  4. Extract Evicted Chunks – The K and V tensors for frames falling outside the retention window are sliced out:

    evicted_k = kv_cache[f"k_{global_idx}"][:, :, evict_start:evict_end, :, :]
    evicted_v = kv_cache[f"v_{global_idx}"][:, :, evict_start:evict_end, :, :]
  5. Preserve Special Tokens – If kv_cache_cross_frame_special is enabled, camera-only or camera-plus-scale tokens from the evicted chunk are copied into a separate special KV cache (*_special). This maintains cross-frame attention capabilities on selected historical tokens.

  6. Re‑assemble Retained Cache – The method constructs the truncated cache using torch.cat with two possible branches:

    • Include scale frames: Concatenate the first scale_frames with the last sliding_window_frames.
    • Exclude scale frames: Keep only the last sliding_window_frames.
  7. Update Dictionaries – The truncated tensors replace the previous entries in kv_cache at keys f"k_{global_idx}" and f"v_{global_idx}".

Practical Configuration Example

Configure the sliding window policy when instantiating a streaming GCT model:

from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    # ... other arguments ...

    kv_cache_sliding_window=8,      # Retain the last 8 keyframes

    kv_cache_scale_frames=2,        # Retain the first 2 scale frames

    kv_cache_include_scale_frames=True,
    kv_cache_cross_frame_special=True,
    kv_cache_camera_only=False,
)

During inference, after processing frame N, the cache automatically holds frames 0‑1 (scale) and frames N‑7 through N (recent), with all intermediate frames removed. This guarantees that memory usage remains constant regardless of stream duration.

Summary

  • The sliding window eviction policy prevents unbounded memory growth by enforcing a fixed cache size of kv_cache_sliding_window + kv_cache_scale_frames frames.
  • Eviction occurs in _apply_kv_cache_eviction within lingbot_map/layers/attention.py through precise tensor slicing and concatenation operations.
  • Scale frames provide stable global context by preserving the oldest frames, while the sliding window retains the most recent temporal context.
  • Special token preservation allows the model to maintain cross-frame attention on critical evicted tokens via a secondary cache.
  • The policy activates automatically during streaming when cache thresholds are exceeded, requiring no manual intervention.

Frequently Asked Questions

What triggers the sliding window eviction in Ling‑Bot?

Eviction triggers when the cached K tensor contains more than one token and the total number of cached frames exceeds the sum of kv_cache_sliding_window and kv_cache_scale_frames. This conditional check ensures the policy does not activate on single-token sequences or when the cache remains within configured bounds.

How does the policy differentiate between scale frames and sliding window frames?

The policy treats kv_cache_scale_frames as the oldest frames to preserve (indices 0 to scale_frames-1) and kv_cache_sliding_window as the most recent frames to preserve (indices num_cached_frames - sliding_window_frames to end). The region between these two groups is evicted. If kv_cache_include_scale_frames is False, the scale frames are not retained in the main cache.

Can evicted tokens still participate in attention calculations?

Yes. When kv_cache_cross_frame_special is enabled, the system extracts a subset of tokens from the evicted region (typically camera-only or camera-plus-scale) and stores them in a separate special KV cache. This allows the model to attend to selected historical tokens even after they leave the primary sliding window, enabling long-range dependencies without full cache retention.

Where is the eviction logic implemented in the codebase?

The core eviction algorithm resides in the _apply_kv_cache_eviction method of lingbot_map/layers/attention.py. The streaming model implementations in lingbot_map/models/gct_stream.py, gct_stream_window.py, and gct_stream_window_v2.py instantiate the attention layers and forward the KV cache configuration parameters to enable this behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →