# How the Sliding Window Eviction Policy Works in Ling‑Bot’s KV Cache Manager

> Learn how the sliding window eviction policy in LingBot's KV Cache Manager bounds memory usage during streaming inference by retaining recent frames and discarding older entries.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: internals
- Published: 2026-07-25

---

**The sliding window eviction policy bounds memory usage during streaming inference by retaining only a configurable number of recent frames plus optional "scale" frames, automatically discarding older entries through tensor slicing in the `_apply_kv_cache_eviction` method.**

Ling‑Bot caches key (K) and value (V) tensors from transformer blocks to avoid recomputing attention for previous tokens during long streaming sessions. Without eviction, this KV cache grows unbounded, eventually exhausting GPU memory. The repository implements a **sliding window eviction policy** that enforces a hard limit on cached frames while preserving critical temporal context.

## Core Configuration Parameters

The eviction behavior is governed by hyperparameters passed to the attention layer during model construction:

- **`kv_cache_sliding_window`** – The number of most recent frames to retain in the cache.
- **`kv_cache_scale_frames`** – The number of oldest frames to preserve as "scale" frames for global context.
- **`kv_cache_include_scale_frames`** – Boolean flag determining whether to concatenate scale frames with the recent window or discard them.
- **`kv_cache_cross_frame_special`** – Enables preservation of a subset of evicted tokens in a separate special cache for cross-frame attention.

## The `_apply_kv_cache_eviction` Method

The eviction logic lives in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) inside the private method `_apply_kv_cache_eviction`. This method inspects the cached tensor dimensions and conditionally truncates history when the total frame count exceeds the configured limits.

### Step‑by‑Step Eviction Process

The method executes the following algorithmic steps:

1. **Determine Limits** – The policy reads `sliding_window_frames = self.kv_cache_sliding_window` and `scale_frames = self.kv_cache_scale_frames` to establish retention boundaries.

2. **Check Eviction Necessity** – Eviction proceeds only if the cached K tensor has more than one token (`shape[3] > 1`) and the total cached frame count exceeds `sliding_window_frames + scale_frames`.

3. **Compute Slice Bounds** – The algorithm calculates the indices for the evicted region:
   ```python
   evict_start = scale_frames
   evict_end = num_cached_frames - sliding_window_frames
   ```

4. **Extract Evicted Chunks** – The K and V tensors for frames falling outside the retention window are sliced out:
   ```python
   evicted_k = kv_cache[f"k_{global_idx}"][:, :, evict_start:evict_end, :, :]
   evicted_v = kv_cache[f"v_{global_idx}"][:, :, evict_start:evict_end, :, :]
   ```

5. **Preserve Special Tokens** – If `kv_cache_cross_frame_special` is enabled, camera-only or camera-plus-scale tokens from the evicted chunk are copied into a separate special KV cache (`*_special`). This maintains cross-frame attention capabilities on selected historical tokens.

6. **Re‑assemble Retained Cache** – The method constructs the truncated cache using `torch.cat` with two possible branches:
   - **Include scale frames**: Concatenate the first `scale_frames` with the last `sliding_window_frames`.
   - **Exclude scale frames**: Keep only the last `sliding_window_frames`.

7. **Update Dictionaries** – The truncated tensors replace the previous entries in `kv_cache` at keys `f"k_{global_idx}"` and `f"v_{global_idx}"`.

## Practical Configuration Example

Configure the sliding window policy when instantiating a streaming GCT model:

```python
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    # ... other arguments ...

    kv_cache_sliding_window=8,      # Retain the last 8 keyframes

    kv_cache_scale_frames=2,        # Retain the first 2 scale frames

    kv_cache_include_scale_frames=True,
    kv_cache_cross_frame_special=True,
    kv_cache_camera_only=False,
)

```

During inference, after processing frame *N*, the cache automatically holds frames 0‑1 (scale) and frames *N‑7* through *N* (recent), with all intermediate frames removed. This guarantees that memory usage remains constant regardless of stream duration.

## Summary

- The **sliding window eviction policy** prevents unbounded memory growth by enforcing a fixed cache size of `kv_cache_sliding_window + kv_cache_scale_frames` frames.
- Eviction occurs in `_apply_kv_cache_eviction` within [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) through precise tensor slicing and concatenation operations.
- **Scale frames** provide stable global context by preserving the oldest frames, while the **sliding window** retains the most recent temporal context.
- **Special token preservation** allows the model to maintain cross-frame attention on critical evicted tokens via a secondary cache.
- The policy activates automatically during streaming when cache thresholds are exceeded, requiring no manual intervention.

## Frequently Asked Questions

### What triggers the sliding window eviction in Ling‑Bot?

Eviction triggers when the cached K tensor contains more than one token and the total number of cached frames exceeds the sum of `kv_cache_sliding_window` and `kv_cache_scale_frames`. This conditional check ensures the policy does not activate on single-token sequences or when the cache remains within configured bounds.

### How does the policy differentiate between scale frames and sliding window frames?

The policy treats `kv_cache_scale_frames` as the oldest frames to preserve (indices 0 to `scale_frames-1`) and `kv_cache_sliding_window` as the most recent frames to preserve (indices `num_cached_frames - sliding_window_frames` to end). The region between these two groups is evicted. If `kv_cache_include_scale_frames` is False, the scale frames are not retained in the main cache.

### Can evicted tokens still participate in attention calculations?

Yes. When `kv_cache_cross_frame_special` is enabled, the system extracts a subset of tokens from the evicted region (typically camera-only or camera-plus-scale) and stores them in a separate special KV cache. This allows the model to attend to selected historical tokens even after they leave the primary sliding window, enabling long-range dependencies without full cache retention.

### Where is the eviction logic implemented in the codebase?

The core eviction algorithm resides in the `_apply_kv_cache_eviction` method of [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py). The streaming model implementations in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py), and [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py) instantiate the attention layers and forward the KV cache configuration parameters to enable this behavior.