# How LingBot-Map Uses KV Caching for Efficient Streaming Inference

> Discover how LingBot-Map leverages paged KV-cache with FlashInfer for efficient streaming inference, achieving real-time performance by avoiding full sequence recomputation.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: getting-started
- Published: 2026-07-23

---

**LingBot-Map implements a paged KV-cache using FlashInfer to store key-value pairs from previously processed video frames, enabling linear-time attention complexity and real-time streaming inference without recomputing the full sequence history.**

LingBot-Map is an open-source video understanding framework that performs causal, frame-by-frame inference on long video streams. To achieve efficient streaming inference without the quadratic cost of full self-attention over entire video histories, the repository leverages a sophisticated **KV caching** architecture built on paged memory management. This system, implemented in `Robbyant/lingbot-map`, maintains cached key-value tensors from historical frames while evicting outdated patches according to a sliding-window policy.

## The Paged KV-Cache Architecture

At the core of LingBot-Map’s streaming capability lies a two-stream paged KV-cache managed by the `FlashInferKVCacheManager` class in [[`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py)](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py). This manager allocates GPU memory in pages and tracks which tokens remain visible for attention computation during each forward pass.

### FlashInfer Backend and Cache Manager

The `FlashInferKVCacheManager` class handles all cache operations, including allocation, appending, and eviction. It wraps FlashInfer’s fast CUDA kernels for paged attention, providing **O(1) append time** and **O(n) attention complexity** where *n* is the window size rather than the full sequence length.

Key methods include:

- **`append_frame`**: Writes a new frame’s K/V tensors into the paged structure, placing patch tokens first followed by special tokens.
- **`evict_frames`**: Removes aged window pages while preserving scale pages, optimizing memory usage.
- **`execute_deferred_eviction`** and **`rollback_last_frame`**: Support temporary cache states for flow-based keyframe selection, allowing the system to discard recently added KV pairs if a frame is rejected.

The high-level `AggregatorStream` class in [[`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py)](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) instantiates this manager and exposes cache control methods like `clean_kv_cache` and `_set_skip_append` to the transformer blocks.

### Two-Stream Design: Patch and Special Tokens

The cache separates tokens into two logical streams to optimize memory and retention policies:

1.  **Patch Stream (Recyclable)**: Contains actual image patch embeddings. These are organized into patch pages where the first `kv_cache_scale_frames` pages remain resident permanently (scale tokens), while subsequent pages enter a sliding window subject to eviction.
2.  **Special Stream (Append-Only)**: Holds six fixed special tokens (camera, register, and scale identifiers) per frame. These occupy minimal space and are never evicted, ensuring persistent metadata context across the entire video.

During each forward pass, the manager constructs a **visible-page table** comprising scale pages plus live window pages plus all special pages. FlashInfer then computes attention over this concatenated view in a single batched operation.

## Sliding-Window Eviction Strategy

After processing each frame via `append_frame`, the cache manager invokes `evict_frames` to enforce the configured memory budget. The eviction logic operates as follows:

- **Scale Protection**: Pages belonging to the initial `kv_cache_scale_frames` are never evicted, ensuring bidirectional context remains available for scale-token processing.
- **Window Management**: Only patch pages beyond the scale region are eligible for eviction. When the live window exceeds `kv_cache_sliding_window`, the oldest page is popped from `live_window_patch_pages` and returned to the free list.
- **Special Token Retention**: Special tokens from evicted frames may be retained when `kv_cache_cross_frame_special` is enabled, supporting cross-frame attention mechanisms without holding full patch data.

This approach yields **linear memory growth** with respect to the sliding window size rather than the video duration, enabling processing of arbitrarily long videos on limited GPU memory.

## Flow-Based Keyframe Selection

The streaming inference pipeline supports intelligent keyframe selection through deferred eviction mechanisms. When `inference_streaming` detects motion via optical flow thresholds, it temporarily suspends automatic eviction by calling `_set_defer_eviction(True)`.

If the flow magnitude exceeds the threshold, indicating significant visual change, the frame is committed as a keyframe via `_execute_deferred_eviction`. Otherwise, the system invokes `_rollback_last_frame` to discard the provisional KV tensors, effectively reverting the cache state. This selective retention optimizes memory for static scenes while preserving detail during motion, implemented in [`GCTStream.inference_streaming`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py).

## Configuration Parameters

The KV-cache behavior is controlled through parameters passed to `GCTStream.__init__` in [[`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py)](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py):

| Parameter | Description | Default |
|-----------|-------------|---------|
| **`kv_cache_sliding_window`** | Number of recent patch frames retained in the sliding window. | `64` |
| **`kv_cache_scale_frames`** | Scale frames kept permanently resident; never evicted. | `8` |
| **`kv_cache_cross_frame_special`** | Retain special tokens from evicted frames for cross-frame attention. | `True` |
| **`kv_cache_include_scale_frames`** | Include scale frames in the cache (disable only for specific ablations). | `True` |
| **`kv_cache_camera_only`** | Restrict caching to camera-specific tokens only (experimental). | `False` |

Adjusting `kv_cache_sliding_window` directly trades off between contextual memory (larger window) and GPU memory consumption (smaller window).

## Implementation Examples

### Instantiating a Model with Custom Cache Configuration

```python
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    kv_cache_sliding_window=128,   # Larger window for more context

    kv_cache_scale_frames=4,       # Fewer permanent scale frames

    kv_cache_cross_frame_special=False,
    enable_3d_rope=True,
)
model.eval()

```

### Running Streaming Inference with Automatic Cache Management

```python

# Input shape: [Batch, Sequence, Channels, Height, Width]

imgs = torch.randn(1, 120, 3, 518, 518)  # 120-frame video clip

# Process with fixed keyframe interval

predictions = model.inference_streaming(
    images=imgs,
    keyframe_interval=5,
    output_device=torch.device('cpu'),  # Offload to prevent OOM

)

```

Internally, this clears the cache (`clean_kv_cache`), appends KV tensors per frame, evicts old pages according to the sliding window, and handles deferred eviction for flow-based keyframes.

### Manual Cache Control for Advanced Use Cases

```python

# Reset cache when switching to a new video sequence

model.clean_kv_cache()

# Process a frame without storing its KV (non-keyframe processing)

model._set_skip_append(True)
frame_output = model.forward(frame_tensor)  # shape: [1, 1, 3, H, W]

model._set_skip_append(False)  # Re-enable normal caching

```

### Inspecting Cache Statistics

```python
stats = model.aggregator.kv_cache_manager.get_cache_stats(block_idx=0)
print("Cache utilization:", stats)

# Output: {'frame_count': 42, 'scale_pages': 8, 'live_pages': 34, 

#          'free_pages': 46, 'special_tokens': 252}

```

Set the environment variable `LINGBOT_DEBUG_KV=1` to enable per-frame diagnostic logging via `_log_kv_stats` during inference.

## Summary

- **LingBot-Map** utilizes a **paged KV-cache** via `FlashInferKVCacheManager` in [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py) to achieve efficient streaming inference on long videos.
- The architecture uses a **two-stream design** separating recyclable patch tokens from append-only special tokens, with scale frames permanently resident.
- **Sliding-window eviction** limits memory usage to O(window size) rather than O(video length), while flow-based keyframe selection optimizes which frames enter the cache.
- Configuration parameters like `kv_cache_sliding_window` and `kv_cache_scale_frames` allow precise tuning of the memory-context tradeoff.
- All cache operations integrate seamlessly with the `GCTStream` high-level API in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py).

## Frequently Asked Questions

### What is KV caching and why is it essential for LingBot-Map’s streaming inference?

KV caching stores the key and value tensors computed during self-attention for reuse in subsequent forward passes. In LingBot-Map, this avoids recomputing attention over all previous video frames when processing a new frame, reducing complexity from quadratic to linear with respect to the sequence length. Without KV caching, real-time inference on long videos would be computationally prohibitive.

### How does the two-stream patch and special token design optimize memory usage?

The separation allows the cache to apply different retention policies to different token types. Image patch tokens are numerous and subject to sliding-window eviction, while special tokens (camera, register, scale) are few and kept permanently. This prevents the cache from being dominated by metadata tokens while ensuring critical structural information remains accessible throughout the video.

### What distinguishes scale frames from window frames in the KV cache?

Scale frames represent the initial `kv_cache_scale_frames` (default 8) frames containing scale tokens; their KV pairs are **never evicted** and remain permanently in GPU memory. Window frames comprise all subsequent patch tokens and are managed in a FIFO sliding window of size `kv_cache_sliding_window`. When the window fills, the oldest frame is evicted to free memory, whereas scale frames provide fixed anchor points for attention.

### How does flow-based keyframe selection interact with the cache eviction mechanism?

When optical flow detection is enabled, the cache enters a deferred eviction mode where newly appended KV pairs are marked provisional. If the computed flow exceeds the threshold, the frame is accepted and eviction proceeds normally. If flow is below threshold (indicating static content), `rollback_last_frame` removes the provisional KV tensors, effectively preventing low-information frames from consuming cache capacity while maintaining temporal continuity in the sliding window.