How LingBot-Map Uses KV Caching for Efficient Streaming Inference
LingBot-Map implements a paged KV-cache using FlashInfer to store key-value pairs from previously processed video frames, enabling linear-time attention complexity and real-time streaming inference without recomputing the full sequence history.
LingBot-Map is an open-source video understanding framework that performs causal, frame-by-frame inference on long video streams. To achieve efficient streaming inference without the quadratic cost of full self-attention over entire video histories, the repository leverages a sophisticated KV caching architecture built on paged memory management. This system, implemented in Robbyant/lingbot-map, maintains cached key-value tensors from historical frames while evicting outdated patches according to a sliding-window policy.
The Paged KV-Cache Architecture
At the core of LingBot-Map’s streaming capability lies a two-stream paged KV-cache managed by the FlashInferKVCacheManager class in [lingbot_map/layers/flashinfer_cache.py](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py). This manager allocates GPU memory in pages and tracks which tokens remain visible for attention computation during each forward pass.
FlashInfer Backend and Cache Manager
The FlashInferKVCacheManager class handles all cache operations, including allocation, appending, and eviction. It wraps FlashInfer’s fast CUDA kernels for paged attention, providing O(1) append time and O(n) attention complexity where n is the window size rather than the full sequence length.
Key methods include:
append_frame: Writes a new frame’s K/V tensors into the paged structure, placing patch tokens first followed by special tokens.evict_frames: Removes aged window pages while preserving scale pages, optimizing memory usage.execute_deferred_evictionandrollback_last_frame: Support temporary cache states for flow-based keyframe selection, allowing the system to discard recently added KV pairs if a frame is rejected.
The high-level AggregatorStream class in [lingbot_map/aggregator/stream.py](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) instantiates this manager and exposes cache control methods like clean_kv_cache and _set_skip_append to the transformer blocks.
Two-Stream Design: Patch and Special Tokens
The cache separates tokens into two logical streams to optimize memory and retention policies:
- Patch Stream (Recyclable): Contains actual image patch embeddings. These are organized into patch pages where the first
kv_cache_scale_framespages remain resident permanently (scale tokens), while subsequent pages enter a sliding window subject to eviction. - Special Stream (Append-Only): Holds six fixed special tokens (camera, register, and scale identifiers) per frame. These occupy minimal space and are never evicted, ensuring persistent metadata context across the entire video.
During each forward pass, the manager constructs a visible-page table comprising scale pages plus live window pages plus all special pages. FlashInfer then computes attention over this concatenated view in a single batched operation.
Sliding-Window Eviction Strategy
After processing each frame via append_frame, the cache manager invokes evict_frames to enforce the configured memory budget. The eviction logic operates as follows:
- Scale Protection: Pages belonging to the initial
kv_cache_scale_framesare never evicted, ensuring bidirectional context remains available for scale-token processing. - Window Management: Only patch pages beyond the scale region are eligible for eviction. When the live window exceeds
kv_cache_sliding_window, the oldest page is popped fromlive_window_patch_pagesand returned to the free list. - Special Token Retention: Special tokens from evicted frames may be retained when
kv_cache_cross_frame_specialis enabled, supporting cross-frame attention mechanisms without holding full patch data.
This approach yields linear memory growth with respect to the sliding window size rather than the video duration, enabling processing of arbitrarily long videos on limited GPU memory.
Flow-Based Keyframe Selection
The streaming inference pipeline supports intelligent keyframe selection through deferred eviction mechanisms. When inference_streaming detects motion via optical flow thresholds, it temporarily suspends automatic eviction by calling _set_defer_eviction(True).
If the flow magnitude exceeds the threshold, indicating significant visual change, the frame is committed as a keyframe via _execute_deferred_eviction. Otherwise, the system invokes _rollback_last_frame to discard the provisional KV tensors, effectively reverting the cache state. This selective retention optimizes memory for static scenes while preserving detail during motion, implemented in GCTStream.inference_streaming.
Configuration Parameters
The KV-cache behavior is controlled through parameters passed to GCTStream.__init__ in [lingbot_map/models/gct_stream_window_v2.py](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py):
| Parameter | Description | Default |
|---|---|---|
kv_cache_sliding_window |
Number of recent patch frames retained in the sliding window. | 64 |
kv_cache_scale_frames |
Scale frames kept permanently resident; never evicted. | 8 |
kv_cache_cross_frame_special |
Retain special tokens from evicted frames for cross-frame attention. | True |
kv_cache_include_scale_frames |
Include scale frames in the cache (disable only for specific ablations). | True |
kv_cache_camera_only |
Restrict caching to camera-specific tokens only (experimental). | False |
Adjusting kv_cache_sliding_window directly trades off between contextual memory (larger window) and GPU memory consumption (smaller window).
Implementation Examples
Instantiating a Model with Custom Cache Configuration
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream
model = GCTStream(
img_size=518,
patch_size=14,
embed_dim=1024,
kv_cache_sliding_window=128, # Larger window for more context
kv_cache_scale_frames=4, # Fewer permanent scale frames
kv_cache_cross_frame_special=False,
enable_3d_rope=True,
)
model.eval()
Running Streaming Inference with Automatic Cache Management
# Input shape: [Batch, Sequence, Channels, Height, Width]
imgs = torch.randn(1, 120, 3, 518, 518) # 120-frame video clip
# Process with fixed keyframe interval
predictions = model.inference_streaming(
images=imgs,
keyframe_interval=5,
output_device=torch.device('cpu'), # Offload to prevent OOM
)
Internally, this clears the cache (clean_kv_cache), appends KV tensors per frame, evicts old pages according to the sliding window, and handles deferred eviction for flow-based keyframes.
Manual Cache Control for Advanced Use Cases
# Reset cache when switching to a new video sequence
model.clean_kv_cache()
# Process a frame without storing its KV (non-keyframe processing)
model._set_skip_append(True)
frame_output = model.forward(frame_tensor) # shape: [1, 1, 3, H, W]
model._set_skip_append(False) # Re-enable normal caching
Inspecting Cache Statistics
stats = model.aggregator.kv_cache_manager.get_cache_stats(block_idx=0)
print("Cache utilization:", stats)
# Output: {'frame_count': 42, 'scale_pages': 8, 'live_pages': 34,
# 'free_pages': 46, 'special_tokens': 252}
Set the environment variable LINGBOT_DEBUG_KV=1 to enable per-frame diagnostic logging via _log_kv_stats during inference.
Summary
- LingBot-Map utilizes a paged KV-cache via
FlashInferKVCacheManagerinlingbot_map/layers/flashinfer_cache.pyto achieve efficient streaming inference on long videos. - The architecture uses a two-stream design separating recyclable patch tokens from append-only special tokens, with scale frames permanently resident.
- Sliding-window eviction limits memory usage to O(window size) rather than O(video length), while flow-based keyframe selection optimizes which frames enter the cache.
- Configuration parameters like
kv_cache_sliding_windowandkv_cache_scale_framesallow precise tuning of the memory-context tradeoff. - All cache operations integrate seamlessly with the
GCTStreamhigh-level API inlingbot_map/models/gct_stream_window_v2.py.
Frequently Asked Questions
What is KV caching and why is it essential for LingBot-Map’s streaming inference?
KV caching stores the key and value tensors computed during self-attention for reuse in subsequent forward passes. In LingBot-Map, this avoids recomputing attention over all previous video frames when processing a new frame, reducing complexity from quadratic to linear with respect to the sequence length. Without KV caching, real-time inference on long videos would be computationally prohibitive.
How does the two-stream patch and special token design optimize memory usage?
The separation allows the cache to apply different retention policies to different token types. Image patch tokens are numerous and subject to sliding-window eviction, while special tokens (camera, register, scale) are few and kept permanently. This prevents the cache from being dominated by metadata tokens while ensuring critical structural information remains accessible throughout the video.
What distinguishes scale frames from window frames in the KV cache?
Scale frames represent the initial kv_cache_scale_frames (default 8) frames containing scale tokens; their KV pairs are never evicted and remain permanently in GPU memory. Window frames comprise all subsequent patch tokens and are managed in a FIFO sliding window of size kv_cache_sliding_window. When the window fills, the oldest frame is evicted to free memory, whereas scale frames provide fixed anchor points for attention.
How does flow-based keyframe selection interact with the cache eviction mechanism?
When optical flow detection is enabled, the cache enters a deferred eviction mode where newly appended KV pairs are marked provisional. If the computed flow exceeds the threshold, the frame is accepted and eviction proceeds normally. If flow is below threshold (indicating static content), rollback_last_frame removes the provisional KV tensors, effectively preventing low-information frames from consuming cache capacity while maintaining temporal continuity in the sliding window.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →