How the Paged KV Cache Works in LingBot-Map's Aggregator
LingBot-Map implements a two-stream paged KV cache using FlashInferKVCacheManager to separate recyclable frame patches from persistent special tokens, enabling constant-memory streaming inference across thousands of frames.
The LingBot-Map repository employs a sophisticated paged key-value cache mechanism to handle long-range streaming inference efficiently. Unlike conventional KV caches that grow linearly with sequence length, this implementation uses a fixed-size page pool with intelligent eviction policies to maintain real-time performance across extended video streams. The core logic resides in FlashInferKVCacheManager and integrates directly with the AggregatorStream class to manage attention state.
Two-Stream Paged Architecture
The cache splits physical storage between patch and special streams that share a single page pool per transformer block, as defined in the header of flashinfer_cache.py [lines 1-13].
Patch Stream: Recyclable Frame Storage
The patch stream stores token embeddings for individual video frames in recyclable pages. Each frame occupies one page within a sliding window, plus a fixed number of scale pages that remain permanently cached. When the sliding window limit is exceeded, the oldest non-scale pages return to the free list for immediate reuse.
Special Stream: Persistent Token Buffer
The special stream maintains camera, register, and scale tokens in an append-only structure. These pages are never evicted and grow monotonically as new frames enter the system. During attention computation, the visible set includes all special pages combined with active patch pages, allowing every frame to attend to global special tokens without additional masking overhead [flashinfer_cache.py lines 24-28].
Initialization and Configuration
When AggregatorStream initializes, it selects between SDPA and FlashInfer backends based on the use_sdpa parameter. If FlashInfer is selected (the default), the KV cache instantiates lazily during the first forward pass [stream.py lines 79-86].
The aggregator passes cache configuration parameters to each transformer block during construction:
GlobalBlockCls = SDPABlock if self.use_sdpa else FlashInferBlock
self.global_blocks = nn.ModuleList([
GlobalBlockCls(
...,
kv_cache_sliding_window=self.kv_cache_sliding_window,
kv_cache_scale_frames=self.kv_cache_scale_frames,
kv_cache_cross_frame_special=self.kv_cache_cross_frame_special,
kv_cache_include_scale_frames=self.kv_cache_include_scale_frames,
kv_cache_camera_only=self.kv_cache_camera_only,
)
for _ in range(depth)
])
These parameters control the sliding window size, permanent scale frame count, and special token retention policies [stream.py lines 140-147].
FlashInferKVCacheManager Internals
Page Pool Allocation
The manager pre-allocates physical storage using a fixed-size tensor of shape [max_num_pages, 2, page_size, H, D], where the first max_patch_pages slots serve the recyclable patch stream and remaining slots accommodate the append-only special stream [flashinfer_cache.py lines 41-49]. The allocation strategy reserves space for scale frames, sliding window frames, and headroom for the patch stream, while sizing the special pool based on maximum expected frame count [flashinfer_cache.py lines 30-38].
Eviction and Sliding Window Management
The manager tracks active pages using scale_patch_pages and live_window_patch_pages deques. When total_frames_processed exceeds kv_cache_sliding_window + kv_cache_scale_frames, the evict_frames method removes the oldest patch page from the live window and returns its ID to the free list [flashinfer_cache.py lines 52-60]. This mechanism ensures constant-time memory growth regardless of input length.
Deferred Eviction for Keyframe Selection
For flow-based keyframe selection workflows, the cache supports deferred eviction via the _defer_eviction flag. When enabled, eviction decisions postpone until the caller explicitly commits or rolls back the last frame, facilitating speculative processing without destroying context prematurely [flashinfer_cache.py lines 71-76].
Runtime Operation Flow
The paged KV cache operates through a four-stage pipeline during streaming inference:
-
Manager Creation: On the first frame,
AggregatorStreamcalls_get_flashinfer_managerto instantiateFlashInferKVCacheManagerif not already present. -
Frame Appending: The
append_framemethod writes patch KV tensors to assigned pages and appends special tokens to the special stream. The manager allocates from the free list or recycles evicted pages automatically. -
Eviction Trigger: When live frames exceed the configured window plus scale frames,
evict_framesremoves obsolete patch pages while preserving special tokens and scale frames. -
Attention Computation: Each
FlashInferBlockcallscompute_attention, which constructs indices for visible pages (scale + live window + special) and executes FlashInfer's batched prefilling kernel [attention.py].
Code Examples
Creating a streaming aggregator with custom cache settings:
from lingbot_map.aggregator.stream import AggregatorStream
agg = AggregatorStream(
sliding_window_size=128, # causal sliding window (blocks)
num_frame_for_scale=8, # keep 8 scale frames permanently
kv_cache_sliding_window=64, # KV cache eviction window
kv_cache_scale_frames=8,
kv_cache_cross_frame_special=True,
kv_cache_include_scale_frames=True,
use_sdpa=False, # use FlashInfer (paged KV)
)
Inspecting cache statistics during inference:
# After processing several frames
info = agg.get_kv_cache_info()
print(f"Cached frames: {info['num_cached']}")
print(f"Total tokens per frame: {info['tokens_per_frame']}")
This method reports stored K/V slots by querying the internal cache dict or FlashInfer manager [gct_stream_window_v2.py lines 484-499].
Resetting the cache between videos:
# Reset the KV cache (e.g., when starting a new video)
agg.clean_kv_cache()
This delegates to both the aggregator and camera head to free all pages [gct_stream_window_v2.py lines 388-401].
Summary
- Two-stream design: Separates recyclable patch pages from append-only special token pages to optimize memory usage.
- Constant memory footprint: Sliding window eviction ensures the cache never exceeds pre-allocated page limits, enabling processing of 10,000+ frame sequences.
- FlashInfer integration:
FlashInferKVCacheManagerinflashinfer_cache.pyprovides the core paging logic, whileAggregatorStreaminstream.pyorchestrates initialization and configuration. - Flexible eviction: Supports immediate and deferred eviction modes to accommodate different keyframe selection strategies.
Frequently Asked Questions
What is the difference between the patch and special streams in LingBot-Map's KV cache?
The patch stream stores frame-specific token embeddings in recyclable pages subject to sliding window eviction, while the special stream holds camera, register, and scale tokens in permanent, append-only pages. This separation allows the model to maintain access to global special tokens while cycling through frame patches efficiently, as implemented in flashinfer_cache.py [lines 41-49].
How does the sliding window eviction policy work in the paged KV cache?
The manager maintains deques tracking scale frames (permanent) and live window frames (temporary). When the total frame count exceeds kv_cache_sliding_window + kv_cache_scale_frames, the oldest non-scale patch page returns to the free list. This policy ensures constant memory usage while preserving recent context and essential scale information [flashinfer_cache.py lines 52-60].
When should I use deferred eviction mode in LingBot-Map?
Enable deferred eviction when implementing flow-based keyframe selection algorithms that require speculative evaluation of frames before commitment. Setting _defer_eviction = True prevents immediate context destruction, allowing the application to decide whether to commit or roll back the last frame's cache entries [flashinfer_cache.py lines 71-76].
What backend alternatives exist to FlashInfer for the KV cache?
LingBot-Map provides an SDPA (Scaled Dot-Product Attention) backend that uses a simple dictionary-based cache instead of the paged FlashInfer implementation. Set use_sdpa=True when initializing AggregatorStream to use this alternative, though it lacks the constant-memory guarantees of the paged approach [stream.py lines 79-86].
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →