Sliding Window Attention in LingBot-Map: Efficient Streaming Video Processing with Constant Memory
Sliding window attention in LingBot-Map is a causal self-attention mechanism that restricts each video frame to attend only to the most recent N frames, enabling unbounded sequence processing with constant memory usage through aggressive KV-cache eviction.
LingBot-Map processes arbitrarily long video streams using causal self-attention while keeping compute and memory bounded. The framework implements this through a configurable masking strategy and cache eviction policy defined in the lingbot_map/layers/attention.py source file, allowing the model to handle live video feeds without unbounded memory growth.
How Sliding Window Attention Works in LingBot-Map
Sliding-Window Mask Construction
In CausalAttention.forward, a binary mask is constructed that limits each query frame to attend only to the most recent frames within the configured window. The window size is calculated as sliding_window_size × num_frame_per_block, representing the number of historical frames retained for attention computation.
The implementation (lines 73-86 in lingbot_map/layers/attention.py) builds this mask iteratively for each frame position:
# lingbot_map/layers/attention.py, lines 73-86
if sliding_window_size > 0 and frame_seqlen is not None:
for i in range(num_frames):
window_size_in_frames = sliding_window_size * num_frame_per_block
window_start_frame = max(0, i - window_size_in_frames + 1)
k_start = window_start_frame * frame_seqlen
k_end = (i + 1) * frame_seqlen
sliding_mask[:, :, q_start:q_end, k_start:k_end] = True
mask = mask & sliding_mask
The sliding_mask is AND-ed with the existing block-wise and video masks, ensuring any token outside the temporal window is explicitly prevented from attending. This preserves strict causality while bounding the attention context.
KV-Cache Eviction for Constant Memory
To prevent unbounded cache growth during streaming inference, LingBot-Map implements proactive cache eviction. The _apply_kv_cache_eviction_causal method (lines 98-114) prunes the KV cache before each attention operation, retaining only the most recent kv_cache_sliding_window frames plus optional scale frames.
# lingbot_map/layers/attention.py, lines 98-114
def _apply_kv_cache_eviction_causal(...):
sliding_window_frames = self.kv_cache_sliding_window
if num_cached_frames > sliding_window_frames + scale_frames:
evict_start = scale_frames
evict_end = num_cached_frames - sliding_window_frames
# ... (evict and optionally keep special tokens)
This eviction strategy ensures the cache size remains constant regardless of stream length, enabling processing of infinite video sequences on fixed hardware.
Configurable Window Parameters
LingBot-Map exposes sliding window controls at three hierarchical levels:
- Model-level:
sliding_window_sizeinGCTStreamcontrols the attention mask (measured in blocks) - KV-cache level:
kv_cache_sliding_windowdetermines physical cache retention (measured in frames) - Command-line: The benchmark script accepts
--sliding-windowto override defaults
These parameters are defined in the streaming model constructors:
# lingbot_map/models/gct_stream_window.py, lines 46-58
sliding_window_size: int = -1, # -1 → full causal (no mask)
kv_cache_sliding_window: int = 64, # default eviction window
Implementation Examples
Instantiate a Streaming Model with Custom Window Size
Configure a 32-block sliding window (32 × num_frame_per_block frames) with 64-frame KV-cache retention:
from lingbot_map.models.gct_stream_window import GCTStream
model = GCTStream(
img_size=224,
patch_size=16,
sliding_window_size=32, # 32 blocks → 32 × num_frame_per_block frames
kv_cache_sliding_window=64, # keep 64 frames in KV cache
enable_stream_inference=True,
)
Run Streaming Inference on Video Sequences
The model automatically handles KV-cache updates and eviction during streaming:
# Assume `frames` is a list of [B, 3, H, W] tensors
kv_cache = {}
outputs = []
for t, frame in enumerate(frames):
out, kv_cache = model(frame.unsqueeze(1), kv_cache=kv_cache)
outputs.append(out)
Override Window Size at Runtime
Temporarily adjust the sliding window for specific segments without reinitializing the model:
# Use a larger 48-block window for this specific step
out, kv_cache = model(
frame.unsqueeze(1),
kv_cache=kv_cache,
sliding_window_size=48,
)
Benchmark Memory Usage
Profile memory consumption with different window configurations using the provided benchmark script:
python scripts/benchmark_gct_memory.py \
--sliding-window 32 \
--seq-len 200 \
--batch-size 1
Why Sliding Window Attention Matters for Streaming Video
By restricting attention to a bounded recent context, LingBot-Map achieves three critical capabilities for real-time robotics and SLAM applications:
- Constant memory usage: The KV cache never grows beyond
kv_cache_sliding_window + scale_frames, regardless of video length - Reduced computational cost: Scaled dot-product attention multiplies queries with a limited key set proportional to the window size rather than the full sequence
- Preserved causality: The mask enforces that each frame cannot attend to future frames, which is essential for online inference on live video streams
Summary
- Sliding window attention in LingBot-Map uses a configurable binary mask in
CausalAttention.forwardto limit each query to recent frames only - The KV cache is actively pruned via
_apply_kv_cache_eviction_causalto maintain constant memory during unbounded streaming - Window sizes are controlled by
sliding_window_size(attention mask) andkv_cache_sliding_window(cache retention) parameters - The mechanism enables processing of arbitrarily long video sequences on fixed hardware while preserving strict temporal causality
Frequently Asked Questions
What is the difference between sliding_window_size and kv_cache_sliding_window?
sliding_window_size controls the attention mask logic in CausalAttention, measured in blocks (where each block contains num_frame_per_block frames), determining which historical frames the model can attend to. kv_cache_sliding_window controls physical memory management, measured in frames, determining how many past key-value tensors are retained in the cache. The former affects attention computation scope, while the latter affects GPU memory allocation.
How does LingBot-Map maintain causality with sliding window attention?
Causality is enforced through the binary mask construction in lingbot_map/layers/attention.py, where the sliding window mask is AND-ed with the base causal mask. This ensures that position i can only attend to positions within the window [max(0, i - window_size + 1), i], preventing any information flow from future frames regardless of window size configuration.
Can I change the sliding window size during inference?
Yes, you can override sliding_window_size at runtime by passing it as a keyword argument to the model's forward method. This allows dynamic adjustment of the attention context for specific segments without model reinitialization, though the KV cache must be managed appropriately to reflect the new window constraints.
Where is the sliding window mask implemented in the source code?
The sliding window mask logic is implemented in lingbot_map/layers/attention.py within the CausalAttention class (lines 73-86). The complementary KV-cache eviction logic is in the same file (lines 98-114), while configuration parameters are defined in lingbot_map/models/gct_stream_window.py. Additional flash-infer cache handling resides in lingbot_map/layers/flashinfer_cache.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →