Sliding Window Attention in LingBot-Map: Efficient Streaming Video Processing with Constant Memory

Sliding window attention in LingBot-Map is a causal self-attention mechanism that restricts each video frame to attend only to the most recent N frames, enabling unbounded sequence processing with constant memory usage through aggressive KV-cache eviction.

LingBot-Map processes arbitrarily long video streams using causal self-attention while keeping compute and memory bounded. The framework implements this through a configurable masking strategy and cache eviction policy defined in the lingbot_map/layers/attention.py source file, allowing the model to handle live video feeds without unbounded memory growth.

How Sliding Window Attention Works in LingBot-Map

Sliding-Window Mask Construction

In CausalAttention.forward, a binary mask is constructed that limits each query frame to attend only to the most recent frames within the configured window. The window size is calculated as sliding_window_size × num_frame_per_block, representing the number of historical frames retained for attention computation.

The implementation (lines 73-86 in lingbot_map/layers/attention.py) builds this mask iteratively for each frame position:


# lingbot_map/layers/attention.py, lines 73-86

if sliding_window_size > 0 and frame_seqlen is not None:
    for i in range(num_frames):
        window_size_in_frames = sliding_window_size * num_frame_per_block
        window_start_frame = max(0, i - window_size_in_frames + 1)
        k_start = window_start_frame * frame_seqlen
        k_end   = (i + 1) * frame_seqlen
        sliding_mask[:, :, q_start:q_end, k_start:k_end] = True
    mask = mask & sliding_mask

The sliding_mask is AND-ed with the existing block-wise and video masks, ensuring any token outside the temporal window is explicitly prevented from attending. This preserves strict causality while bounding the attention context.

KV-Cache Eviction for Constant Memory

To prevent unbounded cache growth during streaming inference, LingBot-Map implements proactive cache eviction. The _apply_kv_cache_eviction_causal method (lines 98-114) prunes the KV cache before each attention operation, retaining only the most recent kv_cache_sliding_window frames plus optional scale frames.


# lingbot_map/layers/attention.py, lines 98-114

def _apply_kv_cache_eviction_causal(...):
    sliding_window_frames = self.kv_cache_sliding_window
    if num_cached_frames > sliding_window_frames + scale_frames:
        evict_start = scale_frames
        evict_end   = num_cached_frames - sliding_window_frames
        # ... (evict and optionally keep special tokens)

This eviction strategy ensures the cache size remains constant regardless of stream length, enabling processing of infinite video sequences on fixed hardware.

Configurable Window Parameters

LingBot-Map exposes sliding window controls at three hierarchical levels:

  • Model-level: sliding_window_size in GCTStream controls the attention mask (measured in blocks)
  • KV-cache level: kv_cache_sliding_window determines physical cache retention (measured in frames)
  • Command-line: The benchmark script accepts --sliding-window to override defaults

These parameters are defined in the streaming model constructors:


# lingbot_map/models/gct_stream_window.py, lines 46-58

sliding_window_size: int = -1,          # -1 → full causal (no mask)

kv_cache_sliding_window: int = 64,      # default eviction window

Implementation Examples

Instantiate a Streaming Model with Custom Window Size

Configure a 32-block sliding window (32 × num_frame_per_block frames) with 64-frame KV-cache retention:

from lingbot_map.models.gct_stream_window import GCTStream

model = GCTStream(
    img_size=224,
    patch_size=16,
    sliding_window_size=32,      # 32 blocks → 32 × num_frame_per_block frames

    kv_cache_sliding_window=64, # keep 64 frames in KV cache

    enable_stream_inference=True,
)

Run Streaming Inference on Video Sequences

The model automatically handles KV-cache updates and eviction during streaming:


# Assume `frames` is a list of [B, 3, H, W] tensors

kv_cache = {}
outputs = []
for t, frame in enumerate(frames):
    out, kv_cache = model(frame.unsqueeze(1), kv_cache=kv_cache)
    outputs.append(out)

Override Window Size at Runtime

Temporarily adjust the sliding window for specific segments without reinitializing the model:


# Use a larger 48-block window for this specific step

out, kv_cache = model(
    frame.unsqueeze(1),
    kv_cache=kv_cache,
    sliding_window_size=48,
)

Benchmark Memory Usage

Profile memory consumption with different window configurations using the provided benchmark script:

python scripts/benchmark_gct_memory.py \
    --sliding-window 32 \
    --seq-len 200 \
    --batch-size 1

Why Sliding Window Attention Matters for Streaming Video

By restricting attention to a bounded recent context, LingBot-Map achieves three critical capabilities for real-time robotics and SLAM applications:

  • Constant memory usage: The KV cache never grows beyond kv_cache_sliding_window + scale_frames, regardless of video length
  • Reduced computational cost: Scaled dot-product attention multiplies queries with a limited key set proportional to the window size rather than the full sequence
  • Preserved causality: The mask enforces that each frame cannot attend to future frames, which is essential for online inference on live video streams

Summary

  • Sliding window attention in LingBot-Map uses a configurable binary mask in CausalAttention.forward to limit each query to recent frames only
  • The KV cache is actively pruned via _apply_kv_cache_eviction_causal to maintain constant memory during unbounded streaming
  • Window sizes are controlled by sliding_window_size (attention mask) and kv_cache_sliding_window (cache retention) parameters
  • The mechanism enables processing of arbitrarily long video sequences on fixed hardware while preserving strict temporal causality

Frequently Asked Questions

What is the difference between sliding_window_size and kv_cache_sliding_window?

sliding_window_size controls the attention mask logic in CausalAttention, measured in blocks (where each block contains num_frame_per_block frames), determining which historical frames the model can attend to. kv_cache_sliding_window controls physical memory management, measured in frames, determining how many past key-value tensors are retained in the cache. The former affects attention computation scope, while the latter affects GPU memory allocation.

How does LingBot-Map maintain causality with sliding window attention?

Causality is enforced through the binary mask construction in lingbot_map/layers/attention.py, where the sliding window mask is AND-ed with the base causal mask. This ensures that position i can only attend to positions within the window [max(0, i - window_size + 1), i], preventing any information flow from future frames regardless of window size configuration.

Can I change the sliding window size during inference?

Yes, you can override sliding_window_size at runtime by passing it as a keyword argument to the model's forward method. This allows dynamic adjustment of the attention context for specific segments without model reinitialization, though the KV cache must be managed appropriately to reflect the new window constraints.

Where is the sliding window mask implemented in the source code?

The sliding window mask logic is implemented in lingbot_map/layers/attention.py within the CausalAttention class (lines 73-86). The complementary KV-cache eviction logic is in the same file (lines 98-114), while configuration parameters are defined in lingbot_map/models/gct_stream_window.py. Additional flash-infer cache handling resides in lingbot_map/layers/flashinfer_cache.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →