# Sliding Window Attention in LingBot-Map: Efficient Streaming Video Processing with Constant Memory

> Discover sliding window attention in LingBot-Map, an efficient streaming video processing technique. Achieve constant memory usage and unbounded sequence processing with this causal self-attention mechanism.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-28

---

**Sliding window attention in LingBot-Map is a causal self-attention mechanism that restricts each video frame to attend only to the most recent N frames, enabling unbounded sequence processing with constant memory usage through aggressive KV-cache eviction.**

LingBot-Map processes arbitrarily long video streams using causal self-attention while keeping compute and memory bounded. The framework implements this through a configurable masking strategy and cache eviction policy defined in the [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) source file, allowing the model to handle live video feeds without unbounded memory growth.

## How Sliding Window Attention Works in LingBot-Map

### Sliding-Window Mask Construction

In `CausalAttention.forward`, a binary mask is constructed that limits each query frame to attend only to the most recent frames within the configured window. The window size is calculated as `sliding_window_size × num_frame_per_block`, representing the number of historical frames retained for attention computation.

The implementation (lines 73-86 in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py)) builds this mask iteratively for each frame position:

```python

# lingbot_map/layers/attention.py, lines 73-86

if sliding_window_size > 0 and frame_seqlen is not None:
    for i in range(num_frames):
        window_size_in_frames = sliding_window_size * num_frame_per_block
        window_start_frame = max(0, i - window_size_in_frames + 1)
        k_start = window_start_frame * frame_seqlen
        k_end   = (i + 1) * frame_seqlen
        sliding_mask[:, :, q_start:q_end, k_start:k_end] = True
    mask = mask & sliding_mask

```

The `sliding_mask` is AND-ed with the existing block-wise and video masks, ensuring any token outside the temporal window is explicitly prevented from attending. This preserves strict causality while bounding the attention context.

### KV-Cache Eviction for Constant Memory

To prevent unbounded cache growth during streaming inference, LingBot-Map implements proactive cache eviction. The `_apply_kv_cache_eviction_causal` method (lines 98-114) prunes the KV cache before each attention operation, retaining only the most recent `kv_cache_sliding_window` frames plus optional scale frames.

```python

# lingbot_map/layers/attention.py, lines 98-114

def _apply_kv_cache_eviction_causal(...):
    sliding_window_frames = self.kv_cache_sliding_window
    if num_cached_frames > sliding_window_frames + scale_frames:
        evict_start = scale_frames
        evict_end   = num_cached_frames - sliding_window_frames
        # ... (evict and optionally keep special tokens)

```

This eviction strategy ensures the cache size remains constant regardless of stream length, enabling processing of infinite video sequences on fixed hardware.

### Configurable Window Parameters

LingBot-Map exposes sliding window controls at three hierarchical levels:

- **Model-level**: `sliding_window_size` in `GCTStream` controls the attention mask (measured in blocks)
- **KV-cache level**: `kv_cache_sliding_window` determines physical cache retention (measured in frames)
- **Command-line**: The benchmark script accepts `--sliding-window` to override defaults

These parameters are defined in the streaming model constructors:

```python

# lingbot_map/models/gct_stream_window.py, lines 46-58

sliding_window_size: int = -1,          # -1 → full causal (no mask)

kv_cache_sliding_window: int = 64,      # default eviction window

```

## Implementation Examples

### Instantiate a Streaming Model with Custom Window Size

Configure a 32-block sliding window (32 × `num_frame_per_block` frames) with 64-frame KV-cache retention:

```python
from lingbot_map.models.gct_stream_window import GCTStream

model = GCTStream(
    img_size=224,
    patch_size=16,
    sliding_window_size=32,      # 32 blocks → 32 × num_frame_per_block frames

    kv_cache_sliding_window=64, # keep 64 frames in KV cache

    enable_stream_inference=True,
)

```

### Run Streaming Inference on Video Sequences

The model automatically handles KV-cache updates and eviction during streaming:

```python

# Assume `frames` is a list of [B, 3, H, W] tensors

kv_cache = {}
outputs = []
for t, frame in enumerate(frames):
    out, kv_cache = model(frame.unsqueeze(1), kv_cache=kv_cache)
    outputs.append(out)

```

### Override Window Size at Runtime

Temporarily adjust the sliding window for specific segments without reinitializing the model:

```python

# Use a larger 48-block window for this specific step

out, kv_cache = model(
    frame.unsqueeze(1),
    kv_cache=kv_cache,
    sliding_window_size=48,
)

```

### Benchmark Memory Usage

Profile memory consumption with different window configurations using the provided benchmark script:

```bash
python scripts/benchmark_gct_memory.py \
    --sliding-window 32 \
    --seq-len 200 \
    --batch-size 1

```

## Why Sliding Window Attention Matters for Streaming Video

By restricting attention to a bounded recent context, LingBot-Map achieves three critical capabilities for real-time robotics and SLAM applications:

- **Constant memory usage**: The KV cache never grows beyond `kv_cache_sliding_window + scale_frames`, regardless of video length
- **Reduced computational cost**: Scaled dot-product attention multiplies queries with a limited key set proportional to the window size rather than the full sequence
- **Preserved causality**: The mask enforces that each frame cannot attend to future frames, which is essential for online inference on live video streams

## Summary

- Sliding window attention in LingBot-Map uses a configurable binary mask in `CausalAttention.forward` to limit each query to recent frames only
- The KV cache is actively pruned via `_apply_kv_cache_eviction_causal` to maintain constant memory during unbounded streaming
- Window sizes are controlled by `sliding_window_size` (attention mask) and `kv_cache_sliding_window` (cache retention) parameters
- The mechanism enables processing of arbitrarily long video sequences on fixed hardware while preserving strict temporal causality

## Frequently Asked Questions

### What is the difference between sliding_window_size and kv_cache_sliding_window?

**`sliding_window_size`** controls the attention mask logic in `CausalAttention`, measured in blocks (where each block contains `num_frame_per_block` frames), determining which historical frames the model can attend to. **`kv_cache_sliding_window`** controls physical memory management, measured in frames, determining how many past key-value tensors are retained in the cache. The former affects attention computation scope, while the latter affects GPU memory allocation.

### How does LingBot-Map maintain causality with sliding window attention?

Causality is enforced through the binary mask construction in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py), where the sliding window mask is AND-ed with the base causal mask. This ensures that position *i* can only attend to positions within the window `[max(0, i - window_size + 1), i]`, preventing any information flow from future frames regardless of window size configuration.

### Can I change the sliding window size during inference?

Yes, you can override `sliding_window_size` at runtime by passing it as a keyword argument to the model's forward method. This allows dynamic adjustment of the attention context for specific segments without model reinitialization, though the KV cache must be managed appropriately to reflect the new window constraints.

### Where is the sliding window mask implemented in the source code?

The sliding window mask logic is implemented in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) within the `CausalAttention` class (lines 73-86). The complementary KV-cache eviction logic is in the same file (lines 98-114), while configuration parameters are defined in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py). Additional flash-infer cache handling resides in [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py).