# What Caused the SDPA KV Cache Bug in LingBot-Map and How It Was Fixed

> Learn how the SDPA KV cache bug in LingBot-Map caused memory issues and how the latest update fixed it with proper eviction and synchronization.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: bug-fix-deep-dive
- Published: 2026-07-31

---

**The SDPA KV cache bug was caused by a dictionary-based cache that never evicted stale key/value entries during sliding-window operations, leading to unbounded memory growth and degraded reconstruction quality; the 2026-06-28 update fixed this by implementing proper eviction pathways including defer-eviction flags and frame-counter synchronization.**

The `Robbyant/lingbot-map` repository provides streaming video processing capabilities using multiple attention backends. The June 28, 2026 update addressed a critical **SDPA KV cache bug** that affected long-sequence inference when using the scaled-dot-product attention backend instead of FlashInfer.

## Root Cause of the SDPA KV Cache Bug

The original implementation relied on a simple dictionary-based KV cache that lacked proper eviction mechanisms when processing streaming video with key-frame intervals greater than 1.

### Unbounded Dictionary Growth

In [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), the cache was implemented as a standard Python dictionary. When the model processed long sequences, stale key/value entries accumulated because the cache was never correctly **evicted** during sliding-window trims or key-frame interval operations. This caused the cache to grow without bound, eventually exhausting GPU memory on extended video sequences.

### Stale Attention Context

Beyond memory issues, the accumulated entries created **incorrect attention** patterns. Later frames attended to outdated keys from previous windows because obsolete entries remained active in the cache, degrading reconstruction quality on long videos.

## The 2026-06-28 Fix: Proper Eviction Pathways

The fix introduced a comprehensive eviction system that aligns the SDPA backend's cache management with the FlashInfer backend's semantics.

### Defer-Eviction Flag Implementation

A new boolean flag controls when eviction occurs to prevent race conditions during window operations. In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) (lines ≈ 426‑440), the `set_defer_eviction` method toggles this flag for both backends:

```python

# Called internally during window sliding operations

def set_defer_eviction(self, defer: bool):
    self._defer_eviction = defer
    # Synchronizes behavior between FlashInfer and SDPA caches

```

### Aggregator Frame Counter Updates

The `StreamAggregator` class in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) (lines ≈ 180‑200) now properly decrements its internal frame counter when evicting entries. The updated logic pops the oldest entry using `popitem(last=False)` and adjusts bookkeeping to maintain sequence integrity.

### Trim-Last-Frame Logic

Explicit cleanup occurs in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) (lines ≈ 455‑460) via the `_trim_last_frame` method. This function removes the last cached frame along dimension 2 when sliding the window:

```python
def _trim_last_frame(self):
    # Removes oldest entry from the dict-based cache

    self.kv_cache.popitem(last=False)

```

### SDPA Block Alignment

The `SDPABlock` class in [`lingbot_map/layers/block.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/block.py) now respects the new eviction flag. Additionally, `SDPAAttention` in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) mirrors FlashInfer's cache-management semantics, ensuring identical behavior across backends when handling the KV cache.

## Code Implementation Details

To run inference with the fixed SDPA backend:

```bash
python demo.py --model_path /path/to/lingbot-map.pt \
    --image_folder example/loop \
    --use_sdpa

```

You can manually control eviction behavior when building custom pipelines:

```python
from lingbot_map.aggregator.stream import StreamAggregator

agg = StreamAggregator(use_sdpa=True)
agg.set_defer_eviction(True)   # Enable safe eviction during sliding-window ops

```

## Impact and Recommendations

After the 2026-06-28 update, the SDPA backend maintains stable memory consumption during long-sequence processing and guarantees that each frame attends only to the intended recent context. While the SDPA implementation now performs reliably, the repository maintainers still recommend the **FlashInfer backend** for maximum performance.

## Summary

- The **SDPA KV cache bug** originated from a dictionary-based cache without eviction logic, causing memory exhaustion and attention errors on long videos.
- The fix added a **defer-eviction flag** (`set_defer_eviction`) to safely manage cache cleanup timing.
- **Frame counter synchronization** in the aggregator ensures accurate bookkeeping when entries are removed.
- The **trim-last-frame logic** explicitly removes obsolete entries using `popitem(last=False)`.
- **SDPA attention layers** now align with FlashInfer's cache-management semantics for consistent behavior across backends.

## Frequently Asked Questions

### What exactly triggered the SDPA KV cache memory leak?

The memory leak occurred because the dictionary-based KV cache in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) accumulated entries indefinitely when processing sequences with key-frame intervals greater than 1 or when sliding windows advanced. Without explicit eviction calls, the cache size grew linearly with sequence length until GPU memory was exhausted.

### How does the defer-eviction flag prevent cache corruption?

The defer-eviction flag delays cache eviction until the window operations complete, preventing the removal of entries that might still be referenced by ongoing attention computations. This ensures thread-safe cache management when the sliding window moves forward or when key-frame sampling skips positions.

### Should I migrate from FlashInfer to SDPA after this fix?

While the 2026-06-28 update makes the SDPA backend functional for long sequences, the repository documentation still recommends FlashInfer for optimal performance. Use SDPA when FlashInfer is unavailable or when debugging attention patterns, as both backends now produce identical results but FlashInfer remains faster.

### Which source files were modified in the 2026-06-28 update?

The primary changes occurred in four locations: [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) added eviction control and trimming logic; [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) updated frame counter handling; [`lingbot_map/layers/block.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/block.py) modified `SDPABlock` to respect eviction flags; and [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) aligned `SDPAAttention` with FlashInfer's cache semantics.