# How Keyframe Interval Affects KV Cache Memory in Streaming Inference

> Discover how keyframe interval impacts KV cache memory in streaming inference. Learn to reduce memory usage by retaining only essential frames while maintaining context for efficient processing.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-26

---

**Increasing the `keyframe_interval` sparsifies the KV cache by retaining only every Nth frame's key-value pairs, reducing memory consumption proportionally by roughly the factor `1/keyframe_interval` while allowing all frames to attend to the cached history.**

In the `Robbyant/lingbot-map` repository, the `keyframe_interval` parameter controls a critical memory optimization for transformer-based streaming video models. By selectively storing only keyframe KV tensors and discarding intermediate frame computations, the system prevents the linear growth of cache memory that would otherwise occur during long sequence inference.

## What Is the Keyframe Interval?

The `keyframe_interval` is an integer parameter that determines which frames become **keyframes** in the KV (key-value) cache during streaming inference. When set to `1` (default), every frame functions as a keyframe, meaning the model stores its KV representation indefinitely. When set to values greater than `1`, the system adopts a sparse caching strategy where only specific frames retain their KV tensors in memory.

Non-key frames still participate in the forward pass and attend to all cached KV data, but their own computed KV pairs are immediately discarded. This architectural choice creates a trade-off between **temporal granularity** (how many frames retain full context) and **memory efficiency** (how much GPU RAM the cache consumes).

## How Keyframe Interval Reduces KV Cache Memory

### The Fixed-Interval Keyframe Mechanism

In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), the keyframe decision occurs inside the per-frame inference loop at lines 665-666:

```python

# Fixed-interval keyframe mode

is_keyframe = (keyframe_interval <= 1) or ((i - scale_frames) % keyframe_interval == 0)
if not is_keyframe:
    self._set_skip_append(True)   # do not append KV for this frame

```

When `is_keyframe` evaluates to `False`, the model invokes `_set_skip_append(True)`, which signals the KV cache manager to withhold the current frame's key-value tensors from permanent storage. After the forward pass completes, the model restores normal appending behavior for subsequent frames.

### Memory Scaling Characteristics

According to the implementation documentation at lines 530-536 of the same file, the memory reduction follows an inverse proportional relationship with the interval value. Because only one out of every `keyframe_interval` frames contributes entries to the KV cache, the total cache size grows at approximately **1/keyframe_interval** of the rate observed when caching every frame.

For example, with `keyframe_interval=4`, the cache retains KV data for roughly 25% of the frames compared to the default configuration. This sub-linear growth pattern proves essential for processing long video sequences where a linear cache would exhaust available GPU memory.

## Flow-Based Keyframe Override

The repository provides an alternative keyframe selection mechanism based on optical flow magnitude. When `flow_threshold` is set to a positive value (e.g., `0.5` pixels), it overrides the fixed `keyframe_interval` logic.

In flow-based mode, the model calculates pixel motion between frames. If the motion exceeds `flow_threshold` or the gap since the last keyframe reaches `max_non_keyframe_gap`, the current frame becomes a keyframe regardless of the interval setting. This dynamic approach uses the same `_set_skip_append(True)` mechanics but adapts keyframe density to actual content motion rather than fixed temporal intervals.

## Practical Configuration Examples

Configure the `keyframe_interval` parameter when calling `inference_streaming()` to balance memory constraints against temporal resolution:

```python

# Example 1: Default behavior (every frame is a keyframe)

model.inference_streaming(images, keyframe_interval=1)

# KV cache grows linearly with every frame → highest memory usage, full temporal precision.

```

```python

# Example 2: Sparse keyframing every 4th frame

model.inference_streaming(images, keyframe_interval=4)

# Only frames 0, 4, 8, ... retain KV entries → ~75% memory reduction.

```

```python

# Example 3: Flow-based keyframe selection (overrides interval)

model.inference_streaming(
    images,
    keyframe_interval=10,          # Ignored because flow_threshold > 0

    flow_threshold=0.5,           # Keyframe triggered by motion > 0.5 pixels

    max_non_keyframe_gap=30,
)

# KV cache size varies based on motion in the video content.

```

## Summary

- **`keyframe_interval`** controls sparse caching in streaming transformer models by determining which frames retain KV tensors.
- **Memory reduction** scales approximately as `1/keyframe_interval`, preventing linear cache growth during long sequences.
- **Implementation** resides in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) (lines 530-536 and 665-666), utilizing `_set_skip_append(True)` to prevent non-keyframe KV storage.
- **Flow-based mode** can override fixed intervals when `flow_threshold > 0`, using optical flow magnitude to trigger keyframe selection.
- Supporting files include [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py) (v1 implementation), [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py) (base streaming utilities), and [`gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_base.py) (KV cache manager definitions).

## Frequently Asked Questions

### What is the default keyframe interval behavior?

By default, `keyframe_interval=1` treats every processed frame as a keyframe. In this configuration, the KV cache grows continuously as the model appends key-value tensors for each new frame without discarding any intermediate computations.

### How does keyframe interval impact inference quality?

Non-key frames still receive full predictions and attend to all cached KV data from previous keyframes; they merely do not contribute their own KV pairs to the cache. This maintains inference quality for the current frame while potentially reducing temporal coherence for very long dependencies, as intermediate frames cannot be directly attended to in future steps.

### Can I use both keyframe interval and flow threshold together?

No, the parameters are mutually exclusive. When `flow_threshold` is set to a value greater than `0`, the flow-based logic takes precedence and the fixed `keyframe_interval` is ignored. The system uses optical flow magnitude to determine keyframe placement instead of arithmetic modulo operations.

### Which files control KV cache management in lingbot-map?

The primary implementation resides in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), specifically lines 530-536 (documentation and initialization) and 665-666 (per-frame keyframe logic). The base KV cache manager is defined in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py), while [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) contains shared streaming utilities. An analogous implementation exists in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) for the non-v2 model variant.