# What Is the keyframe_interval Parameter and How Does It Reduce Memory Usage in LingBot-Map

> Understand the keyframe_interval parameter in LingBot-Map. Learn how it reduces GPU memory by controlling frame storage in the attention KV-cache during inference.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-25

---

**The `keyframe_interval` parameter controls how frequently frames are stored in the attention KV-cache during streaming inference, reducing GPU memory consumption linearly by approximately 1/N when set to N.**

LingBot-Map is an open-source streaming inference framework designed for processing long visual sequences with efficient memory management. The `keyframe_interval` runtime flag determines how often the model retains key-value pairs in the transformer cache, directly controlling the **KV-cache memory footprint** during windowed attention operations. This parameter is particularly critical when handling sequences that exceed the model's ~320-frame RoPE training range.

## Understanding the keyframe_interval Parameter

### Default Behavior and Configuration Values

The `--keyframe_interval` flag accepts integer values that dictate cache retention frequency:

- **`1` (default)**: Every incoming frame becomes a keyframe, requiring the KV-cache to store one block per frame
- **`N > 1`**: Only every N-th frame (after the initial scale frames) is preserved as a keyframe

When using values greater than 1, intervening frames still process through the model but their KV-entries are immediately discarded after the forward pass completes. This behavior is documented in the repository's README (lines 37-40) and registered in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (lines 369-376).

### Command-Line Interface Registration

The parameter is exposed through the demo script's argument parser:

```python

# demo.py (lines 369-376)

parser.add_argument(
    '--keyframe_interval',
    type=int,
    default=1,
    help='Interval between keyframes in streaming mode (higher = less memory)'
)

```

## Memory Usage Impact and Calculation

### Linear Memory Reduction Relationship

The effect on GPU memory is directly proportional to the interval value. The number of cached blocks decreases by roughly **1 / keyframe_interval**, making this the primary mechanism for handling long sequences without encountering out-of-memory errors.

### Cache Memory Formula Implementation

According to the source code in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) (lines 667-746), the cache memory calculation follows this precise formula:

```python

# lingbot_map/models/gct_stream_window.py

total_elements = num_cached * self.num_heads * self.head_dim * self.block_size
cache_memory_mb = (total_elements * 2) / (1024 * 1024)   # bf16 → 2 bytes per element

```

Where:
- `num_cached` represents the number of keyframe blocks currently stored
- `self.num_heads` and `self.head_dim` are model architecture constants
- `self.block_size` defines the cache block dimensions
- The multiplication by 2 accounts for bf16 precision (2 bytes per element)

Thus, setting `--keyframe_interval 10` reduces the `num_cached` value to approximately one-tenth of the baseline, cutting `cache_memory_mb` by roughly 90%.

## Windowed Mode Frame Processing

In windowed streaming mode, the relationship between processed frames and cache growth follows this formula:

```

actual_frames = scale_frames + (window_size - scale_frames) * keyframe_interval

```

This equation ensures the KV-cache expands only with the **effective number of keyframes**, not the raw frame count. The `scale_frames` parameter defines initial frames that always become keyframes, while subsequent frames adhere to the interval sampling strategy. This architectural decision prevents memory from scaling linearly with sequence length during long-form video processing.

## Practical Implementation Examples

### Basic Command-Line Usage

Run inference with default caching (every frame stored):

```bash
python demo.py --image_folder example/university --model_path lingbot-map.pt

```

Reduce KV-cache memory by storing only every 4th frame:

```bash
python demo.py --image_folder example/university \
    --model_path lingbot-map.pt \
    --keyframe_interval 4

```

For maximum memory efficiency on very long sequences (storing every 10th frame):

```bash
python demo.py --image_folder example/university \
    --model_path lingbot-map.pt \
    --keyframe_interval 10

```

### Programmatic Cache Inspection

Monitor memory usage within Python scripts using the `GCTStreamWindow` class:

```python
from lingbot_map.models.gct_stream_window import GCTStreamWindow

model = GCTStreamWindow(...)
model.clean_kv_cache()

# Run inference here...

info = model.get_cache_memory_info()
print(f"Cached blocks: {info['num_cached_blocks']}, "
      f"Cache memory: {info['cache_memory_mb']} MiB")

```

## Summary

- **`keyframe_interval`** determines how frequently frames are retained in the transformer KV-cache during LingBot-Map streaming inference
- **Default value `1`** stores every frame, while higher values (N) reduce cache size by approximately 1/N
- **Memory calculation** occurs in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py) using the formula: `num_cached * num_heads * head_dim * block_size * 2 / (1024^2)` MB
- **Windowed mode** uses the formula `actual_frames = scale_frames + (window_size - scale_frames) * keyframe_interval` to separate processing scope from cache growth
- **Recommended usage** for sequences exceeding the ~320-frame RoPE training range to prevent GPU out-of-memory errors

## Frequently Asked Questions

### What is the default value of keyframe_interval in LingBot-Map?

The default value is `1`, meaning every incoming frame becomes a keyframe and persists in the KV-cache. This configuration provides maximum accuracy but requires the most memory, as implemented in the argument parser within [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (lines 369-376).

### How much memory does setting keyframe_interval to 10 actually save?

Setting `keyframe_interval` to 10 reduces the KV-cache memory footprint to approximately **10% of the baseline** (a 90% reduction). According to the calculation logic in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py), this stores only every 10th frame while discarding the intermediate KV-entries immediately after processing.

### Does increasing the keyframe_interval affect inference accuracy?

While the source code does not explicitly quantify accuracy degradation, the README documentation suggests using this parameter specifically for sequences longer than the ~320-frame RoPE training range. The intervening frames still participate in the forward pass and contribute to predictions, but their attention states are not retained for future context windows.

### Where is the keyframe logic implemented in the LingBot-Map source code?

The core implementation resides in **[`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py)** (lines 667-746), which contains the cache memory calculation formula and keyframe retention logic. The CLI argument handling appears in **[`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py)** (lines 369-376), while user-facing documentation exists in the **README.md** (lines 37-40).