# How to Configure keyframe_interval to Optimize Memory for Very Long Video Sequences in LingBot-Map

> Optimize GPU memory for long video sequences in LingBot Map by configuring keyframe_interval. Learn how to set this parameter to avoid out-of-memory errors.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-31

---

**Set `--keyframe_interval` to a value greater than 1 (e.g., 10 or 20) or use `auto` mode to subsample cached frames and prevent GPU out-of-memory errors on sequences exceeding 320 frames.**

LingBot-Map processes video through a **causal transformer** architecture that maintains a **KV-cache** of past observations. By default, every frame occupies a cache slot, causing memory to grow linearly with sequence length. For videos exceeding the training horizon of approximately 320 frames, the cache becomes the dominant memory consumer. Configuring the `_keyframe_interval` parameter allows you to store only every *N*-th frame, dramatically reducing GPU RAM usage on long sequences.

## How the KV-Cache Drives Memory Consumption

The transformer in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py) retains attention matrices for historical frames in a sliding-window KV-cache. Each stored keyframe preserves the full attention context for all heads, which dominates GPU memory allocation during inference.

When processing very long videos (>10,000 frames), an interval of `1` (every frame cached) exhausts VRAM and triggers out-of-memory (OOM) failures. Increasing the interval linearly reduces the cache size by storing only subsampled reference frames while still generating predictions for every input frame.

## Understanding the keyframe_interval Configuration

The `_keyframe_interval` field (exposed via CLI as `--keyframe_interval`) controls cache subsampling behavior:

- **`1`** (default for short sequences): Every frame becomes a keyframe. This provides maximum temporal context but highest memory usage.
- **`N > 1`**: Only every *N*-th frame is cached. Intermediate frames generate pose and depth predictions but do not occupy permanent cache slots.
- **`auto` (or omitted)**: The library automatically selects an interval based on sequence length. If total frames ≤ `auto_keyframe_threshold` (default 320), it uses `1`; otherwise it calculates `ceil(num_frames / threshold)`.

### Automatic Resolution Logic

The resolution logic resides in `_resolve_keyframe_interval` at lines 24-37 of [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py). When the config value is `"auto"`, `None`, or `0`, the function computes:

```python
interval = ceil(num_frames / auto_keyframe_threshold)

```

For a 25,000-frame video, this returns `79` (ceil(25000/320)), reducing the cache from 25,000 slots to approximately 320 slots. When you supply a manual integer, the resolver bypasses auto-selection and uses your value directly.

## Memory Optimization Strategies by Use Case

### Streaming Mode for Continuous Video

For sequences exceeding 10,000 frames in streaming mode (`--mode streaming`), manually set a fixed interval to cap memory usage:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path very_long_video.mp4 \
    --mode streaming \
    --keyframe_interval 10

```

**Additional optimizations:** Add `--offload_to_cpu` to move evicted cache entries to system RAM, or reduce `--num_scale_frames` to minimize the initial activation peak.

### Windowed Inference with Overlap

For batch processing of long videos in windowed mode, combine interval spacing with cache resets:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path long_video.mp4 \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 10 \
    --overlap_keyframes 8

```

The `--overlap_keyframes` parameter preserves temporal continuity between windows by carrying over the final 8 keyframes from the previous window into the next, mitigating accuracy loss from the reduced interval.

### Critical Memory Constraints

On GPUs with limited VRAM, aggressive subsampling may be necessary:

| Scenario | Recommended Interval | Companion Flags |
|----------|---------------------|-----------------|
| **>20k frames, 16GB GPU** | `--keyframe_interval 20` | `--num_scale_frames 2` |
| **>50k frames, 24GB GPU** | `--keyframe_interval 30` | `--offload_to_cpu` |
| **Windowed >3k frames** | `--keyframe_interval 10` | `--window_size 128 --overlap_keyframes 8` |

## Implementation Examples

Resolve the automatic interval programmatically before inference:

```python
from benchmark.methods.lingbot_map import _resolve_keyframe_interval

num_frames = 25000
cfg_val = "auto"  # Could also be None, 0, or an integer

interval = _resolve_keyframe_interval(cfg_val, num_frames)
print(interval)  # Output: 79

```

Explicit configuration via YAML in [`benchmark/configs/methods/lingbot_map.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/configs/methods/lingbot_map.yaml):

```yaml
_keyframe_interval: 10  # Fixed interval

_auto_keyframe_threshold: 320  # Threshold for auto mode

```

## Summary

- The **KV-cache** in LingBot-Map stores attention matrices for keyframes, consuming GPU RAM linearly with the number of cached frames.
- Configure **`_keyframe_interval`** to subsample cached frames; values greater than 1 reduce memory proportionally while maintaining inference on all frames.
- Use **`auto` mode** to automatically calculate intervals for sequences longer than 320 frames, or manually specify intervals (10-30) for very long videos.
- Combine interval configuration with **windowed mode** and **`--overlap_keyframes`** to balance memory efficiency and temporal accuracy.
- Reference the resolver implementation in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py) (lines 24-37) to understand automatic selection logic.

## Frequently Asked Questions

### What happens if I set keyframe_interval too high?

Setting the interval too high (e.g., 50+ on a standard video) reduces temporal context available to the transformer, potentially degrading pose estimation accuracy and depth consistency. The model sees fewer past reference frames, which is particularly problematic for fast camera motion. For sequences far beyond the training horizon, the accuracy loss is usually acceptable compared to the alternative of OOM crashes.

### Does keyframe_interval affect inference speed?

Yes, but indirectly. A larger interval reduces the KV-cache size, which decreases the computational cost of attention operations over long sequences. However, the model still processes every input frame to generate predictions; it simply does not cache the intermediate states. The primary benefit is memory reduction rather than throughput increase.

### How does windowed mode interact with keyframe_interval?

In windowed mode (`--mode windowed`), the KV-cache resets every `--window_size` frames, creating a hard boundary regardless of interval. The keyframe interval determines how many slots are occupied *within* each window. Use `--overlap_keyframes` to carry context between windows, ensuring that the last *N* keyframes from window *i* seed the cache for window *i+1*, maintaining trajectory continuity despite periodic cache resets.

### Where is the default keyframe_interval defined?

Default values reside in [`benchmark/configs/methods/lingbot_map.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/configs/methods/lingbot_map.yaml) under the `_keyframe_interval` and `_auto_keyframe_threshold` fields. The CLI flag `--keyframe_interval` in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) overrides these defaults. When omitted, the `_resolve_keyframe_interval` function in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py) applies the auto-resolution logic based on the input frame count.