How to Configure keyframe_interval to Optimize Memory for Very Long Video Sequences in LingBot-Map

Set --keyframe_interval to a value greater than 1 (e.g., 10 or 20) or use auto mode to subsample cached frames and prevent GPU out-of-memory errors on sequences exceeding 320 frames.

LingBot-Map processes video through a causal transformer architecture that maintains a KV-cache of past observations. By default, every frame occupies a cache slot, causing memory to grow linearly with sequence length. For videos exceeding the training horizon of approximately 320 frames, the cache becomes the dominant memory consumer. Configuring the _keyframe_interval parameter allows you to store only every N-th frame, dramatically reducing GPU RAM usage on long sequences.

How the KV-Cache Drives Memory Consumption

The transformer in benchmark/methods/lingbot_map.py retains attention matrices for historical frames in a sliding-window KV-cache. Each stored keyframe preserves the full attention context for all heads, which dominates GPU memory allocation during inference.

When processing very long videos (>10,000 frames), an interval of 1 (every frame cached) exhausts VRAM and triggers out-of-memory (OOM) failures. Increasing the interval linearly reduces the cache size by storing only subsampled reference frames while still generating predictions for every input frame.

Understanding the keyframe_interval Configuration

The _keyframe_interval field (exposed via CLI as --keyframe_interval) controls cache subsampling behavior:

  • 1 (default for short sequences): Every frame becomes a keyframe. This provides maximum temporal context but highest memory usage.
  • N > 1: Only every N-th frame is cached. Intermediate frames generate pose and depth predictions but do not occupy permanent cache slots.
  • auto (or omitted): The library automatically selects an interval based on sequence length. If total frames ≤ auto_keyframe_threshold (default 320), it uses 1; otherwise it calculates ceil(num_frames / threshold).

Automatic Resolution Logic

The resolution logic resides in _resolve_keyframe_interval at lines 24-37 of benchmark/methods/lingbot_map.py. When the config value is "auto", None, or 0, the function computes:

interval = ceil(num_frames / auto_keyframe_threshold)

For a 25,000-frame video, this returns 79 (ceil(25000/320)), reducing the cache from 25,000 slots to approximately 320 slots. When you supply a manual integer, the resolver bypasses auto-selection and uses your value directly.

Memory Optimization Strategies by Use Case

Streaming Mode for Continuous Video

For sequences exceeding 10,000 frames in streaming mode (--mode streaming), manually set a fixed interval to cap memory usage:

python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path very_long_video.mp4 \
    --mode streaming \
    --keyframe_interval 10

Additional optimizations: Add --offload_to_cpu to move evicted cache entries to system RAM, or reduce --num_scale_frames to minimize the initial activation peak.

Windowed Inference with Overlap

For batch processing of long videos in windowed mode, combine interval spacing with cache resets:

python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path long_video.mp4 \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 10 \
    --overlap_keyframes 8

The --overlap_keyframes parameter preserves temporal continuity between windows by carrying over the final 8 keyframes from the previous window into the next, mitigating accuracy loss from the reduced interval.

Critical Memory Constraints

On GPUs with limited VRAM, aggressive subsampling may be necessary:

Scenario Recommended Interval Companion Flags
>20k frames, 16GB GPU --keyframe_interval 20 --num_scale_frames 2
>50k frames, 24GB GPU --keyframe_interval 30 --offload_to_cpu
Windowed >3k frames --keyframe_interval 10 --window_size 128 --overlap_keyframes 8

Implementation Examples

Resolve the automatic interval programmatically before inference:

from benchmark.methods.lingbot_map import _resolve_keyframe_interval

num_frames = 25000
cfg_val = "auto"  # Could also be None, 0, or an integer

interval = _resolve_keyframe_interval(cfg_val, num_frames)
print(interval)  # Output: 79

Explicit configuration via YAML in benchmark/configs/methods/lingbot_map.yaml:

_keyframe_interval: 10  # Fixed interval

_auto_keyframe_threshold: 320  # Threshold for auto mode

Summary

  • The KV-cache in LingBot-Map stores attention matrices for keyframes, consuming GPU RAM linearly with the number of cached frames.
  • Configure _keyframe_interval to subsample cached frames; values greater than 1 reduce memory proportionally while maintaining inference on all frames.
  • Use auto mode to automatically calculate intervals for sequences longer than 320 frames, or manually specify intervals (10-30) for very long videos.
  • Combine interval configuration with windowed mode and --overlap_keyframes to balance memory efficiency and temporal accuracy.
  • Reference the resolver implementation in benchmark/methods/lingbot_map.py (lines 24-37) to understand automatic selection logic.

Frequently Asked Questions

What happens if I set keyframe_interval too high?

Setting the interval too high (e.g., 50+ on a standard video) reduces temporal context available to the transformer, potentially degrading pose estimation accuracy and depth consistency. The model sees fewer past reference frames, which is particularly problematic for fast camera motion. For sequences far beyond the training horizon, the accuracy loss is usually acceptable compared to the alternative of OOM crashes.

Does keyframe_interval affect inference speed?

Yes, but indirectly. A larger interval reduces the KV-cache size, which decreases the computational cost of attention operations over long sequences. However, the model still processes every input frame to generate predictions; it simply does not cache the intermediate states. The primary benefit is memory reduction rather than throughput increase.

How does windowed mode interact with keyframe_interval?

In windowed mode (--mode windowed), the KV-cache resets every --window_size frames, creating a hard boundary regardless of interval. The keyframe interval determines how many slots are occupied within each window. Use --overlap_keyframes to carry context between windows, ensuring that the last N keyframes from window i seed the cache for window i+1, maintaining trajectory continuity despite periodic cache resets.

Where is the default keyframe_interval defined?

Default values reside in benchmark/configs/methods/lingbot_map.yaml under the _keyframe_interval and _auto_keyframe_threshold fields. The CLI flag --keyframe_interval in demo.py overrides these defaults. When omitted, the _resolve_keyframe_interval function in benchmark/methods/lingbot_map.py applies the auto-resolution logic based on the input frame count.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →