What Is the keyframe_interval Parameter and How Does It Reduce Memory Usage in LingBot-Map

The keyframe_interval parameter controls how frequently frames are stored in the attention KV-cache during streaming inference, reducing GPU memory consumption linearly by approximately 1/N when set to N.

LingBot-Map is an open-source streaming inference framework designed for processing long visual sequences with efficient memory management. The keyframe_interval runtime flag determines how often the model retains key-value pairs in the transformer cache, directly controlling the KV-cache memory footprint during windowed attention operations. This parameter is particularly critical when handling sequences that exceed the model's ~320-frame RoPE training range.

Understanding the keyframe_interval Parameter

Default Behavior and Configuration Values

The --keyframe_interval flag accepts integer values that dictate cache retention frequency:

  • 1 (default): Every incoming frame becomes a keyframe, requiring the KV-cache to store one block per frame
  • N > 1: Only every N-th frame (after the initial scale frames) is preserved as a keyframe

When using values greater than 1, intervening frames still process through the model but their KV-entries are immediately discarded after the forward pass completes. This behavior is documented in the repository's README (lines 37-40) and registered in demo.py (lines 369-376).

Command-Line Interface Registration

The parameter is exposed through the demo script's argument parser:


# demo.py (lines 369-376)

parser.add_argument(
    '--keyframe_interval',
    type=int,
    default=1,
    help='Interval between keyframes in streaming mode (higher = less memory)'
)

Memory Usage Impact and Calculation

Linear Memory Reduction Relationship

The effect on GPU memory is directly proportional to the interval value. The number of cached blocks decreases by roughly 1 / keyframe_interval, making this the primary mechanism for handling long sequences without encountering out-of-memory errors.

Cache Memory Formula Implementation

According to the source code in lingbot_map/models/gct_stream_window.py (lines 667-746), the cache memory calculation follows this precise formula:


# lingbot_map/models/gct_stream_window.py

total_elements = num_cached * self.num_heads * self.head_dim * self.block_size
cache_memory_mb = (total_elements * 2) / (1024 * 1024)   # bf16 → 2 bytes per element

Where:

  • num_cached represents the number of keyframe blocks currently stored
  • self.num_heads and self.head_dim are model architecture constants
  • self.block_size defines the cache block dimensions
  • The multiplication by 2 accounts for bf16 precision (2 bytes per element)

Thus, setting --keyframe_interval 10 reduces the num_cached value to approximately one-tenth of the baseline, cutting cache_memory_mb by roughly 90%.

Windowed Mode Frame Processing

In windowed streaming mode, the relationship between processed frames and cache growth follows this formula:


actual_frames = scale_frames + (window_size - scale_frames) * keyframe_interval

This equation ensures the KV-cache expands only with the effective number of keyframes, not the raw frame count. The scale_frames parameter defines initial frames that always become keyframes, while subsequent frames adhere to the interval sampling strategy. This architectural decision prevents memory from scaling linearly with sequence length during long-form video processing.

Practical Implementation Examples

Basic Command-Line Usage

Run inference with default caching (every frame stored):

python demo.py --image_folder example/university --model_path lingbot-map.pt

Reduce KV-cache memory by storing only every 4th frame:

python demo.py --image_folder example/university \
    --model_path lingbot-map.pt \
    --keyframe_interval 4

For maximum memory efficiency on very long sequences (storing every 10th frame):

python demo.py --image_folder example/university \
    --model_path lingbot-map.pt \
    --keyframe_interval 10

Programmatic Cache Inspection

Monitor memory usage within Python scripts using the GCTStreamWindow class:

from lingbot_map.models.gct_stream_window import GCTStreamWindow

model = GCTStreamWindow(...)
model.clean_kv_cache()

# Run inference here...

info = model.get_cache_memory_info()
print(f"Cached blocks: {info['num_cached_blocks']}, "
      f"Cache memory: {info['cache_memory_mb']} MiB")

Summary

  • keyframe_interval determines how frequently frames are retained in the transformer KV-cache during LingBot-Map streaming inference
  • Default value 1 stores every frame, while higher values (N) reduce cache size by approximately 1/N
  • Memory calculation occurs in gct_stream_window.py using the formula: num_cached * num_heads * head_dim * block_size * 2 / (1024^2) MB
  • Windowed mode uses the formula actual_frames = scale_frames + (window_size - scale_frames) * keyframe_interval to separate processing scope from cache growth
  • Recommended usage for sequences exceeding the ~320-frame RoPE training range to prevent GPU out-of-memory errors

Frequently Asked Questions

What is the default value of keyframe_interval in LingBot-Map?

The default value is 1, meaning every incoming frame becomes a keyframe and persists in the KV-cache. This configuration provides maximum accuracy but requires the most memory, as implemented in the argument parser within demo.py (lines 369-376).

How much memory does setting keyframe_interval to 10 actually save?

Setting keyframe_interval to 10 reduces the KV-cache memory footprint to approximately 10% of the baseline (a 90% reduction). According to the calculation logic in gct_stream_window.py, this stores only every 10th frame while discarding the intermediate KV-entries immediately after processing.

Does increasing the keyframe_interval affect inference accuracy?

While the source code does not explicitly quantify accuracy degradation, the README documentation suggests using this parameter specifically for sequences longer than the ~320-frame RoPE training range. The intervening frames still participate in the forward pass and contribute to predictions, but their attention states are not retained for future context windows.

Where is the keyframe logic implemented in the LingBot-Map source code?

The core implementation resides in lingbot_map/models/gct_stream_window.py (lines 667-746), which contains the cache memory calculation formula and keyframe retention logic. The CLI argument handling appears in demo.py (lines 369-376), while user-facing documentation exists in the README.md (lines 37-40).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →