How the Keyframe Strategy Controls Memory Usage in LingBot-Map
The keyframe strategy in LingBot-Map reduces GPU memory growth by storing only a subset of frame key-value (KV) caches, making memory usage roughly proportional to 1/keyframe_interval.
LingBot-Map implements a selective caching mechanism to manage the attention KV cache during long video sequences. By treating only specific frames as keyframes, the model discards intermediate frame caches immediately after the forward pass, preventing the linear memory explosion typical of standard transformer architectures.
How Keyframe Interval Shapes KV Cache Growth
The keyframe_interval parameter in lingbot_map/models/gct_stream_window_v2.py determines which frames retain their KV tensors. When set to N, the model marks every N-th frame (after the initial scale frames) as a keyframe.
Keyframe vs. Non-Keyframe Handling
Keyframes persist their KV pairs in the GPU cache, allowing subsequent frames to attend to them during attention operations. Non-keyframes trigger the _skip_append flag, causing their KV pairs to be computed but immediately discarded rather than stored. This selective retention is the core mechanism behind the memory savings.
Memory Growth Characteristics
Memory consumption scales inversely with the interval setting:
keyframe_interval = 1: Every frame is cached (default behavior), causing linear memory growth with sequence length.keyframe_interval = 4: Memory growth reduces to approximately 25% of the baseline.keyframe_interval = 8: Memory growth drops to roughly 12.5%.
As documented in the source code, this design "reduces KV cache memory growth by ~1/keyframe_interval"【https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L531-L536】.
Configuring the Keyframe Strategy
LingBot-Map supports two modes for selecting keyframes: fixed interval and optical flow-based detection.
Fixed Interval Mode
The standard approach uses a regular sampling interval controlled via the --keyframe_interval argument. This mode is predictable and works uniformly across static and dynamic scenes.
# Cache only every 4th frame
python demo.py --model_path lingbot-map.pt \
--image_folder example/courthouse \
--keyframe_interval 4
Flow-Based Keyframe Selection
When --flow_threshold is set greater than 0, the model uses optical flow magnitude to determine keyframes dynamically. Frames exhibiting significant motion are retained, while static frames are discarded. Despite the selection criteria change, the memory-saving principle remains identical: only designated keyframes retain KV entries【https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L37-L41】.
Monitoring Memory Usage
The model provides built-in utilities to inspect cache consumption. The get_kv_cache_info() method in gct_stream_window_v2.py returns the estimated cache size in megabytes, assuming bfloat16 tensor elements:
"Approximate memory usage in MB"【https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L84-L92】
For empirical validation, the repository includes scripts/benchmark_gct_memory.py, which measures peak CUDA memory across varying sequence lengths and keyframe intervals.
# Baseline: cache every frame (highest memory usage)
python scripts/benchmark_gct_memory.py \
--height 384 --width 518 \
--frame-counts 64 128 256 512 1024 2048 \
--keyframe_interval 1
# Optimized: cache every 8th frame (reduced memory footprint)
python scripts/benchmark_gct_memory.py \
--height 384 --width 518 \
--frame-counts 64 128 256 512 1024 2048 \
--keyframe_interval 8
The script outputs a CSV containing peak memory measurements for each configuration【https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py#L64-L66】, allowing you to observe the linear reduction in memory consumption as the interval increases.
Summary
- Mechanism: Only keyframes retain KV cache entries; non-keyframes set
_skip_appendand discard their tensors immediately. - Scaling: Memory usage grows proportionally to
1/keyframe_interval, enabling processing of tens of thousands of frames without OOM errors. - Configuration: Use
--keyframe_intervalfor fixed sampling or--flow_thresholdfor motion-adaptive selection. - Monitoring: Call
get_kv_cache_info()to inspect current cache size, or usebenchmark_gct_memory.pyfor systematic profiling.
Frequently Asked Questions
What is the default keyframe interval in LingBot-Map?
The default keyframe_interval is 1, meaning every frame is treated as a keyframe and stored in the KV cache. This provides maximum temporal fidelity but results in linear GPU memory growth with sequence length.
How does increasing the keyframe interval affect model accuracy?
Raising the interval reduces the temporal resolution of the attention mechanism, as non-keyframes cannot be attended to directly. The trade-off is between memory efficiency (higher intervals) and fine-grained temporal detail (lower intervals). For most long-sequence applications, the accuracy degradation is minimal compared to the memory savings gained.
Can I use different keyframe strategies for different video regions?
The current implementation in gct_stream_window_v2.py applies a global keyframe policy across the entire sequence. While the flow-based mode (--flow_threshold) adapts to local motion, it does not support region-specific keyframe densities within a single frame.
How do I estimate GPU memory requirements before running inference?
Use the get_kv_cache_info() method to query the approximate cache size in megabytes during a short test run. Combine this with the 1/keyframe_interval scaling factor to project memory needs for longer sequences. For systematic planning, run scripts/benchmark_gct_memory.py with your target frame counts and interval settings to generate precise peak memory measurements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →