How Keyframe Interval Affects KV Cache Memory in Streaming Inference
Increasing the keyframe_interval sparsifies the KV cache by retaining only every Nth frame's key-value pairs, reducing memory consumption proportionally by roughly the factor 1/keyframe_interval while allowing all frames to attend to the cached history.
In the Robbyant/lingbot-map repository, the keyframe_interval parameter controls a critical memory optimization for transformer-based streaming video models. By selectively storing only keyframe KV tensors and discarding intermediate frame computations, the system prevents the linear growth of cache memory that would otherwise occur during long sequence inference.
What Is the Keyframe Interval?
The keyframe_interval is an integer parameter that determines which frames become keyframes in the KV (key-value) cache during streaming inference. When set to 1 (default), every frame functions as a keyframe, meaning the model stores its KV representation indefinitely. When set to values greater than 1, the system adopts a sparse caching strategy where only specific frames retain their KV tensors in memory.
Non-key frames still participate in the forward pass and attend to all cached KV data, but their own computed KV pairs are immediately discarded. This architectural choice creates a trade-off between temporal granularity (how many frames retain full context) and memory efficiency (how much GPU RAM the cache consumes).
How Keyframe Interval Reduces KV Cache Memory
The Fixed-Interval Keyframe Mechanism
In lingbot_map/models/gct_stream_window_v2.py, the keyframe decision occurs inside the per-frame inference loop at lines 665-666:
# Fixed-interval keyframe mode
is_keyframe = (keyframe_interval <= 1) or ((i - scale_frames) % keyframe_interval == 0)
if not is_keyframe:
self._set_skip_append(True) # do not append KV for this frame
When is_keyframe evaluates to False, the model invokes _set_skip_append(True), which signals the KV cache manager to withhold the current frame's key-value tensors from permanent storage. After the forward pass completes, the model restores normal appending behavior for subsequent frames.
Memory Scaling Characteristics
According to the implementation documentation at lines 530-536 of the same file, the memory reduction follows an inverse proportional relationship with the interval value. Because only one out of every keyframe_interval frames contributes entries to the KV cache, the total cache size grows at approximately 1/keyframe_interval of the rate observed when caching every frame.
For example, with keyframe_interval=4, the cache retains KV data for roughly 25% of the frames compared to the default configuration. This sub-linear growth pattern proves essential for processing long video sequences where a linear cache would exhaust available GPU memory.
Flow-Based Keyframe Override
The repository provides an alternative keyframe selection mechanism based on optical flow magnitude. When flow_threshold is set to a positive value (e.g., 0.5 pixels), it overrides the fixed keyframe_interval logic.
In flow-based mode, the model calculates pixel motion between frames. If the motion exceeds flow_threshold or the gap since the last keyframe reaches max_non_keyframe_gap, the current frame becomes a keyframe regardless of the interval setting. This dynamic approach uses the same _set_skip_append(True) mechanics but adapts keyframe density to actual content motion rather than fixed temporal intervals.
Practical Configuration Examples
Configure the keyframe_interval parameter when calling inference_streaming() to balance memory constraints against temporal resolution:
# Example 1: Default behavior (every frame is a keyframe)
model.inference_streaming(images, keyframe_interval=1)
# KV cache grows linearly with every frame → highest memory usage, full temporal precision.
# Example 2: Sparse keyframing every 4th frame
model.inference_streaming(images, keyframe_interval=4)
# Only frames 0, 4, 8, ... retain KV entries → ~75% memory reduction.
# Example 3: Flow-based keyframe selection (overrides interval)
model.inference_streaming(
images,
keyframe_interval=10, # Ignored because flow_threshold > 0
flow_threshold=0.5, # Keyframe triggered by motion > 0.5 pixels
max_non_keyframe_gap=30,
)
# KV cache size varies based on motion in the video content.
Summary
keyframe_intervalcontrols sparse caching in streaming transformer models by determining which frames retain KV tensors.- Memory reduction scales approximately as
1/keyframe_interval, preventing linear cache growth during long sequences. - Implementation resides in
lingbot_map/models/gct_stream_window_v2.py(lines 530-536 and 665-666), utilizing_set_skip_append(True)to prevent non-keyframe KV storage. - Flow-based mode can override fixed intervals when
flow_threshold > 0, using optical flow magnitude to trigger keyframe selection. - Supporting files include
gct_stream_window.py(v1 implementation),gct_stream.py(base streaming utilities), andgct_base.py(KV cache manager definitions).
Frequently Asked Questions
What is the default keyframe interval behavior?
By default, keyframe_interval=1 treats every processed frame as a keyframe. In this configuration, the KV cache grows continuously as the model appends key-value tensors for each new frame without discarding any intermediate computations.
How does keyframe interval impact inference quality?
Non-key frames still receive full predictions and attend to all cached KV data from previous keyframes; they merely do not contribute their own KV pairs to the cache. This maintains inference quality for the current frame while potentially reducing temporal coherence for very long dependencies, as intermediate frames cannot be directly attended to in future steps.
Can I use both keyframe interval and flow threshold together?
No, the parameters are mutually exclusive. When flow_threshold is set to a value greater than 0, the flow-based logic takes precedence and the fixed keyframe_interval is ignored. The system uses optical flow magnitude to determine keyframe placement instead of arithmetic modulo operations.
Which files control KV cache management in lingbot-map?
The primary implementation resides in lingbot_map/models/gct_stream_window_v2.py, specifically lines 530-536 (documentation and initialization) and 665-666 (per-frame keyframe logic). The base KV cache manager is defined in lingbot_map/models/gct_base.py, while lingbot_map/models/gct_stream.py contains shared streaming utilities. An analogous implementation exists in lingbot_map/models/gct_stream_window.py for the non-v2 model variant.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →