How `keyframe_interval` Controls KV Cache Memory and Reconstruction Quality in lingbot-map

The keyframe_interval parameter determines which frames store their Key-Value tensors in the cache during streaming inference, directly reducing memory usage by a factor of 1/N where N is the interval, while maintaining per-frame reconstruction quality through attention to cached keyframes.

The keyframe_interval parameter in the lingbot-map repository governs the critical memory-efficiency trade-off in streaming vision transformers. By selectively controlling which frames append their Key-Value (KV) representations to the persistent cache, this setting allows developers to scale inference to arbitrarily long video sequences without proportional GPU memory growth. Understanding how this parameter interacts with the GCTStreamWindowV2 implementation is essential for optimizing both hardware utilization and reconstruction fidelity.

How keyframe_interval Works

In lingbot_map/models/gct_stream_window_v2.py, the keyframe_interval acts as a selective gate for the KV cache append operation. When the model processes a streaming sequence, it evaluates each frame against this interval to determine its status as either a keyframe (stored permanently) or a non-keyframe (processed then discarded).

Keyframe Selection Logic

The logic follows a modulo pattern applied after the initial scale frames (controlled by num_scale_frames). When the frame index modulo keyframe_interval equals zero, the frame is designated as a keyframe. For these frames, the full KV tensors are appended to the persistent cache via the standard cache update path. Non-keyframes trigger the helper method _set_skip_append(True), which prevents the model from appending the current frame's KV tensors to the growing cache, as implemented in lines 28-33 of scripts/benchmark_gct_memory.py.

Impact on KV Cache Memory Usage

The memory implications are linear and precisely calculable, as documented in the model's docstring on lines 30-36 of gct_stream_window_v2.py. The cache growth rate is inversely proportional to the interval value.

Linear Growth vs. Subsampled Storage

When keyframe_interval=1, every frame becomes a keyframe, causing the KV cache to grow linearly with sequence length—exhibiting O(N) memory complexity. Setting keyframe_interval=4 reduces cache growth by approximately 75%, storing only 25% of the frames' KV representations. For an interval of 8, memory usage drops to roughly 12.5% of the baseline, enabling processing of thousands of frames on consumer GPUs.

Cache Skip Implementation

The implementation toggles cache behavior through the internal helpers defined in lines 4-22 of gct_stream_window_v2.py. The _set_skip_append(True) method signals the attention mechanism to perform the forward pass without updating the persistent KV store. This occurs only for non-keyframes; keyframes proceed with _set_skip_append(False), ensuring their KV tensors remain available for future attention operations.

Impact on Reconstruction Quality

Despite aggressive memory reduction, reconstruction quality remains robust because every frame—regardless of keyframe status—receives a full forward pass prediction for depth, pose, and other output modalities.

Per-Frame Prediction Guarantee

Non-keyframes still generate complete predictions. During processing, the attention mechanism utilizes the cached KV from the most recent keyframe combined with the current frame's ephemeral KV tensors. The model computes the full output using this composite attention context, then discards only the current frame's KV after the forward pass completes. This architecture ensures that depth maps and pose estimates maintain high fidelity even when the underlying KV cache stores only 10% of the frames.

Temporal Consistency Trade-offs

Quality degradation manifests primarily in long-range temporal consistency when using very large intervals. Since non-keyframes never contribute their KV tensors to the permanent cache, distant future frames cannot attend directly to intermediate non-keyframe content. In practice, intervals up to 8 show negligible perceptual quality loss, while the default keyframe_interval=1 provides maximum temporal coherence for applications requiring perfect consistency across extended sequences.

Configuration Examples and Benchmarking

To empirically verify memory savings, use the provided benchmark script or configure the model directly in Python:


# Example: run streaming inference with different keyframe intervals

from lingbot_map.models.gct_stream_window_v2 import GCTStreamWindowV2
import torch

model = GCTStreamWindowV2(...)
images = torch.randn(1, 50, 3, 480, 640)          # 50 frames

# 1️⃣ Every frame is a keyframe (default)

out_all_key = model.inference_streaming(
    images,
    num_scale_frames=5,
    keyframe_interval=1,          # <-- every frame stored in KV‑cache

    output_device=torch.device('cpu')
)

# 2️⃣ Store only every 4th frame in the KV‑cache

out_sparse_key = model.inference_streaming(
    images,
    num_scale_frames=5,
    keyframe_interval=4,          # <-- 1 in 4 frames kept

    output_device=torch.device('cpu')
)

print("Memory usage (approx.)")
print(model.get_kv_cache_info())

Run the dedicated benchmark to observe proportional memory reduction:

python scripts/benchmark_gct_memory.py \
    --frame-counts 1000 \
    --keyframe-interval 1   # full KV cache

    
python scripts/benchmark_gct_memory.py \
    --frame-counts 1000 \
    --keyframe-interval 8   # 8‑times less KV memory

The CSV output from benchmark_gct_memory.py shows peak GPU memory shrinking proportionally to the interval while the inference time remains stable, confirming that the memory reduction comes without computational penalty.

Summary

  • keyframe_interval=N reduces KV cache memory by approximately 1/N by storing only every Nth frame's Key-Value tensors in lingbot_map/aggregator/base.py.
  • Non-keyframes still produce full per-frame predictions using cached keyframe KV plus ephemeral current KV, then discard the current KV after the forward pass.
  • Memory control is implemented via _set_skip_append(True) in gct_stream_window_v2.py, which gates the cache update logic during inference_streaming().
  • Default value of 1 provides highest temporal consistency; larger values trade minor long-range coherence for significant memory savings.
  • Monitoring: Use model.get_kv_cache_info() to inspect current cache utilization after streaming inference.

Frequently Asked Questions

What is the default keyframe_interval value in lingbot-map?

The default value is 1, meaning every frame is treated as a keyframe and stored in the KV cache. This provides maximum temporal consistency at the cost of linear memory growth with sequence length. According to the source code in gct_stream_window_v2.py, this default ensures that the model maintains the richest possible attention context for applications where memory is not constrained.

Does increasing keyframe_interval reduce inference speed?

No, increasing the interval does not significantly reduce inference latency. While the cache append operation is skipped for non-keyframes, the model still performs a full forward pass for every frame to generate depth and pose predictions. The computation cost remains nearly constant because the attention mechanism still processes the current frame's KV tensors during the forward pass—it simply discards them immediately afterward rather than storing them permanently.

How does the model maintain quality for non-keyframe predictions without storing their KV tensors?

During processing of a non-keyframe, the attention mechanism in GCTStreamWindowV2 utilizes the cached KV from the most recent keyframe combined with the current frame's temporary KV tensors generated during that specific forward pass. This allows the model to maintain contextual awareness and produce accurate estimates without permanently storing every frame's representation. The ephemeral current KV provides immediate information, while the cached keyframe KV provides historical context.

Where is the KV cache physically stored in the architecture?

The KV cache resides in the aggregator class defined in lingbot_map/aggregator/base.py, which maintains the persistent tensor storage across the streaming sequence. The streaming model interfaces with this aggregator through helper methods like _set_skip_append() and _set_defer_eviction() (defined in lines 4-22 of gct_stream_window_v2.py) to control which frames populate the cache during the inference_streaming() execution loop.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →