Understanding the Trade‑offs Between `--keyframe_interval` and `--image_stride` for Long‑Sequence Processing in LingBot

While --keyframe_interval controls how often the KV cache stores full‑resolution attention keys during streaming inference, --image_stride (exposed as --stride in the CLI) subsamples the input sequence before processing, creating a direct tension between GPU memory efficiency, computational throughput, and temporal reconstruction fidelity.

The Robbyant/lingbot-map repository implements a streaming visual SLAM system that processes extended image sequences using transformer‑based global context transformers. When running long‑duration mapping tasks on limited hardware, configuring the trade‑offs between --keyframe_interval and --image_stride for long sequence processing correctly prevents out‑of‑memory errors while maintaining map accuracy.

What --keyframe_interval Controls

The --keyframe_interval parameter determines how frequently the system stores a full‑resolution keyframe in the KV cache during streaming inference. In lingbot_map/models/gct_stream_window_v2.py (lines 530‑540), this flag drives the is_keyframe decision logic that dictates whether the current frame’s attention keys and values are retained for future frames or discarded after processing.

Memory Impact of Keyframe Selection

Larger interval values (e.g., --keyframe_interval 8) reduce KV‑cache growth roughly by a factor of the interval setting. According to the implementation at lines 650‑667 in gct_stream_window_v2.py, the is_keyframe calculation for fixed‑interval mode ensures that only every N‑th frame triggers cache storage, while non‑keyframes immediately free their KV pairs.

Compute Characteristics

Non‑keyframes still attend to the cached KV of the most recent keyframe but skip the expensive KV‑generation step entirely. This yields compute savings proportional to the interval value without sacrificing the ability to reference prior context.

What --image_stride Controls

The --image_stride parameter (parsed as --stride in demo.py lines 340‑360) subsamples the input image sequence before any model processing occurs. The dataset loading logic in benchmark/dataset/base.py (lines 205‑216) applies this stride when building the image list, effectively decimating the temporal resolution before it reaches the transformer.

Memory and Compute Reduction

Higher stride values (e.g., --stride 4) lower overall memory consumption because fewer frames are ever loaded into GPU memory, shrinking both the activation tensors and the KV cache. This produces near‑linear speedup by cutting the total number of forward passes required to process the sequence.

Temporal Resolution Cost

Aggressive image‑stride can miss fast motions entirely, leading to gaps in the reconstructed map or motion‑blur artifacts in depth estimation. Small strides preserve temporal detail but increase both memory footprint and processing time.

Comparative Trade‑offs for Long Sequences

When configuring the pipeline for extended trajectories, consider these specific resource trade‑offs:

  • Memory efficiency: --keyframe_interval optimizes KV‑cache growth rate by reducing stored keyframe frequency, while --image_stride reduces absolute frame count and total activation memory
  • Computational cost: Keyframe intervals skip KV computation for intermediate frames, whereas stride values skip entire forward passes
  • Accuracy risks: Large keyframe intervals risk stale context during rapid scene changes, while high stride values risk missing transient motions entirely

Interaction Effects and Joint Tuning

These parameters interact significantly in production deployments. If you apply a high --image_stride (e.g., 4), you can safely increase --keyframe_interval because fewer total frames enter the pipeline, preventing KV‑cache overflow. Conversely, with --stride 1 (dense sampling), you must use smaller keyframe intervals to avoid exhausting GPU memory.

The scripts/benchmark_gct_memory.py script (lines 64‑71) exposes both flags specifically for empirical testing of these interactions on your target hardware.

Practical Configuration Examples

Run a balanced configuration for moderate sequences:

python -m lingbot_map.demo \
    --image_folder /path/to/images/ \
    --mode streaming \
    --keyframe_interval 4 \
    --stride 2 \
    --image_size 518 \
    --fps 30

Process very long sequences on limited GPU memory:

python -m lingbot_map.demo \
    --image_folder /path/to/long_sequence/ \
    --mode streaming \
    --keyframe_interval 8 \
    --stride 4 \
    --image_size 480

Enable flow‑based adaptive keyframe detection to protect accuracy during high‑motion segments:

python -m lingbot_map.demo \
    --image_folder /path/to/sequence/ \
    --mode streaming \
    --keyframe_interval 6 \
    --flow_threshold 0.5

Summary

  • --keyframe_interval manages internal KV‑cache frequency in gct_stream_window_v2.py, controlling memory growth during streaming inference by determining which frames retain their attention keys
  • --image_stride (CLI: --stride) filters input data at the loading stage in demo.py and benchmark/dataset/base.py, reducing the absolute number of frames processed and shrinking activation tensors
  • High keyframe intervals risk temporal fidelity degradation during rapid scene changes, while aggressive image‑stride risks missing fast motion entirely
  • These parameters must be tuned jointly: high stride permits larger intervals, while dense sampling requires smaller intervals to prevent cache overflow
  • For limited GPU memory on long sequences, combine --keyframe_interval 8 with --stride 4 as validated in scripts/benchmark_gct_memory.py

Frequently Asked Questions

Can I use --keyframe_interval and --image_stride independently?

Yes, but they interact closely. You can set either parameter independently, but as implemented in the LingBot pipeline, high --image_stride values (which reduce input frames) allow you to safely increase --keyframe_interval without overwhelming the KV cache. Conversely, with --stride 1 (dense sampling), you typically need smaller interval values to maintain memory bounds and prevent out‑of‑memory errors during long sequences.

Which parameter has greater impact on GPU memory usage?

--image_stride typically provides larger absolute memory savings because it reduces the total tensor count before processing begins, whereas --keyframe_interval only optimizes the KV‑cache growth rate. According to the implementation in gct_stream_window_v2.py, the KV cache scales with the number of keyframes, but the overall activation memory scales with the total frames ingested, making stride the more aggressive memory reducer.

How does the flow‑based keyframe detection interact with these settings?

When --keyframe_interval is greater than 1, the system in gct_stream_window_v2.py (lines 530‑540) uses optical flow to override the fixed interval if motion exceeds --flow_threshold. This means rapid motion can force additional keyframes beyond the specified interval, protecting mapping accuracy at the cost of unpredictable memory usage spikes that violate the nominal interval budget.

For real‑time processing on limited laptop hardware, use --stride 2 or --stride 4 to maintain interactive frame rates, combined with --keyframe_interval 4 or 6. This configuration, as exposed in scripts/benchmark_gct_memory.py, keeps the KV cache manageable while skipping computationally expensive forward passes on intermediate frames, ensuring the pipeline runs within typical mobile GPU memory constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →