# Understanding the Trade‑offs Between `--keyframe_interval` and `--image_stride` for Long‑Sequence Processing in LingBot

> Explore the trade-offs between --keyframe_interval and --image_stride for long sequence processing in LingBot. Optimize GPU memory, throughput, and fidelity for streaming inference.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-31

---

**While `--keyframe_interval` controls how often the KV cache stores full‑resolution attention keys during streaming inference, `--image_stride` (exposed as `--stride` in the CLI) subsamples the input sequence before processing, creating a direct tension between GPU memory efficiency, computational throughput, and temporal reconstruction fidelity.**

The `Robbyant/lingbot-map` repository implements a streaming visual SLAM system that processes extended image sequences using transformer‑based global context transformers. When running long‑duration mapping tasks on limited hardware, configuring the **trade‑offs between `--keyframe_interval` and `--image_stride` for long sequence processing** correctly prevents out‑of‑memory errors while maintaining map accuracy.

## What `--keyframe_interval` Controls

The `--keyframe_interval` parameter determines how frequently the system stores a **full‑resolution keyframe** in the KV cache during streaming inference. In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) (lines 530‑540), this flag drives the `is_keyframe` decision logic that dictates whether the current frame’s attention keys and values are retained for future frames or discarded after processing.

### Memory Impact of Keyframe Selection

Larger interval values (e.g., `--keyframe_interval 8`) reduce KV‑cache growth roughly by a factor of the interval setting. According to the implementation at lines 650‑667 in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py), the `is_keyframe` calculation for fixed‑interval mode ensures that only every N‑th frame triggers cache storage, while non‑keyframes immediately free their KV pairs.

### Compute Characteristics

Non‑keyframes still attend to the cached KV of the most recent keyframe but skip the expensive KV‑generation step entirely. This yields compute savings proportional to the interval value without sacrificing the ability to reference prior context.

## What `--image_stride` Controls

The `--image_stride` parameter (parsed as `--stride` in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) lines 340‑360) subsamples the **input image sequence** before any model processing occurs. The dataset loading logic in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py) (lines 205‑216) applies this stride when building the image list, effectively decimating the temporal resolution before it reaches the transformer.

### Memory and Compute Reduction

Higher stride values (e.g., `--stride 4`) lower overall memory consumption because fewer frames are ever loaded into GPU memory, shrinking both the activation tensors and the KV cache. This produces near‑linear speedup by cutting the total number of forward passes required to process the sequence.

### Temporal Resolution Cost

Aggressive image‑stride can miss fast motions entirely, leading to gaps in the reconstructed map or motion‑blur artifacts in depth estimation. Small strides preserve temporal detail but increase both memory footprint and processing time.

## Comparative Trade‑offs for Long Sequences

When configuring the pipeline for extended trajectories, consider these specific resource trade‑offs:

- **Memory efficiency**: `--keyframe_interval` optimizes KV‑cache growth rate by reducing stored keyframe frequency, while `--image_stride` reduces absolute frame count and total activation memory
- **Computational cost**: Keyframe intervals skip KV computation for intermediate frames, whereas stride values skip entire forward passes
- **Accuracy risks**: Large keyframe intervals risk stale context during rapid scene changes, while high stride values risk missing transient motions entirely

## Interaction Effects and Joint Tuning

These parameters interact significantly in production deployments. If you apply a high `--image_stride` (e.g., 4), you can safely increase `--keyframe_interval` because fewer total frames enter the pipeline, preventing KV‑cache overflow. Conversely, with `--stride 1` (dense sampling), you must use smaller keyframe intervals to avoid exhausting GPU memory.

The [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py) script (lines 64‑71) exposes both flags specifically for empirical testing of these interactions on your target hardware.

## Practical Configuration Examples

Run a balanced configuration for moderate sequences:

```bash
python -m lingbot_map.demo \
    --image_folder /path/to/images/ \
    --mode streaming \
    --keyframe_interval 4 \
    --stride 2 \
    --image_size 518 \
    --fps 30

```

Process very long sequences on limited GPU memory:

```bash
python -m lingbot_map.demo \
    --image_folder /path/to/long_sequence/ \
    --mode streaming \
    --keyframe_interval 8 \
    --stride 4 \
    --image_size 480

```

Enable flow‑based adaptive keyframe detection to protect accuracy during high‑motion segments:

```bash
python -m lingbot_map.demo \
    --image_folder /path/to/sequence/ \
    --mode streaming \
    --keyframe_interval 6 \
    --flow_threshold 0.5

```

## Summary

- `--keyframe_interval` manages internal KV‑cache frequency in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py), controlling memory growth during streaming inference by determining which frames retain their attention keys
- `--image_stride` (CLI: `--stride`) filters input data at the loading stage in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py), reducing the absolute number of frames processed and shrinking activation tensors
- High keyframe intervals risk temporal fidelity degradation during rapid scene changes, while aggressive image‑stride risks missing fast motion entirely
- These parameters must be tuned jointly: high stride permits larger intervals, while dense sampling requires smaller intervals to prevent cache overflow
- For limited GPU memory on long sequences, combine `--keyframe_interval 8` with `--stride 4` as validated in [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py)

## Frequently Asked Questions

### Can I use `--keyframe_interval` and `--image_stride` independently?

Yes, but they interact closely. You can set either parameter independently, but as implemented in the LingBot pipeline, high `--image_stride` values (which reduce input frames) allow you to safely increase `--keyframe_interval` without overwhelming the KV cache. Conversely, with `--stride 1` (dense sampling), you typically need smaller interval values to maintain memory bounds and prevent out‑of‑memory errors during long sequences.

### Which parameter has greater impact on GPU memory usage?

`--image_stride` typically provides larger absolute memory savings because it reduces the total tensor count before processing begins, whereas `--keyframe_interval` only optimizes the KV‑cache growth rate. According to the implementation in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py), the KV cache scales with the number of keyframes, but the overall activation memory scales with the total frames ingested, making stride the more aggressive memory reducer.

### How does the flow‑based keyframe detection interact with these settings?

When `--keyframe_interval` is greater than 1, the system in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py) (lines 530‑540) uses optical flow to override the fixed interval if motion exceeds `--flow_threshold`. This means rapid motion can force additional keyframes beyond the specified interval, protecting mapping accuracy at the cost of unpredictable memory usage spikes that violate the nominal interval budget.

### What are recommended values for real‑time laptop demos?

For real‑time processing on limited laptop hardware, use `--stride 2` or `--stride 4` to maintain interactive frame rates, combined with `--keyframe_interval 4` or `6`. This configuration, as exposed in [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py), keeps the KV cache manageable while skipping computationally expensive forward passes on intermediate frames, ensuring the pipeline runs within typical mobile GPU memory constraints.