# Understanding the overlap_keyframes Parameter in Windowed Inference for LingBot-MAP

> Learn how overlap_keyframes parameter ensures temporal consistency and efficient KV-cache reuse in LingBot-MAP windowed inference for improved 3-D reconstruction.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: internals
- Published: 2026-07-26

---

**The `overlap_keyframes` parameter specifies how many keyframes are shared between consecutive windows in LingBot-MAP's windowed inference mode, ensuring temporal consistency in 3-D reconstruction while enabling efficient KV-cache reuse.**

LingBot-MAP processes long video sequences through **windowed inference**, a technique that divides footage into manageable chunks to handle memory constraints. The `overlap_keyframes` argument in the `inference_windowed` method determines the overlap strategy between these windows, directly impacting reconstruction continuity and computational efficiency.

## What is the overlap_keyframes Parameter?

In the LingBot-MAP architecture, windowed inference addresses the challenge of processing extended videos that exceed GPU memory limits. Instead of processing an entire sequence at once, the system divides the video into windows of a specified size.

The **`overlap_keyframes`** parameter controls how many frames from the end of one window are reused as the beginning of the next window. These overlapping keyframes serve as anchor points that maintain spatial context and geometric consistency across window boundaries.

## Implementation in the Source Code

The parameter is defined in the `inference_windowed` method of the `GCTStream` class, located in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py). The method signature reveals how the argument integrates into the inference pipeline:

```python
def inference_windowed(
    self,
    images,
    depth_maps,
    intrinsics,
    poses,
    window_size,
    overlap_keyframes=0,
    keyframe_interval=1,
    ...):

```

When `inference_windowed` executes, it uses `overlap_keyframes` to determine the stride between consecutive windows. A positive value causes the next window to start earlier in the sequence, reusing the specified number of keyframes from the previous window's tail.

## Impact of Different Configuration Values

The value assigned to `overlap_keyframes` fundamentally changes the behavior of the reconstruction pipeline.

### Setting overlap_keyframes to 0

When **`overlap_keyframes = 0`**, windows are strictly disjoint. Each window starts where the previous ended, with no shared frames between boundaries. This configuration:

- Eliminates temporal redundancy but risks geometric discontinuities at window boundaries
- Requires full recomputation of geometry for every frame
- Increases memory pressure due to lack of KV-cache reuse between windows

### Setting overlap_keyframes Greater Than 0

When **`overlap_keyframes > 0`**, the specified number of keyframes from the end of the current window become the first frames of the next window. This approach:

- Preserves spatial context across window transitions, ensuring smooth 3-D reconstruction
- Allows the KV-cache to persist for overlapping frames, reducing latency and computational overhead
- Maintains consistency in camera pose estimation across sequential chunks

## Practical Usage Examples

You can configure `overlap_keyframes` through both the command-line interface and programmatically via the Python API.

### Command-Line Interface

The [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) script exposes the parameter for windowed inference mode:

```bash
python demo.py \
    --model_path checkpoints/lingbot_map.pt \
    --video_path my_video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 64 \
    --overlap_keyframes 8

```

### Python API Implementation

For custom pipelines, instantiate `GCTStream` and call `inference_windowed` directly:

```python
from lingbot_map.models.gct_stream_window import GCTStream

model = GCTStream(
    img_size=518,
    patch_size=14,
    max_frame_num=1000,
    kv_cache_sliding_window=True,
    kv_cache_scale_frames=4,
    # other args …

)

# images, depth, intrinsics, poses are pre-processed tensors

preds = model.inference_windowed(
    images,
    depth_maps,
    intrinsics,
    poses,
    window_size=64,
    overlap_keyframes=8,          # share 8 keyframes between windows

    keyframe_interval=6,
)

```

The parameter also propagates through the benchmark wrapper in [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py), ensuring consistent behavior across evaluation scenarios.

## Summary

- The **`overlap_keyframes`** parameter in LingBot-MAP controls keyframe sharing between consecutive inference windows
- Setting the value to **0** creates disjoint windows, risking reconstruction discontinuities
- Positive values enable **KV-cache reuse** and maintain **3-D reconstruction consistency** across window boundaries
- The implementation resides in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) within the `inference_windowed` method
- Both CLI ([`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py)) and Python API support configuration of this parameter for windowed video processing

## Frequently Asked Questions

### What happens if I set overlap_keyframes to 0?

Setting `overlap_keyframes` to 0 processes windows as completely separate chunks with no shared frames. According to the LingBot-MAP source code, this eliminates temporal continuity between windows, potentially causing geometric discontinuities in the 3-D reconstruction and preventing KV-cache reuse, which increases memory consumption.

### How does overlap_keyframes affect memory usage?

Positive values for `overlap_keyframes` reduce memory pressure by allowing the KV-cache to persist across window boundaries. When keyframes overlap, the model reuses cached attention computations for those frames rather than recomputing geometry from scratch, significantly lowering latency for long video sequences.

### Where is the overlap_keyframes parameter defined in the codebase?

The parameter is defined in the `inference_windowed` method signature within [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) at approximately line 1022. It also appears in the alternative implementation file [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) and is forwarded through the benchmark wrapper at [`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py).

### Can overlap_keyframes be larger than window_size?

No, `overlap_keyframes` must be smaller than `window_size` to maintain valid window processing logic. If the overlap equaled or exceeded the window size, the sliding window would fail to advance through the sequence, causing infinite loops or processing errors in the `GCTStream` inference pipeline.