# How to Manage Overlapping Windows for Sequences Over 3000 Frames with Windowed Inference

> Learn to manage overlapping windows for sequences over 3000 frames using windowed inference. Align keyframes accurately across boundaries for efficient video stream processing.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-27

---

**LingBot-Map processes arbitrarily long video streams by splitting them into overlapping windows with fresh KV-caches per window, where the effective overlap must be calculated in keyframes rather than raw frames to ensure accurate alignment across window boundaries.**

LingBot-Map reconstructs long video sequences through **windowed inference**, a technique that divides inputs into manageable chunks to bound GPU memory usage. For sequences exceeding 3000 frames, configuring the overlap between windows correctly is critical to prevent trajectory misalignment and Out-Of-Memory (OOM) errors. The implementation resides in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), which handles KV-cache management, keyframe sampling, and pairwise alignment stitching.

## Understanding Window Size and Keyframe Interval

LingBot-Map uses a sparse KV-cache that only stores **keyframes**, not every frame. Three parameters define how many actual frames each window covers:

- **`window_size`**: The number of keyframe slots allocated in the KV-cache per window (includes scale frames).
- **`keyframe_interval`**: Every N-th frame after the scale phase is retained as a keyframe (N=1 means every frame, N=2 means every second frame).
- **`num_scale_frames`** (default: `8`): The first M frames in each window receiving bidirectional attention, where no keyframe interval is applied.

The actual frame count per window is calculated as:

```python
actual_frames = num_scale_frames + (window_size - num_scale_frames) * keyframe_interval

```

For example, with `window_size=128`, `keyframe_interval=2`, and the default `num_scale_frames=8`, each window covers `8 + (120 * 2) = 248` actual frames.

## How Overlap is Calculated in the Source Code

In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), the `inference_windowed` method computes effective overlap between lines 1093–1110. The logic ensures the overlap region always contains at least one **paired keyframe** required by `_pairwise_alignment`:

```python

# Resolve overlap in *actual frames*

if overlap_keyframes is not None:
    kf = max(keyframe_interval, 1)                     # guard against 0

    eff_overlap = max(ws, overlap_keyframes * kf)      # ws = num_scale_frames

elif overlap_size is not None:
    eff_overlap = overlap_size
else:
    eff_overlap = ws                                   # default = scale frames

eff_overlap = min(eff_overlap, S - 1) if S > 1 else 0

```

**Critical insight**: When `keyframe_interval > 1`, specifying `--overlap_size` in raw frames might create an overlap with **no keyframes**, causing alignment to fall back to a less accurate first-frame heuristic. Always use `--overlap_keyframes` instead; the code multiplies this by `keyframe_interval` and takes the maximum with `num_scale_frames` to guarantee paired keyframes exist in the overlap zone.

## Configuring Window Strategy for 3000+ Frames

Follow this workflow to process sequences longer than 3000 frames without memory exhaustion or alignment drift:

1. **Select `keyframe_interval`** based on GPU memory. A value of `2` halves the KV-cache size compared to storing every frame.

2. **Choose `window_size`** (64–128 is typical). With `keyframe_interval=2`, `window_size=128` yields 248 actual frames per window.

3. **Set `overlap_keyframes`** to ensure sufficient alignment data. The effective overlap in actual frames becomes:
   
   ```python
   eff_overlap = max(8, overlap_keyframes * keyframe_interval)
   ```

   
   Setting `overlap_keyframes=16` with `keyframe_interval=2` produces a 32-frame overlap, providing 16 keyframes for the alignment routine to estimate similarity transforms.

4. **Calculate window count**:
   
   ```python
   window_stride = actual_frames_per_window - eff_overlap
   windows_needed = ceil(total_frames / window_stride)
   ```

   
   For 3000 frames with 248-frame windows and 32-frame overlap: `ceil(3000 / 216) = 14` windows.

## CLI and Python API Examples

### Command-Line Usage

Run [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) with the windowed inference mode for a 3000-frame video:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path my_long_video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 16 \
    --keyframe_interval 2 \
    --output_device cpu

```

* `--mode windowed` triggers `inference_windowed`.
* `--output_device cpu` offloads per-frame predictions to system RAM, preventing GPU OOM during long sequences.

### Python API Implementation

For programmatic control, instantiate `GCTStream` and call `inference_windowed`:

```python
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

model = GCTStream(
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    enable_3d_rope=True,
).eval().to('cuda')

# imgs shape: [B, S, 3, H, W] where S=3000

imgs = torch.load('my_frames.pt')  

pred = model.inference_windowed(
    images=imgs,
    window_size=128,
    overlap_keyframes=16,
    keyframe_interval=2,
    output_device=torch.device('cpu')
)

poses = pred['pose_enc']  # shape [1, 3000, 9]

```

## Troubleshooting Common Pitfalls

| Symptom | Root Cause | Solution |
|---------|------------|----------|
| **Mis-aligned trajectories** at window boundaries | `eff_overlap` smaller than `num_scale_frames` or no paired keyframes in overlap region | Increase `--overlap_keyframes` (e.g., 8→16) or decrease `--keyframe_interval` |
| **GPU OOM** despite windowed mode | `window_size` too large or `keyframe_interval` too small | Reduce `window_size` to 64 or increase `keyframe_interval` to 4 |
| **Severe performance degradation** | Per-frame tensors accumulating on GPU | Specify `--output_device cpu` to move results off-GPU immediately |
| **Alignment fallback to first frame** | `overlap_keyframes=0` while `keyframe_interval>1` | Set `--overlap_keyframes` ≥ `num_scale_frames / keyframe_interval` |

## Summary

- **LingBot-Map** handles sequences over 3000 frames via `inference_windowed` in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py), which splits video into overlapping chunks with independent KV-caches.
- **Always specify overlap in keyframes** (`--overlap_keyframes`) rather than raw pixels or frames when `keyframe_interval > 1` to ensure the alignment routine finds paired keyframes.
- **Effective overlap** is calculated as `max(num_scale_frames, overlap_keyframes * keyframe_interval)`, guaranteeing at least one shared keyframe between windows.
- **Memory management** relies on the sparsity of the KV-cache; increasing `keyframe_interval` reduces memory linearly but requires proportional increases in overlap keyframes to maintain alignment accuracy.
- **Offload outputs** to CPU via `--output_device cpu` to prevent GPU memory accumulation during inference of very long sequences.

## Frequently Asked Questions

### What is the difference between `overlap_size` and `overlap_keyframes` in LingBot-Map?

`overlap_size` specifies the overlap in raw frame count, while `overlap_keyframes` specifies it in terms of keyframe indices. When `keyframe_interval > 1`, using `overlap_size` risks creating an overlap region containing no actual keyframes, forcing `_pairwise_alignment` to use inaccurate first-frame matching. The `--overlap_keyframes` flag ensures the code multiplies by `keyframe_interval` and enforces a minimum overlap of `num_scale_frames`, guaranteeing valid alignment anchors.

### Why does my inference fail with mis-aligned trajectories at window boundaries?

Mis-alignment occurs when the effective overlap contains fewer frames than `num_scale_frames` (default 8) or lacks keyframes present in both windows. According to the source code in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py), the overlap must be at least `max(num_scale_frames, overlap_keyframes * keyframe_interval)` actual frames. Increase `--overlap_keyframes` until the overlap region contains sufficient keyframes for the similarity transform estimation.

### How do I prevent GPU Out-Of-Memory errors when processing 3000+ frames?

Reduce the KV-cache footprint by increasing `--keyframe_interval` (e.g., to 2 or 4) to store fewer frames per window, or decrease `--window_size` to process fewer keyframes simultaneously. Additionally, set `--output_device cpu` to move prediction tensors off the GPU immediately after each window processes, preventing accumulation of per-frame pose encodings in video memory.

### Can I use windowed inference with the Python API instead of the CLI?

Yes. Import `GCTStream` from `lingbot_map.models.gct_stream_window_v2`, initialize the model with desired cache settings, and call the `inference_windowed` method with explicit `window_size`, `overlap_keyframes`, and `keyframe_interval` arguments. Pass `output_device=torch.device('cpu')` to manage memory for long sequences programmatically without modifying global device settings.