# When to Use Windowed Inference Mode in LingBot-Map: A Complete Guide

> Learn when to use windowed inference mode in LingBot-Map for long videos or limited GPU memory. Prevent KV-cache overflow and pose drift effectively.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-27

---

**Use windowed inference mode when processing video sequences longer than approximately 300 frames or when GPU memory is limited, as it prevents KV-cache overflow and pose drift by resetting the cache in overlapping windows.**

LingBot-Map processes video streams using a causal transformer that maintains a **KV cache** for all past keyframes. During training, the model sees at most **≈ 320 frames** (the RoPE range). When inference runs on longer sequences, the cache grows linearly and can cause out-of-memory errors or trajectory collapse. The windowed inference mode solves this by processing videos in overlapping chunks with periodic cache resets.

## Why Windowed Inference Matters

The causal transformer architecture in LingBot-Map keeps a running cache of key and value tensors for every keyframe processed. This design enables efficient streaming inference but creates two critical bottlenecks on long sequences:

1. **KV-cache overflow** – Memory usage grows linearly with sequence length, eventually exceeding GPU VRAM and causing OOM errors.
2. **Pose drift** – The model has never learned to attend to more than 320 views during training, so estimated camera trajectories collapse when the cache exceeds this range.

According to the source code in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), the `inference_windowed` method addresses these issues by creating **independent windows** with fresh caches and handling **cross-window alignment** via scale frames and overlap (see lines 1022–1046).

## When to Enable Windowed Inference Mode

Enable windowed inference in the following situations:

- **Sequences longer than ≈ 300 frames** – The README explicitly recommends windowed mode for "long sequences, > 3000 frames" and for any sequence where you observe pose collapse (see README.md lines 44–52).
- **Limited GPU memory** – Each window processes at most `window_size` keyframes, bounding peak memory usage regardless of total video length.
- **Deterministic memory requirements** – Fixing `window_size` and `keyframe_interval` guarantees an upper bound on cache size for deployment on fixed hardware.
- **Fallback attention implementations** – When running without FlashInfer (using SDPA fallback), memory consumption is significantly higher, making windowed mode essential.
- **Speed vs. accuracy trade-offs** – Increasing `keyframe_interval` skips caching many frames while maintaining predictions; windowed mode keeps this optimization within memory limits.

## How Windowed Inference Works

The implementation in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) processes videos through overlapping temporal windows rather than a single continuous stream.

### Independent Windows with Fresh Caches

Each window initializes a new KV cache, preventing the linear growth that causes OOM errors. The `inference_windowed` method (lines 1022–1046) manages the stitching of trajectories across these disjoint windows.

### Scale Frames and Overlap

Cross-window alignment relies on **scale frames** and overlapping keyframes:

- The first `num_scale_frames` slots in each window are reserved for scale frames that anchor the coordinate system.
- Overlap between consecutive windows (controlled by `overlap_keyframes` or `overlap_size`) guarantees that pose information propagates across window boundaries, maintaining global trajectory consistency.

### Window Size Interpretation

The `window_size` parameter counts **KV-cache slots** (keyframes) rather than raw frames. As documented in lines 1042–1048 of the source file, the cache layout reserves the first `num_scale_frames` slots for scale frames, with the remainder allocated to keyframes.

### Overlap Calculation

Overlap can be specified in keyframe units via `overlap_keyframes`. The code converts this to actual frame overlap using the formula:

```

actual_overlap = max(num_scale_frames, overlap_keyframes * keyframe_interval)

```

This logic appears in lines 1049–1062 of [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py), ensuring sufficient temporal context for alignment regardless of keyframe sampling rate.

## Configuration Parameters

Control memory usage and accuracy through these key arguments to `inference_windowed`:

- **`window_size`** – Number of KV-cache slots (keyframes) per window. Default is 128; reduce to 64 or 32 for extreme memory constraints.
- **`overlap_keyframes`** – Number of keyframes shared between consecutive windows. Default is 16; higher values improve alignment at the cost of increased computation.
- **`keyframe_interval`** – Sampling rate for the KV cache (e.g., every 2nd or 4th frame). Higher intervals reduce memory but may decrease accuracy.
- **`num_scale_frames`** – Reserved slots for cross-window alignment frames. Default is 8; can be reduced to 4 for memory savings.

## Implementation Examples

### Command-Line Interface

Process long videos using the windowed mode flags demonstrated in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and documented in the README:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 16 \
    --keyframe_interval 2

```

### Python API

Call `inference_windowed` directly for programmatic control:

```python
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

# Load model

model = GCTStream()
model.load_state_dict(torch.load("lingbot-map.pt"))
model.eval()

# Simulate a long sequence: 5000 frames

images = torch.randn(5000, 3, 720, 1280)

# Run windowed inference with default parameters

out = model.inference_windowed(
    images,
    window_size=128,          # KV-cache slots per window

    overlap_keyframes=16,     # Shared keyframes between windows

    keyframe_interval=2,      # Cache every 2nd frame

    num_scale_frames=8,       # Scale frames for alignment

    output_device=torch.device("cpu")  # Offload to save GPU RAM

)

print(out["pose_enc"].shape)  # [B, S, 9]

```

### Memory-Constrained Configuration

For GPUs with limited VRAM, reduce all cache dimensions:

```python
out = model.inference_windowed(
    images,
    window_size=64,           # Smaller KV cache per window

    overlap_keyframes=8,      # Minimal overlap

    keyframe_interval=4,      # Sparse keyframe sampling

    num_scale_frames=4,       # Fewer scale frames

    offload_to_cpu=True       # Immediate CPU offload

)

```

## Summary

- Enable **windowed inference mode** for any sequence longer than ~300 frames to prevent KV-cache overflow and pose drift.
- The mode processes videos in independent windows with periodic cache resets, using overlapping keyframes and scale frames to maintain trajectory consistency.
- Key parameters include `window_size` (cache slots per window), `overlap_keyframes` (shared context), and `keyframe_interval` (sampling rate).
- Default values (`window_size=128`, `overlap_keyframes=16`, `keyframe_interval=2`) work well for most GPUs; reduce these for extreme memory constraints.
- The implementation resides in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), specifically the `inference_windowed` method (lines 1022–1062).

## Frequently Asked Questions

### What is the maximum sequence length without windowed mode?

Without windowed mode, practical limits are approximately **300–320 frames**, matching the RoPE training range. Beyond this, the model exhibits pose drift and memory usage becomes unbounded, potentially causing GPU OOM errors.

### How does windowed mode affect reconstruction accuracy?

Windowed mode maintains high accuracy through **cross-window alignment** using scale frames and overlapping keyframes. While aggressive settings (very small `window_size` or high `keyframe_interval`) can degrade results, the default parameters (`window_size=128`, `overlap_keyframes=16`) preserve trajectory consistency comparable to full-sequence inference.

### Can I use windowed inference on short videos?

Yes, but it is unnecessary for sequences under **300 frames**. The overhead of window management provides no benefit for short clips and may slightly increase processing time due to overlapping computations between windows.

### What is the difference between `overlap_keyframes` and `overlap_size`?

`overlap_keyframes` specifies overlap in **keyframe units** (e.g., 16 keyframes), while `overlap_size` specifies overlap in **raw frames**. The code converts `overlap_keyframes` to actual frames using `max(num_scale_frames, overlap_keyframes * keyframe_interval)` to ensure sufficient temporal context for alignment.