# When to Use Windowed Inference Mode in LingBot-Map: Complete Guide

> Learn when to use windowed inference mode in LingBot-Map for long videos or limited GPU memory. Avoid KV-cache overflow and pose drift with this processing technique.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-28

---

**Enable windowed inference mode in LingBot-Map when processing video sequences longer than approximately 300 frames or when GPU memory is constrained, as it prevents KV-cache overflow and pose drift by processing the video in overlapping windows with periodically reset caches.**

LingBot-Map is an open-source visual localization system that processes video streams using a causal transformer architecture. During inference, the model maintains a KV cache for all past keyframes that grows linearly with sequence length, but the underlying model was only trained on sequences of approximately 320 frames (the RoPE range). The windowed inference mode, implemented in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), solves the scaling limitations by segmenting long videos into overlapping windows with independent cache resets.

## Why Windowed Inference Mode Exists

LingBot-Map's causal transformer retains a **KV cache** for every keyframe encountered during inference. While this enables efficient streaming processing, it creates two critical failure modes when sequences grow beyond the training distribution:

1. **KV-cache overflow** – The cache consumes GPU memory linearly with the number of frames. On long videos, this causes out-of-memory (OOM) errors as the cache grows unbounded.
2. **Pose drift** – The model has never learned to attend to more than 320 views. When the cache exceeds this range, the estimated camera trajectory collapses because the attention mechanism operates outside its trained distribution.

The `inference_windowed` method (lines 1022-1046 in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py)) addresses both issues by creating **independent windows** with fresh caches and handling **cross-window alignment** via scale frames and controlled overlap.

## When to Enable Windowed Inference Mode

### Processing Long Video Sequences (>300 Frames)

You must enable windowed mode when inference sequences exceed the model's RoPE training range of roughly 320 frames. According to the LingBot-Map source code, the README explicitly recommends windowed inference for "long sequences, >3000 frames" and for any scenario where you observe "pose collapse" (see [`README.md`](https://github.com/Robbyant/lingbot-map/blob/main/README.md), lines 44-52). Even sequences approaching 300 frames benefit from windowing to prevent boundary cases where the cache approaches the training limit.

### Limited GPU Memory Environments

Windowed inference bounds peak memory consumption by processing at most `window_size` keyframes simultaneously. Each window initializes a fresh KV cache, ensuring that memory usage remains constant regardless of total video length. This is essential when running on consumer GPUs with limited VRAM.

### Deterministic Memory Budgeting

When you need guaranteed memory limits, windowed mode allows you to fix `window_size` and `keyframe_interval` to calculate an exact upper bound on cache size. Unlike standard inference where memory grows with video length, windowed inference provides predictable resource consumption required for production deployment or edge devices.

### Fallback Attention Backends (Non-FlashInfer)

If your GPU does not support FlashInfer and falls back to the SDPA (Scaled Dot-Product Attention) implementation, memory consumption increases significantly. The README's "Performance & Memory" section notes that windowed mode becomes essential in this configuration to stay within VRAM constraints.

## Technical Architecture of Windowed Inference

### KV Cache Slots and Window Size

The `window_size` parameter does not count raw frames—it specifies the number of **KV-cache slots** (keyframes) maintained per window. As documented in lines 1042-1048 of [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py), the first `num_scale_frames` slots are reserved for scale frames used for cross-window alignment, while the remaining slots contain actual keyframes. This architecture ensures that each window maintains a fixed memory footprint regardless of input video length.

### Cross-Window Alignment and Overlap

To maintain global trajectory consistency, consecutive windows share pose information through overlapping keyframes. You can specify overlap using `overlap_keyframes`, which the implementation converts to an actual frame overlap calculated as at least `max(num_scale_frames, overlap_keyframes * keyframe_interval)` (see lines 1049-1062). The scale frames preserved at the beginning of each window facilitate alignment between the current window's coordinate system and the previous window's trajectory.

## Implementing Windowed Inference in Practice

### Command-Line Interface

The [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) script exposes windowed inference via CLI flags. The README recommends these defaults for long sequences:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 16 \
    --keyframe_interval 2

```

### Python API Usage

Call `inference_windowed` directly on a `GCTStream` instance for programmatic control:

```python
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

# Initialize model

model = GCTStream()
model.load_state_dict(torch.load("lingbot-map.pt"))
model.eval()

# Long video tensor: [S, 3, H, W]

images = torch.randn(5000, 3, 720, 1280)

# Run windowed inference with default parameters

out = model.inference_windowed(
    images,
    window_size=128,          # KV-cache slots per window

    overlap_keyframes=16,     # Shared keyframes between windows

    keyframe_interval=2,      # Cache every 2nd frame

    num_scale_frames=8,       # Scale frames for alignment

    output_device=torch.device("cpu")
)

# Results contain merged trajectory

print(out["pose_enc"].shape)  # [B, S, 9]

```

### Extreme Memory Constraints

For GPUs with severe memory limitations, reduce all cache-related parameters:

```python
out = model.inference_windowed(
    images,
    window_size=64,           # Smaller cache footprint

    overlap_keyframes=8,      # Minimal overlap

    keyframe_interval=4,      # Sparser keyframe sampling

    num_scale_frames=4,       # Fewer scale frames

    offload_to_cpu=True       # Immediate CPU offloading

)

```

## Summary

- **Windowed inference mode** resets the KV cache periodically to prevent unbounded memory growth and pose drift on long sequences.
- **Enable it** when processing videos longer than 300 frames, running on memory-constrained GPUs, or using fallback attention implementations.
- The `window_size` parameter controls keyframe slots (not raw frames), with reserved slots for scale frames that enable cross-window alignment.
- **Overlap parameters** ensure trajectory consistency between windows by sharing keyframe information across boundaries.
- Default parameters (`--window_size 128`, `--overlap_keyframes 16`) work for most scenarios, with smaller values available for extreme memory constraints.

## Frequently Asked Questions

### What is the relationship between window_size and GPU memory usage?

The `window_size` parameter directly determines the maximum size of the KV cache allocated during inference. Since the cache stores key and value tensors for every attention head and every keyframe slot, peak GPU memory scales linearly with `window_size`. Setting `window_size=64` uses approximately half the cache memory of `window_size=128`, making it the primary lever for fitting long videos into limited VRAM.

### How do I choose the right overlap_keyframes value?

The `overlap_keyframes` value balances trajectory consistency against computational overhead. Larger overlaps (e.g., 32 keyframes) provide smoother stitching between windows but increase redundant computation. For most applications, the default of 16 keyframes provides sufficient alignment, but increase this value if you observe discontinuities in the estimated camera trajectory at window boundaries.

### Can windowed inference recover from pose drift within a single window?

No. If the model begins drifting within a single window, the windowed mechanism cannot correct it until the next cache reset. However, because each window initializes a fresh cache, drift accumulated in previous windows is discarded. Keep `window_size` below 300 keyframes to ensure each individual window stays within the model's trained RoPE range and avoids intra-window drift.

### Is windowed inference slower than standard inference?

Yes, windowed inference incurs a small computational overhead due to overlapping processing and cross-window alignment calculations. However, this penalty is negligible compared to the alternative—processing long sequences without windows causes OOM errors that halt inference entirely. The overlap frames represent redundant computation, so minimizing `overlap_keyframes` while maintaining trajectory quality optimizes the speed/accuracy tradeoff.