When to Use Windowed Inference Mode in LingBot-Map: A Complete Guide

Use windowed inference mode when processing video sequences longer than approximately 300 frames or when GPU memory is limited, as it prevents KV-cache overflow and pose drift by resetting the cache in overlapping windows.

LingBot-Map processes video streams using a causal transformer that maintains a KV cache for all past keyframes. During training, the model sees at most ≈ 320 frames (the RoPE range). When inference runs on longer sequences, the cache grows linearly and can cause out-of-memory errors or trajectory collapse. The windowed inference mode solves this by processing videos in overlapping chunks with periodic cache resets.

Why Windowed Inference Matters

The causal transformer architecture in LingBot-Map keeps a running cache of key and value tensors for every keyframe processed. This design enables efficient streaming inference but creates two critical bottlenecks on long sequences:

  1. KV-cache overflow – Memory usage grows linearly with sequence length, eventually exceeding GPU VRAM and causing OOM errors.
  2. Pose drift – The model has never learned to attend to more than 320 views during training, so estimated camera trajectories collapse when the cache exceeds this range.

According to the source code in lingbot_map/models/gct_stream_window_v2.py, the inference_windowed method addresses these issues by creating independent windows with fresh caches and handling cross-window alignment via scale frames and overlap (see lines 1022–1046).

When to Enable Windowed Inference Mode

Enable windowed inference in the following situations:

  • Sequences longer than ≈ 300 frames – The README explicitly recommends windowed mode for "long sequences, > 3000 frames" and for any sequence where you observe pose collapse (see README.md lines 44–52).
  • Limited GPU memory – Each window processes at most window_size keyframes, bounding peak memory usage regardless of total video length.
  • Deterministic memory requirements – Fixing window_size and keyframe_interval guarantees an upper bound on cache size for deployment on fixed hardware.
  • Fallback attention implementations – When running without FlashInfer (using SDPA fallback), memory consumption is significantly higher, making windowed mode essential.
  • Speed vs. accuracy trade-offs – Increasing keyframe_interval skips caching many frames while maintaining predictions; windowed mode keeps this optimization within memory limits.

How Windowed Inference Works

The implementation in lingbot_map/models/gct_stream_window_v2.py processes videos through overlapping temporal windows rather than a single continuous stream.

Independent Windows with Fresh Caches

Each window initializes a new KV cache, preventing the linear growth that causes OOM errors. The inference_windowed method (lines 1022–1046) manages the stitching of trajectories across these disjoint windows.

Scale Frames and Overlap

Cross-window alignment relies on scale frames and overlapping keyframes:

  • The first num_scale_frames slots in each window are reserved for scale frames that anchor the coordinate system.
  • Overlap between consecutive windows (controlled by overlap_keyframes or overlap_size) guarantees that pose information propagates across window boundaries, maintaining global trajectory consistency.

Window Size Interpretation

The window_size parameter counts KV-cache slots (keyframes) rather than raw frames. As documented in lines 1042–1048 of the source file, the cache layout reserves the first num_scale_frames slots for scale frames, with the remainder allocated to keyframes.

Overlap Calculation

Overlap can be specified in keyframe units via overlap_keyframes. The code converts this to actual frame overlap using the formula:


actual_overlap = max(num_scale_frames, overlap_keyframes * keyframe_interval)

This logic appears in lines 1049–1062 of gct_stream_window_v2.py, ensuring sufficient temporal context for alignment regardless of keyframe sampling rate.

Configuration Parameters

Control memory usage and accuracy through these key arguments to inference_windowed:

  • window_size – Number of KV-cache slots (keyframes) per window. Default is 128; reduce to 64 or 32 for extreme memory constraints.
  • overlap_keyframes – Number of keyframes shared between consecutive windows. Default is 16; higher values improve alignment at the cost of increased computation.
  • keyframe_interval – Sampling rate for the KV cache (e.g., every 2nd or 4th frame). Higher intervals reduce memory but may decrease accuracy.
  • num_scale_frames – Reserved slots for cross-window alignment frames. Default is 8; can be reduced to 4 for memory savings.

Implementation Examples

Command-Line Interface

Process long videos using the windowed mode flags demonstrated in demo.py and documented in the README:

python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 16 \
    --keyframe_interval 2

Python API

Call inference_windowed directly for programmatic control:

import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

# Load model

model = GCTStream()
model.load_state_dict(torch.load("lingbot-map.pt"))
model.eval()

# Simulate a long sequence: 5000 frames

images = torch.randn(5000, 3, 720, 1280)

# Run windowed inference with default parameters

out = model.inference_windowed(
    images,
    window_size=128,          # KV-cache slots per window

    overlap_keyframes=16,     # Shared keyframes between windows

    keyframe_interval=2,      # Cache every 2nd frame

    num_scale_frames=8,       # Scale frames for alignment

    output_device=torch.device("cpu")  # Offload to save GPU RAM

)

print(out["pose_enc"].shape)  # [B, S, 9]

Memory-Constrained Configuration

For GPUs with limited VRAM, reduce all cache dimensions:

out = model.inference_windowed(
    images,
    window_size=64,           # Smaller KV cache per window

    overlap_keyframes=8,      # Minimal overlap

    keyframe_interval=4,      # Sparse keyframe sampling

    num_scale_frames=4,       # Fewer scale frames

    offload_to_cpu=True       # Immediate CPU offload

)

Summary

  • Enable windowed inference mode for any sequence longer than ~300 frames to prevent KV-cache overflow and pose drift.
  • The mode processes videos in independent windows with periodic cache resets, using overlapping keyframes and scale frames to maintain trajectory consistency.
  • Key parameters include window_size (cache slots per window), overlap_keyframes (shared context), and keyframe_interval (sampling rate).
  • Default values (window_size=128, overlap_keyframes=16, keyframe_interval=2) work well for most GPUs; reduce these for extreme memory constraints.
  • The implementation resides in lingbot_map/models/gct_stream_window_v2.py, specifically the inference_windowed method (lines 1022–1062).

Frequently Asked Questions

What is the maximum sequence length without windowed mode?

Without windowed mode, practical limits are approximately 300–320 frames, matching the RoPE training range. Beyond this, the model exhibits pose drift and memory usage becomes unbounded, potentially causing GPU OOM errors.

How does windowed mode affect reconstruction accuracy?

Windowed mode maintains high accuracy through cross-window alignment using scale frames and overlapping keyframes. While aggressive settings (very small window_size or high keyframe_interval) can degrade results, the default parameters (window_size=128, overlap_keyframes=16) preserve trajectory consistency comparable to full-sequence inference.

Can I use windowed inference on short videos?

Yes, but it is unnecessary for sequences under 300 frames. The overhead of window management provides no benefit for short clips and may slightly increase processing time due to overlapping computations between windows.

What is the difference between overlap_keyframes and overlap_size?

overlap_keyframes specifies overlap in keyframe units (e.g., 16 keyframes), while overlap_size specifies overlap in raw frames. The code converts overlap_keyframes to actual frames using max(num_scale_frames, overlap_keyframes * keyframe_interval) to ensure sufficient temporal context for alignment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →