When to Use Windowed Inference Mode in LingBot-Map: Complete Guide
Enable windowed inference mode in LingBot-Map when processing video sequences longer than approximately 300 frames or when GPU memory is constrained, as it prevents KV-cache overflow and pose drift by processing the video in overlapping windows with periodically reset caches.
LingBot-Map is an open-source visual localization system that processes video streams using a causal transformer architecture. During inference, the model maintains a KV cache for all past keyframes that grows linearly with sequence length, but the underlying model was only trained on sequences of approximately 320 frames (the RoPE range). The windowed inference mode, implemented in lingbot_map/models/gct_stream_window_v2.py, solves the scaling limitations by segmenting long videos into overlapping windows with independent cache resets.
Why Windowed Inference Mode Exists
LingBot-Map's causal transformer retains a KV cache for every keyframe encountered during inference. While this enables efficient streaming processing, it creates two critical failure modes when sequences grow beyond the training distribution:
- KV-cache overflow – The cache consumes GPU memory linearly with the number of frames. On long videos, this causes out-of-memory (OOM) errors as the cache grows unbounded.
- Pose drift – The model has never learned to attend to more than 320 views. When the cache exceeds this range, the estimated camera trajectory collapses because the attention mechanism operates outside its trained distribution.
The inference_windowed method (lines 1022-1046 in lingbot_map/models/gct_stream_window_v2.py) addresses both issues by creating independent windows with fresh caches and handling cross-window alignment via scale frames and controlled overlap.
When to Enable Windowed Inference Mode
Processing Long Video Sequences (>300 Frames)
You must enable windowed mode when inference sequences exceed the model's RoPE training range of roughly 320 frames. According to the LingBot-Map source code, the README explicitly recommends windowed inference for "long sequences, >3000 frames" and for any scenario where you observe "pose collapse" (see README.md, lines 44-52). Even sequences approaching 300 frames benefit from windowing to prevent boundary cases where the cache approaches the training limit.
Limited GPU Memory Environments
Windowed inference bounds peak memory consumption by processing at most window_size keyframes simultaneously. Each window initializes a fresh KV cache, ensuring that memory usage remains constant regardless of total video length. This is essential when running on consumer GPUs with limited VRAM.
Deterministic Memory Budgeting
When you need guaranteed memory limits, windowed mode allows you to fix window_size and keyframe_interval to calculate an exact upper bound on cache size. Unlike standard inference where memory grows with video length, windowed inference provides predictable resource consumption required for production deployment or edge devices.
Fallback Attention Backends (Non-FlashInfer)
If your GPU does not support FlashInfer and falls back to the SDPA (Scaled Dot-Product Attention) implementation, memory consumption increases significantly. The README's "Performance & Memory" section notes that windowed mode becomes essential in this configuration to stay within VRAM constraints.
Technical Architecture of Windowed Inference
KV Cache Slots and Window Size
The window_size parameter does not count raw frames—it specifies the number of KV-cache slots (keyframes) maintained per window. As documented in lines 1042-1048 of gct_stream_window_v2.py, the first num_scale_frames slots are reserved for scale frames used for cross-window alignment, while the remaining slots contain actual keyframes. This architecture ensures that each window maintains a fixed memory footprint regardless of input video length.
Cross-Window Alignment and Overlap
To maintain global trajectory consistency, consecutive windows share pose information through overlapping keyframes. You can specify overlap using overlap_keyframes, which the implementation converts to an actual frame overlap calculated as at least max(num_scale_frames, overlap_keyframes * keyframe_interval) (see lines 1049-1062). The scale frames preserved at the beginning of each window facilitate alignment between the current window's coordinate system and the previous window's trajectory.
Implementing Windowed Inference in Practice
Command-Line Interface
The demo.py script exposes windowed inference via CLI flags. The README recommends these defaults for long sequences:
python demo.py \
--model_path /path/to/lingbot-map.pt \
--video_path video.mp4 \
--fps 10 \
--mode windowed \
--window_size 128 \
--overlap_keyframes 16 \
--keyframe_interval 2
Python API Usage
Call inference_windowed directly on a GCTStream instance for programmatic control:
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream
# Initialize model
model = GCTStream()
model.load_state_dict(torch.load("lingbot-map.pt"))
model.eval()
# Long video tensor: [S, 3, H, W]
images = torch.randn(5000, 3, 720, 1280)
# Run windowed inference with default parameters
out = model.inference_windowed(
images,
window_size=128, # KV-cache slots per window
overlap_keyframes=16, # Shared keyframes between windows
keyframe_interval=2, # Cache every 2nd frame
num_scale_frames=8, # Scale frames for alignment
output_device=torch.device("cpu")
)
# Results contain merged trajectory
print(out["pose_enc"].shape) # [B, S, 9]
Extreme Memory Constraints
For GPUs with severe memory limitations, reduce all cache-related parameters:
out = model.inference_windowed(
images,
window_size=64, # Smaller cache footprint
overlap_keyframes=8, # Minimal overlap
keyframe_interval=4, # Sparser keyframe sampling
num_scale_frames=4, # Fewer scale frames
offload_to_cpu=True # Immediate CPU offloading
)
Summary
- Windowed inference mode resets the KV cache periodically to prevent unbounded memory growth and pose drift on long sequences.
- Enable it when processing videos longer than 300 frames, running on memory-constrained GPUs, or using fallback attention implementations.
- The
window_sizeparameter controls keyframe slots (not raw frames), with reserved slots for scale frames that enable cross-window alignment. - Overlap parameters ensure trajectory consistency between windows by sharing keyframe information across boundaries.
- Default parameters (
--window_size 128,--overlap_keyframes 16) work for most scenarios, with smaller values available for extreme memory constraints.
Frequently Asked Questions
What is the relationship between window_size and GPU memory usage?
The window_size parameter directly determines the maximum size of the KV cache allocated during inference. Since the cache stores key and value tensors for every attention head and every keyframe slot, peak GPU memory scales linearly with window_size. Setting window_size=64 uses approximately half the cache memory of window_size=128, making it the primary lever for fitting long videos into limited VRAM.
How do I choose the right overlap_keyframes value?
The overlap_keyframes value balances trajectory consistency against computational overhead. Larger overlaps (e.g., 32 keyframes) provide smoother stitching between windows but increase redundant computation. For most applications, the default of 16 keyframes provides sufficient alignment, but increase this value if you observe discontinuities in the estimated camera trajectory at window boundaries.
Can windowed inference recover from pose drift within a single window?
No. If the model begins drifting within a single window, the windowed mechanism cannot correct it until the next cache reset. However, because each window initializes a fresh cache, drift accumulated in previous windows is discarded. Keep window_size below 300 keyframes to ensure each individual window stays within the model's trained RoPE range and avoids intra-window drift.
Is windowed inference slower than standard inference?
Yes, windowed inference incurs a small computational overhead due to overlapping processing and cross-window alignment calculations. However, this penalty is negligible compared to the alternative—processing long sequences without windows causes OOM errors that halt inference entirely. The overlap frames represent redundant computation, so minimizing overlap_keyframes while maintaining trajectory quality optimizes the speed/accuracy tradeoff.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →