How to Manage Overlapping Windows for Sequences Over 3000 Frames with Windowed Inference

LingBot-Map processes arbitrarily long video streams by splitting them into overlapping windows with fresh KV-caches per window, where the effective overlap must be calculated in keyframes rather than raw frames to ensure accurate alignment across window boundaries.

LingBot-Map reconstructs long video sequences through windowed inference, a technique that divides inputs into manageable chunks to bound GPU memory usage. For sequences exceeding 3000 frames, configuring the overlap between windows correctly is critical to prevent trajectory misalignment and Out-Of-Memory (OOM) errors. The implementation resides in lingbot_map/models/gct_stream_window_v2.py, which handles KV-cache management, keyframe sampling, and pairwise alignment stitching.

Understanding Window Size and Keyframe Interval

LingBot-Map uses a sparse KV-cache that only stores keyframes, not every frame. Three parameters define how many actual frames each window covers:

  • window_size: The number of keyframe slots allocated in the KV-cache per window (includes scale frames).
  • keyframe_interval: Every N-th frame after the scale phase is retained as a keyframe (N=1 means every frame, N=2 means every second frame).
  • num_scale_frames (default: 8): The first M frames in each window receiving bidirectional attention, where no keyframe interval is applied.

The actual frame count per window is calculated as:

actual_frames = num_scale_frames + (window_size - num_scale_frames) * keyframe_interval

For example, with window_size=128, keyframe_interval=2, and the default num_scale_frames=8, each window covers 8 + (120 * 2) = 248 actual frames.

How Overlap is Calculated in the Source Code

In lingbot_map/models/gct_stream_window_v2.py, the inference_windowed method computes effective overlap between lines 1093–1110. The logic ensures the overlap region always contains at least one paired keyframe required by _pairwise_alignment:


# Resolve overlap in *actual frames*

if overlap_keyframes is not None:
    kf = max(keyframe_interval, 1)                     # guard against 0

    eff_overlap = max(ws, overlap_keyframes * kf)      # ws = num_scale_frames

elif overlap_size is not None:
    eff_overlap = overlap_size
else:
    eff_overlap = ws                                   # default = scale frames

eff_overlap = min(eff_overlap, S - 1) if S > 1 else 0

Critical insight: When keyframe_interval > 1, specifying --overlap_size in raw frames might create an overlap with no keyframes, causing alignment to fall back to a less accurate first-frame heuristic. Always use --overlap_keyframes instead; the code multiplies this by keyframe_interval and takes the maximum with num_scale_frames to guarantee paired keyframes exist in the overlap zone.

Configuring Window Strategy for 3000+ Frames

Follow this workflow to process sequences longer than 3000 frames without memory exhaustion or alignment drift:

  1. Select keyframe_interval based on GPU memory. A value of 2 halves the KV-cache size compared to storing every frame.

  2. Choose window_size (64–128 is typical). With keyframe_interval=2, window_size=128 yields 248 actual frames per window.

  3. Set overlap_keyframes to ensure sufficient alignment data. The effective overlap in actual frames becomes:

    eff_overlap = max(8, overlap_keyframes * keyframe_interval)

    Setting overlap_keyframes=16 with keyframe_interval=2 produces a 32-frame overlap, providing 16 keyframes for the alignment routine to estimate similarity transforms.

  4. Calculate window count:

    window_stride = actual_frames_per_window - eff_overlap
    windows_needed = ceil(total_frames / window_stride)

    For 3000 frames with 248-frame windows and 32-frame overlap: ceil(3000 / 216) = 14 windows.

CLI and Python API Examples

Command-Line Usage

Run demo.py with the windowed inference mode for a 3000-frame video:

python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --video_path my_long_video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 16 \
    --keyframe_interval 2 \
    --output_device cpu
  • --mode windowed triggers inference_windowed.
  • --output_device cpu offloads per-frame predictions to system RAM, preventing GPU OOM during long sequences.

Python API Implementation

For programmatic control, instantiate GCTStream and call inference_windowed:

import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

model = GCTStream(
    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
    enable_3d_rope=True,
).eval().to('cuda')

# imgs shape: [B, S, 3, H, W] where S=3000

imgs = torch.load('my_frames.pt')  

pred = model.inference_windowed(
    images=imgs,
    window_size=128,
    overlap_keyframes=16,
    keyframe_interval=2,
    output_device=torch.device('cpu')
)

poses = pred['pose_enc']  # shape [1, 3000, 9]

Troubleshooting Common Pitfalls

Symptom Root Cause Solution
Mis-aligned trajectories at window boundaries eff_overlap smaller than num_scale_frames or no paired keyframes in overlap region Increase --overlap_keyframes (e.g., 8→16) or decrease --keyframe_interval
GPU OOM despite windowed mode window_size too large or keyframe_interval too small Reduce window_size to 64 or increase keyframe_interval to 4
Severe performance degradation Per-frame tensors accumulating on GPU Specify --output_device cpu to move results off-GPU immediately
Alignment fallback to first frame overlap_keyframes=0 while keyframe_interval>1 Set --overlap_keyframes ≥ num_scale_frames / keyframe_interval

Summary

  • LingBot-Map handles sequences over 3000 frames via inference_windowed in gct_stream_window_v2.py, which splits video into overlapping chunks with independent KV-caches.
  • Always specify overlap in keyframes (--overlap_keyframes) rather than raw pixels or frames when keyframe_interval > 1 to ensure the alignment routine finds paired keyframes.
  • Effective overlap is calculated as max(num_scale_frames, overlap_keyframes * keyframe_interval), guaranteeing at least one shared keyframe between windows.
  • Memory management relies on the sparsity of the KV-cache; increasing keyframe_interval reduces memory linearly but requires proportional increases in overlap keyframes to maintain alignment accuracy.
  • Offload outputs to CPU via --output_device cpu to prevent GPU memory accumulation during inference of very long sequences.

Frequently Asked Questions

What is the difference between overlap_size and overlap_keyframes in LingBot-Map?

overlap_size specifies the overlap in raw frame count, while overlap_keyframes specifies it in terms of keyframe indices. When keyframe_interval > 1, using overlap_size risks creating an overlap region containing no actual keyframes, forcing _pairwise_alignment to use inaccurate first-frame matching. The --overlap_keyframes flag ensures the code multiplies by keyframe_interval and enforces a minimum overlap of num_scale_frames, guaranteeing valid alignment anchors.

Why does my inference fail with mis-aligned trajectories at window boundaries?

Mis-alignment occurs when the effective overlap contains fewer frames than num_scale_frames (default 8) or lacks keyframes present in both windows. According to the source code in gct_stream_window_v2.py, the overlap must be at least max(num_scale_frames, overlap_keyframes * keyframe_interval) actual frames. Increase --overlap_keyframes until the overlap region contains sufficient keyframes for the similarity transform estimation.

How do I prevent GPU Out-Of-Memory errors when processing 3000+ frames?

Reduce the KV-cache footprint by increasing --keyframe_interval (e.g., to 2 or 4) to store fewer frames per window, or decrease --window_size to process fewer keyframes simultaneously. Additionally, set --output_device cpu to move prediction tensors off the GPU immediately after each window processes, preventing accumulation of per-frame pose encodings in video memory.

Can I use windowed inference with the Python API instead of the CLI?

Yes. Import GCTStream from lingbot_map.models.gct_stream_window_v2, initialize the model with desired cache settings, and call the inference_windowed method with explicit window_size, overlap_keyframes, and keyframe_interval arguments. Pass output_device=torch.device('cpu') to manage memory for long sequences programmatically without modifying global device settings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →