How to Manage Overlapping Windows for Sequences Over 3000 Frames with Windowed Inference
LingBot-Map processes arbitrarily long video streams by splitting them into overlapping windows with fresh KV-caches per window, where the effective overlap must be calculated in keyframes rather than raw frames to ensure accurate alignment across window boundaries.
LingBot-Map reconstructs long video sequences through windowed inference, a technique that divides inputs into manageable chunks to bound GPU memory usage. For sequences exceeding 3000 frames, configuring the overlap between windows correctly is critical to prevent trajectory misalignment and Out-Of-Memory (OOM) errors. The implementation resides in lingbot_map/models/gct_stream_window_v2.py, which handles KV-cache management, keyframe sampling, and pairwise alignment stitching.
Understanding Window Size and Keyframe Interval
LingBot-Map uses a sparse KV-cache that only stores keyframes, not every frame. Three parameters define how many actual frames each window covers:
window_size: The number of keyframe slots allocated in the KV-cache per window (includes scale frames).keyframe_interval: Every N-th frame after the scale phase is retained as a keyframe (N=1 means every frame, N=2 means every second frame).num_scale_frames(default:8): The first M frames in each window receiving bidirectional attention, where no keyframe interval is applied.
The actual frame count per window is calculated as:
actual_frames = num_scale_frames + (window_size - num_scale_frames) * keyframe_interval
For example, with window_size=128, keyframe_interval=2, and the default num_scale_frames=8, each window covers 8 + (120 * 2) = 248 actual frames.
How Overlap is Calculated in the Source Code
In lingbot_map/models/gct_stream_window_v2.py, the inference_windowed method computes effective overlap between lines 1093–1110. The logic ensures the overlap region always contains at least one paired keyframe required by _pairwise_alignment:
# Resolve overlap in *actual frames*
if overlap_keyframes is not None:
kf = max(keyframe_interval, 1) # guard against 0
eff_overlap = max(ws, overlap_keyframes * kf) # ws = num_scale_frames
elif overlap_size is not None:
eff_overlap = overlap_size
else:
eff_overlap = ws # default = scale frames
eff_overlap = min(eff_overlap, S - 1) if S > 1 else 0
Critical insight: When keyframe_interval > 1, specifying --overlap_size in raw frames might create an overlap with no keyframes, causing alignment to fall back to a less accurate first-frame heuristic. Always use --overlap_keyframes instead; the code multiplies this by keyframe_interval and takes the maximum with num_scale_frames to guarantee paired keyframes exist in the overlap zone.
Configuring Window Strategy for 3000+ Frames
Follow this workflow to process sequences longer than 3000 frames without memory exhaustion or alignment drift:
-
Select
keyframe_intervalbased on GPU memory. A value of2halves the KV-cache size compared to storing every frame. -
Choose
window_size(64–128 is typical). Withkeyframe_interval=2,window_size=128yields 248 actual frames per window. -
Set
overlap_keyframesto ensure sufficient alignment data. The effective overlap in actual frames becomes:eff_overlap = max(8, overlap_keyframes * keyframe_interval)Setting
overlap_keyframes=16withkeyframe_interval=2produces a 32-frame overlap, providing 16 keyframes for the alignment routine to estimate similarity transforms. -
Calculate window count:
window_stride = actual_frames_per_window - eff_overlap windows_needed = ceil(total_frames / window_stride)For 3000 frames with 248-frame windows and 32-frame overlap:
ceil(3000 / 216) = 14windows.
CLI and Python API Examples
Command-Line Usage
Run demo.py with the windowed inference mode for a 3000-frame video:
python demo.py \
--model_path /path/to/lingbot-map.pt \
--video_path my_long_video.mp4 \
--fps 10 \
--mode windowed \
--window_size 128 \
--overlap_keyframes 16 \
--keyframe_interval 2 \
--output_device cpu
--mode windowedtriggersinference_windowed.--output_device cpuoffloads per-frame predictions to system RAM, preventing GPU OOM during long sequences.
Python API Implementation
For programmatic control, instantiate GCTStream and call inference_windowed:
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream
model = GCTStream(
kv_cache_sliding_window=64,
kv_cache_scale_frames=8,
enable_3d_rope=True,
).eval().to('cuda')
# imgs shape: [B, S, 3, H, W] where S=3000
imgs = torch.load('my_frames.pt')
pred = model.inference_windowed(
images=imgs,
window_size=128,
overlap_keyframes=16,
keyframe_interval=2,
output_device=torch.device('cpu')
)
poses = pred['pose_enc'] # shape [1, 3000, 9]
Troubleshooting Common Pitfalls
| Symptom | Root Cause | Solution |
|---|---|---|
| Mis-aligned trajectories at window boundaries | eff_overlap smaller than num_scale_frames or no paired keyframes in overlap region |
Increase --overlap_keyframes (e.g., 8→16) or decrease --keyframe_interval |
| GPU OOM despite windowed mode | window_size too large or keyframe_interval too small |
Reduce window_size to 64 or increase keyframe_interval to 4 |
| Severe performance degradation | Per-frame tensors accumulating on GPU | Specify --output_device cpu to move results off-GPU immediately |
| Alignment fallback to first frame | overlap_keyframes=0 while keyframe_interval>1 |
Set --overlap_keyframes ≥ num_scale_frames / keyframe_interval |
Summary
- LingBot-Map handles sequences over 3000 frames via
inference_windowedingct_stream_window_v2.py, which splits video into overlapping chunks with independent KV-caches. - Always specify overlap in keyframes (
--overlap_keyframes) rather than raw pixels or frames whenkeyframe_interval > 1to ensure the alignment routine finds paired keyframes. - Effective overlap is calculated as
max(num_scale_frames, overlap_keyframes * keyframe_interval), guaranteeing at least one shared keyframe between windows. - Memory management relies on the sparsity of the KV-cache; increasing
keyframe_intervalreduces memory linearly but requires proportional increases in overlap keyframes to maintain alignment accuracy. - Offload outputs to CPU via
--output_device cputo prevent GPU memory accumulation during inference of very long sequences.
Frequently Asked Questions
What is the difference between overlap_size and overlap_keyframes in LingBot-Map?
overlap_size specifies the overlap in raw frame count, while overlap_keyframes specifies it in terms of keyframe indices. When keyframe_interval > 1, using overlap_size risks creating an overlap region containing no actual keyframes, forcing _pairwise_alignment to use inaccurate first-frame matching. The --overlap_keyframes flag ensures the code multiplies by keyframe_interval and enforces a minimum overlap of num_scale_frames, guaranteeing valid alignment anchors.
Why does my inference fail with mis-aligned trajectories at window boundaries?
Mis-alignment occurs when the effective overlap contains fewer frames than num_scale_frames (default 8) or lacks keyframes present in both windows. According to the source code in gct_stream_window_v2.py, the overlap must be at least max(num_scale_frames, overlap_keyframes * keyframe_interval) actual frames. Increase --overlap_keyframes until the overlap region contains sufficient keyframes for the similarity transform estimation.
How do I prevent GPU Out-Of-Memory errors when processing 3000+ frames?
Reduce the KV-cache footprint by increasing --keyframe_interval (e.g., to 2 or 4) to store fewer frames per window, or decrease --window_size to process fewer keyframes simultaneously. Additionally, set --output_device cpu to move prediction tensors off the GPU immediately after each window processes, preventing accumulation of per-frame pose encodings in video memory.
Can I use windowed inference with the Python API instead of the CLI?
Yes. Import GCTStream from lingbot_map.models.gct_stream_window_v2, initialize the model with desired cache settings, and call the inference_windowed method with explicit window_size, overlap_keyframes, and keyframe_interval arguments. Pass output_device=torch.device('cpu') to manage memory for long sequences programmatically without modifying global device settings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →