How Windowed Inference Handles Overlapping Windows and Cross-Window Pose Alignment in LingBot-Map
LingBot-Map processes long video streams by dividing them into overlapping temporal windows and applying similarity transforms to align poses across window boundaries, ensuring seamless 3D reconstruction.
The GCTStream architecture in the LingBot-Map repository employs windowed inference to process arbitrarily long video sequences without exceeding GPU memory constraints. By breaking streams into manageable chunks with strategic overlap and geometric alignment, the system maintains spatial consistency across window transitions. This approach is implemented primarily in lingbot_map/models/gct_stream_window_v2.py, where overlapping frames serve as geometric anchors for cross-window pose alignment.
Overlapping Window Strategy
LingBot-Map guarantees temporal continuity by sharing overlap frames between consecutive windows. These shared frames provide bidirectional context that smooths transitions and enable geometric alignment.
Configurable Overlap Calculations
The effective overlap (eff_overlap) is resolved dynamically in inference_windowed (lines 1093‑1110) with the following priority:
# Resolve overlap in *actual frames* (priority: overlap_keyframes → overlap_size → default)
if overlap_keyframes is not None:
kf = max(keyframe_interval, 1)
eff_overlap = max(ws, overlap_keyframes * kf)
elif overlap_size is not None:
eff_overlap = overlap_size
else:
eff_overlap = ws # default: number of scale frames
eff_overlap = min(eff_overlap, S - 1) if S > 1 else 0
Here, ws represents the number of scale frames per window. The default behavior ensures the overlap always contains at least the scale frames, guaranteeing that the next window can reuse the same bidirectional context.
Scale Frames as Bidirectional Anchors
Each window is processed with a fresh KV cache (self.clean_kv_cache()) to maintain constant memory usage. The first num_scale_frames frames in every window receive bidirectional attention, while subsequent frames use causal KV-cache mechanics. When windows overlap, the shared frames are processed as scale frames in the subsequent window, providing the model with geometric memory of the previous context.
Cross-Window Pose Alignment Pipeline
After independent window inference, LingBot-Map warps predictions into a unified coordinate frame using similarity transforms (scale s, rotation R, translation t) estimated from paired keyframes within overlap regions.
Pairwise Similarity Transform Estimation
The _pairwise_alignment method (lines 842‑886) computes the relative transformation between consecutive windows:
- Keyframe Pairing: Identify paired keyframes inside the overlap region where both windows label the frame as a keyframe. If no paired keyframes exist, the first overlap frame serves as fallback.
- Pose Extraction: Convert quaternion encodings to rotation matrices using
quat_to_matfromlingbot_map/utils/rotation.py. - Relative Rotation: Compute
R_ab = Ra · Rbᵀbetween window a and window b. - Scale Estimation: Calculate the median depth ratio across all paired keyframes via
_depth_ratio_scale. - Translation Computation: Derive
t_ab = ca – s·R_ab·cb, where ca and cb are camera centers.
The function returns the tuple (s_ab, R_ab, t_ab) that aligns the current window to the previous window's reference frame.
Predictive Warping
The _warp_predictions function (lines 909‑945) applies the similarity transform to every prediction field:
- Pose encodings: Centers are scaled and rotated; quaternions are rotated via
R_ab. - Depth maps: Multiplied by the estimated scale factor
s. - World points: Transformed as
s·R·p + t. - Intrinsics: Remain unchanged as they are camera-intrinsic properties.
All other prediction keys pass through unchanged, preserving semantic consistency while adjusting geometric coordinates.
Temporal Stitching with Deduplication
The _align_and_stitch_windows orchestrator (lines 957‑979) chains these operations, while _stitch_windows (lines 743‑791) constructs a slice table for each temporal key (pose, depth, world points). For non-final windows, the slice ends overlap frames before the window's terminus, discarding duplicated overlap regions. The remaining slices are concatenated via torch.cat along the temporal axis, yielding a seamless prediction tensor.
Practical Implementation Examples
Basic Windowed Inference with Default Overlap
The default overlap automatically equals the number of scale frames (num_scale_frames):
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream
model = GCTStream(...)
images = torch.randn(1, 200, 3, 518, 518) # B×S×C×H×W
# Run windowed inference with 16 keyframes per window, default overlap
preds = model.inference_windowed(
images,
window_size=16, # 16 keyframes per window (includes scale frames)
num_scale_frames=4, # 4 scale frames per window
keyframe_interval=1, # every frame is a keyframe
flow_threshold=0.0, # disable flow-based keyframe detection
)
print(preds["pose_enc"].shape) # → torch.Size([1, 200, 9])
Custom Overlap Expressed in Keyframes
Specify overlap duration using overlap_keyframes for precise geometric alignment:
preds = model.inference_windowed(
images,
window_size=20,
overlap_keyframes=2, # 2 keyframes of overlap
keyframe_interval=4, # every 4th frame is a keyframe
)
This configuration yields eff_overlap = max(4, 8) = 8 actual frames, ensuring sufficient keyframes exist in the overlap for robust similarity transform estimation.
Flow-Based Dynamic Windows
Enable optical flow keyframe detection for adaptive windowing:
preds = model.inference_windowed(
images,
flow_threshold=1.5, # trigger new keyframe when mean optical flow > 1.5 px
max_non_keyframe_gap=30,
)
When flow_threshold > 0, the method enters the flow-based branch (lines 1069‑1082), dynamically creating windows while maintaining the same overlap and alignment pipeline.
Inspecting Alignment Metadata
Access per-window transformation parameters for debugging or downstream processing:
scales = preds["chunk_scales"] # (B, N_windows, 1) – per-window scale factors
transforms = preds["chunk_transforms"] # (B, N_windows, 4, 4) – similarity matrices
print(scales.shape, transforms.shape)
These tensors originate from _align_and_stitch_windows (lines 1014‑1018) and record the alignment mode as "scaled".
Summary
- Windowed inference in LingBot-Map processes long videos by chunking them into temporal windows with configurable overlap (defaulting to the scale frame count).
- Overlap handling ensures geometric continuity by processing shared frames as scale frames in subsequent windows, providing bidirectional context.
- Cross-window pose alignment estimates similarity transforms (scale, rotation, translation) from paired keyframes in overlap regions via
_pairwise_alignment. - Predictive warping applies these transforms to poses, depth, and world points using
_warp_predictionsbefore stitching. - Memory efficiency is maintained through fresh KV caches per window and deduplication during the
_stitch_windowsconcatenation phase.
Frequently Asked Questions
What is the default overlap size in windowed inference?
The default overlap equals the number of scale frames (num_scale_frames or ws). This guarantees that the subsequent window receives the same bidirectional context as the previous window's trailing frames, ensuring smooth geometric continuity without requiring manual overlap configuration.
How does the system handle alignment when no keyframes exist in the overlap region?
If _pairwise_alignment cannot identify paired keyframes within the overlap, it falls back to using the first overlap frame as the alignment anchor. While less robust than multi-keyframe estimation, this ensures the similarity transform computation never fails, maintaining pipeline continuity even in low-texture or static video segments.
Why are scale frames critical for overlapping windows?
Scale frames receive bidirectional attention during inference, allowing the model to attend to both past and future context within the window. When used as overlap frames, they provide the subsequent window with geometric memory of the previous window's coordinate system, enabling the alignment algorithms to estimate accurate relative poses and scales.
Can windowed inference process videos longer than GPU memory allows?
Yes. By processing video streams in independent windows with fresh KV caches (self.clean_kv_cache()), LingBot-Map maintains constant GPU memory usage regardless of video length. The stitching mechanism concatenates results after CPU-side alignment, enabling theoretically unlimited sequence lengths limited only by storage rather than VRAM.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →