How the `overlap_keyframes` Parameter Ensures Stable Pose Alignment Across Sliding Windows
The overlap_keyframes parameter guarantees stable pose alignment by converting keyframe units into actual frame counts, ensuring every overlap region contains at least one mutually-present keyframe that enables computation of a robust similarity transform between windows.
In the lingbot-map repository, windowed inference splits long video sequences into manageable chunks to process them independently. The overlap_keyframes parameter in lingbot_map/models/gct_stream_window.py provides a high-level mechanism to specify overlap between consecutive sliding windows using keyframe units rather than raw frame counts, directly preventing pose drift when stitching predictions back together.
The Alignment Challenge in Windowed Inference
When processing video sequences longer than the model's maximum window size, the system divides frames into overlapping segments. Each window generates independent pose predictions, and these must be aligned into a single global coordinate system. Stable alignment requires sufficient overlap between windows and keyframe pairing—specifically, the overlap region must contain at least one cached keyframe present in both the previous and current windows. Without this shared reference point, the similarity transform (scale and translation) computed between windows becomes unreliable, causing trajectory drift.
How overlap_keyframes Converts to Effective Frame Overlap
The implementation in lingbot_map/models/gct_stream_window.py (lines 31-40) handles the conversion from keyframe units to actual frames through a priority-based calculation that ensures robust alignment conditions.
Parameter Priority and Conversion Logic
The system checks overlap_keyframes first, taking precedence over the legacy overlap_size parameter:
if overlap_keyframes is not None:
# Conversion logic executes here
elif overlap_size is not None:
# Legacy frame-based overlap
else:
# Default handling
The conversion multiplies the requested keyframe overlap by the keyframe_interval (the spacing between cached keyframes) to yield an effective overlap in actual frames:
kf = max(keyframe_interval, 1)
eff_overlap = max(ws, overlap_keyframes * kf)
This formula ensures that requesting 4 keyframes with a keyframe_interval of 2 produces 8 actual frames of overlap, guaranteeing that the region spans enough raw frames to include the cached keyframes from both windows.
Minimum Overlap Enforcement
The calculation enforces a floor using max(ws, ...) where ws represents the number of scale frames—frames used for depth-ratio scaling that are always present at the start of every window. This guarantees the overlap region always contains these critical reference frames regardless of the keyframe calculation.
Sequence Boundary Clamping
The effective overlap is clamped to the sequence length to prevent invalid windows:
eff_overlap = min(eff_overlap, S-1) if S > 1 else 0
Where S represents the total number of frames in the sequence.
The Stitching Process and Keyframe Pairing
Once calculated, eff_overlap passes to _stitch_windows(warped_windows, win_sz, overlap), which performs the actual alignment. The method locates paired keyframes inside the overlap region—keyframes present in both the tail of the previous window and the head of the current window. The keyframe-pair logic, found in the if is_kf_prev is not None ... block within _stitch_windows, computes a robust similarity transform by aggregating depth-scaled transformations across these matched pairs.
Because the overlap region always contains at least one mutually-present keyframe, the pose encoding from the previous window transforms reliably into the next window's coordinate system. The merged output includes "alignment_mode": "scaled" and retains per-window transforms, producing a smooth global trajectory without drift.
Practical Usage Examples
Configure windowed inference with automatic overlap calculation:
pred = model.inference_windowed(
images, # [S,3,H,W] tensor
window_size=16, # 16 keyframes per window
overlap_keyframes=4, # 4 keyframes of overlap (converted to actual frames)
keyframe_interval=2, # cache every 2nd frame
)
Explicitly control the conversion while disabling legacy parameters:
pred = model.inference_windowed(
images,
window_size=16,
overlap_keyframes=3, # 3 keyframes → 3 * 2 = 6 actual frames overlap
keyframe_interval=2,
overlap_size=None, # disable legacy frame-based overlap
)
Both calls compute eff_overlap = max(ws, overlap_keyframes * keyframe_interval), ensuring the overlap region always includes at least one paired keyframe for robust alignment according to the implementation in lingbot_map/models/gct_stream_window.py.
Summary
overlap_keyframesspecifies overlap in keyframe units rather than raw frames, taking precedence overoverlap_sizeinlingbot_map/models/gct_stream_window.py.- Conversion formula multiplies keyframes by
keyframe_intervaland enforces a minimum ofwsscale frames:eff_overlap = max(ws, overlap_keyframes * kf). - Keyframe pairing requires the overlap region to contain at least one cached keyframe present in both consecutive windows to compute a reliable similarity transform.
- Drift prevention occurs because the guaranteed keyframe overlap provides stable reference points for
_stitch_windowsto align pose encodings between independent window predictions.
Frequently Asked Questions
What happens if overlap_keyframes is set to None?
When overlap_keyframes is None, the implementation falls back to checking overlap_size according to the parameter priority logic in lingbot_map/models/gct_stream_window.py. If both are None, the system uses default overlap calculations that may not guarantee keyframe presence in the overlap region, potentially reducing alignment stability.
How does overlap_keyframes differ from overlap_size?
overlap_keyframes measures overlap in units of cached keyframes and automatically converts to actual frames using the keyframe_interval, while overlap_size specifies the overlap directly in raw frame counts without considering keyframe spacing. The keyframe-based approach ensures the overlap region always contains reference frames suitable for computing similarity transforms, whereas raw frame counts might miss keyframes entirely depending on the interval setting.
Why must the overlap region contain at least one keyframe from each window?
The stitching algorithm in _stitch_windows computes a similarity transform (scale and translation) to align pose encodings between windows. This calculation requires corresponding reference points that exist in both coordinate systems. Keyframes serve as these anchors because their hidden states are cached and their positions are known in both the previous window's tail and the current window's head. Without a mutually-present keyframe, the system lacks correspondences to solve for the alignment transformation reliably.
What are scale frames (ws) and why enforce minimum overlap with them?
Scale frames are specific frames used for depth-ratio scaling that appear at the start of every window. The implementation enforces eff_overlap = max(ws, ...) because these frames provide critical geometric information for normalizing depth predictions across windows. Ensuring they remain within the overlap region guarantees that the stitching process has access to consistent scale references when merging window predictions, preventing scale drift in the global trajectory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →