# How Windowed Inference Handles Overlapping Windows and Cross-Window Pose Alignment in LingBot-Map

> Discover how LingBot-Map achieves seamless 3D reconstruction using windowed inference with overlapping windows and cross-window pose alignment for long video streams.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-29

---

**LingBot-Map processes long video streams by dividing them into overlapping temporal windows and applying similarity transforms to align poses across window boundaries, ensuring seamless 3D reconstruction.**

The `GCTStream` architecture in the [LingBot-Map](https://github.com/Robbyant/lingbot-map) repository employs **windowed inference** to process arbitrarily long video sequences without exceeding GPU memory constraints. By breaking streams into manageable chunks with strategic overlap and geometric alignment, the system maintains spatial consistency across window transitions. This approach is implemented primarily in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), where overlapping frames serve as geometric anchors for cross-window pose alignment.

## Overlapping Window Strategy

LingBot-Map guarantees temporal continuity by sharing **overlap frames** between consecutive windows. These shared frames provide bidirectional context that smooths transitions and enable geometric alignment.

### Configurable Overlap Calculations

The effective overlap (`eff_overlap`) is resolved dynamically in `inference_windowed` (lines [1093‑1110](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L1093-L1110)) with the following priority:

```python

# Resolve overlap in *actual frames* (priority: overlap_keyframes → overlap_size → default)

if overlap_keyframes is not None:
    kf = max(keyframe_interval, 1)
    eff_overlap = max(ws, overlap_keyframes * kf)
elif overlap_size is not None:
    eff_overlap = overlap_size
else:
    eff_overlap = ws          # default: number of scale frames

eff_overlap = min(eff_overlap, S - 1) if S > 1 else 0

```

Here, `ws` represents the number of scale frames per window. The default behavior ensures the overlap always contains at least the scale frames, guaranteeing that the next window can reuse the same bidirectional context.

### Scale Frames as Bidirectional Anchors

Each window is processed with a **fresh KV cache** (`self.clean_kv_cache()`) to maintain constant memory usage. The first `num_scale_frames` frames in every window receive bidirectional attention, while subsequent frames use causal KV-cache mechanics. When windows overlap, the shared frames are processed as scale frames in the subsequent window, providing the model with geometric memory of the previous context.

## Cross-Window Pose Alignment Pipeline

After independent window inference, LingBot-Map warps predictions into a unified coordinate frame using **similarity transforms** (scale *s*, rotation *R*, translation *t*) estimated from paired keyframes within overlap regions.

### Pairwise Similarity Transform Estimation

The `_pairwise_alignment` method (lines [842‑886](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L842-L886)) computes the relative transformation between consecutive windows:

1. **Keyframe Pairing**: Identify paired keyframes inside the overlap region where both windows label the frame as a keyframe. If no paired keyframes exist, the first overlap frame serves as fallback.
2. **Pose Extraction**: Convert quaternion encodings to rotation matrices using `quat_to_mat` from [`lingbot_map/utils/rotation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/rotation.py).
3. **Relative Rotation**: Compute `R_ab = Ra · Rbᵀ` between window *a* and window *b*.
4. **Scale Estimation**: Calculate the median depth ratio across all paired keyframes via `_depth_ratio_scale`.
5. **Translation Computation**: Derive `t_ab = ca – s·R_ab·cb`, where *ca* and *cb* are camera centers.

The function returns the tuple `(s_ab, R_ab, t_ab)` that aligns the current window to the previous window's reference frame.

### Predictive Warping

The `_warp_predictions` function (lines [909‑945]((https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L909-L945))) applies the similarity transform to every prediction field:

- **Pose encodings**: Centers are scaled and rotated; quaternions are rotated via `R_ab`.
- **Depth maps**: Multiplied by the estimated scale factor `s`.
- **World points**: Transformed as `s·R·p + t`.
- **Intrinsics**: Remain unchanged as they are camera-intrinsic properties.

All other prediction keys pass through unchanged, preserving semantic consistency while adjusting geometric coordinates.

### Temporal Stitching with Deduplication

The `_align_and_stitch_windows` orchestrator (lines [957‑979](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L957-L979)) chains these operations, while `_stitch_windows` (lines [743‑791](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L743-L791)) constructs a **slice table** for each temporal key (pose, depth, world points). For non-final windows, the slice ends `overlap` frames before the window's terminus, discarding duplicated overlap regions. The remaining slices are concatenated via `torch.cat` along the temporal axis, yielding a seamless prediction tensor.

## Practical Implementation Examples

### Basic Windowed Inference with Default Overlap

The default overlap automatically equals the number of scale frames (`num_scale_frames`):

```python
import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

model = GCTStream(...)
images = torch.randn(1, 200, 3, 518, 518)          # B×S×C×H×W

# Run windowed inference with 16 keyframes per window, default overlap

preds = model.inference_windowed(
    images,
    window_size=16,          # 16 keyframes per window (includes scale frames)

    num_scale_frames=4,      # 4 scale frames per window

    keyframe_interval=1,    # every frame is a keyframe

    flow_threshold=0.0,     # disable flow-based keyframe detection

)
print(preds["pose_enc"].shape)   # → torch.Size([1, 200, 9])

```

### Custom Overlap Expressed in Keyframes

Specify overlap duration using `overlap_keyframes` for precise geometric alignment:

```python
preds = model.inference_windowed(
    images,
    window_size=20,
    overlap_keyframes=2,      # 2 keyframes of overlap

    keyframe_interval=4,     # every 4th frame is a keyframe

)

```

This configuration yields `eff_overlap = max(4, 8) = 8` actual frames, ensuring sufficient keyframes exist in the overlap for robust similarity transform estimation.

### Flow-Based Dynamic Windows

Enable optical flow keyframe detection for adaptive windowing:

```python
preds = model.inference_windowed(
    images,
    flow_threshold=1.5,      # trigger new keyframe when mean optical flow > 1.5 px

    max_non_keyframe_gap=30,
)

```

When `flow_threshold > 0`, the method enters the flow-based branch (lines [1069‑1082](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L1069-L1082)), dynamically creating windows while maintaining the same overlap and alignment pipeline.

### Inspecting Alignment Metadata

Access per-window transformation parameters for debugging or downstream processing:

```python
scales = preds["chunk_scales"]          # (B, N_windows, 1) – per-window scale factors

transforms = preds["chunk_transforms"]  # (B, N_windows, 4, 4) – similarity matrices

print(scales.shape, transforms.shape)

```

These tensors originate from `_align_and_stitch_windows` (lines [1014‑1018](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py#L1014-L1018)) and record the alignment mode as `"scaled"`.

## Summary

- **Windowed inference** in LingBot-Map processes long videos by chunking them into temporal windows with configurable overlap (defaulting to the scale frame count).
- **Overlap handling** ensures geometric continuity by processing shared frames as scale frames in subsequent windows, providing bidirectional context.
- **Cross-window pose alignment** estimates similarity transforms (scale, rotation, translation) from paired keyframes in overlap regions via `_pairwise_alignment`.
- **Predictive warping** applies these transforms to poses, depth, and world points using `_warp_predictions` before stitching.
- **Memory efficiency** is maintained through fresh KV caches per window and deduplication during the `_stitch_windows` concatenation phase.

## Frequently Asked Questions

### What is the default overlap size in windowed inference?

The default overlap equals the number of **scale frames** (`num_scale_frames` or `ws`). This guarantees that the subsequent window receives the same bidirectional context as the previous window's trailing frames, ensuring smooth geometric continuity without requiring manual overlap configuration.

### How does the system handle alignment when no keyframes exist in the overlap region?

If `_pairwise_alignment` cannot identify paired keyframes within the overlap, it falls back to using the **first overlap frame** as the alignment anchor. While less robust than multi-keyframe estimation, this ensures the similarity transform computation never fails, maintaining pipeline continuity even in low-texture or static video segments.

### Why are scale frames critical for overlapping windows?

Scale frames receive **bidirectional attention** during inference, allowing the model to attend to both past and future context within the window. When used as overlap frames, they provide the subsequent window with geometric memory of the previous window's coordinate system, enabling the alignment algorithms to estimate accurate relative poses and scales.

### Can windowed inference process videos longer than GPU memory allows?

Yes. By processing video streams in independent windows with **fresh KV caches** (`self.clean_kv_cache()`), LingBot-Map maintains constant GPU memory usage regardless of video length. The stitching mechanism concatenates results after CPU-side alignment, enabling theoretically unlimited sequence lengths limited only by storage rather than VRAM.