# What Is the Pose-Reference Window in LingBot-Map? Purpose and Implementation

> Explore the purpose of the pose-reference window in LingBot-Map for computing similarity transforms and maintaining 3D trajectory consistency during streaming inference.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-26

---

**The pose-reference window in LingBot-Map is the overlapping set of keyframes shared between consecutive temporal windows that enables the model to compute similarity transforms and maintain globally consistent 3D trajectories during streaming inference.**

In the Robbyant/lingbot-map repository, the pose-reference window solves a critical alignment problem when processing long video sequences in windowed streaming mode. As the `GCTStream` class processes video chunks with limited KV-cache memory, it must stitch together predictions from separate windows into a single coherent coordinate frame without accumulating drift.

## Why LingBot-Map Requires a Pose-Reference Window

Streaming inference splits arbitrarily long videos into fixed-size temporal windows to bound memory usage. Each window processes `window_size` keyframes, but without alignment, the 3D coordinate systems of adjacent windows would drift relative to each other.

The pose-reference window provides the necessary overlap—typically the last few keyframes of one window and the first few of the next—whose pose encodings appear in **both** adjacent windows. By comparing these shared frames, the system computes a similarity transform (scale, rotation, and translation) that warps the later window’s predictions into the coordinate system of the previous window.

This mechanism is implemented in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), where the high-level docstring of `inference_windowed` (lines 758–770) explains that the overlap "guarantees the highest quality predictions at alignment boundaries."

## How the Pose-Reference Window Works

The alignment pipeline involves three coordinated stages that resolve the relative pose between windows using the shared reference frames.

### Window Configuration and Overlap Resolution

The streaming parameters define the size and composition of the pose-reference window:

- **`window_size`**: The number of keyframes retained in the KV-cache for each processing window.
- **`overlap_keyframes`** or **`overlap_size`**: Defines how many frames are shared between successive windows.
- **`num_scale_frames`**: The minimum number of frames used as reference inside each window.

The system automatically converts overlap specifications into an *effective overlap* (`eff_overlap`). According to the overlap resolution logic (lines 31–44), the effective overlap is calculated as:

```python
eff_overlap = max(num_scale_frames, overlap_keyframes * keyframe_interval)

```

This ensures the pose-reference window always contains at least the number of scale frames required for stable alignment, even when keyframe intervals are sparse.

### Computing Alignment with `_pairwise_alignment`

When two windows overlap, the method `_pairwise_alignment` (lines 78–105) extracts the **pose encodings** (`pose_enc`) from the overlapping keyframes (where `is_keyframe` is true). It then computes the relative transformation between windows:

- **Scale estimation** (`_depth_ratio_scale`): Uses depth ratios from the overlapping poses to resolve absolute scale.
- **Rotation matrix** (`R`): Derived from the relative orientation of camera extrinsics.
- **Translation vector** (`t`): Computed from the relative camera positions.

The function utilizes utilities from [`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py), such as `closed_form_inverse_se3`, to handle SE(3) transformations during re-projection.

### Warping Predictions with `_warp_predictions`

After computing the similarity transform (scale `s`, rotation `R`, translation `t`), the system calls `_warp_predictions` to apply this transform to all predictions in the current window. This warps the local coordinate frame of the new window into the global coordinate system established by previous windows, ensuring temporal consistency across the entire sequence.

## Implementing Pose-Reference Windows in Practice

The following example demonstrates how to invoke windowed inference with explicit pose-reference parameters in LingBot-Map:

```python
from lingbot_map.models.gct_stream import GCTStream

# Initialize model

model = GCTStream(...)

# Run windowed inference with pose-reference overlap

preds = model.inference_windowed(
    images,                     # B×S×3×H×W tensor

    window_size=16,             # 16 keyframes per window

    overlap_keyframes=4,        # 4 keyframes shared between windows (pose-reference)

    num_scale_frames=2,         # Minimum 2 scale frames for reference

    flow_threshold=0.0,        # Use fixed-interval keyframe mode

)

# Internal alignment process:

# 1. Compute effective overlap: max(2, 4 * keyframe_interval)

# 2. Extract pose_enc from overlapping keyframes

# 3. s, R, t = self._pairwise_alignment(prev_pred, cur_pred, overlap, ...)

# 4. cur_pred = self._warp_predictions(cur_pred, R, t, s, batch_size)

```

The overlapping pose frames serve as anchors that prevent trajectory drift while allowing the model to process videos of arbitrary length with fixed memory constraints.

## Key Components and File Locations

Understanding the pose-reference window requires familiarity with these specific modules in the LingBot-Map codebase:

| File | Purpose |
|------|---------|
| [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) | Contains `GCTStream.inference_windowed` and the pose-reference window alignment logic. |
| [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) | Defines the streaming inference pipeline and KV-cache management for overlapping windows. |
| [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py) | Provides conversion utilities between camera extrinsics/intrinsics and compact pose encodings (`pose_enc`). |
| [`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py) | Implements geometric helpers like `closed_form_inverse_se3` used for pose alignment transformations. |

## Summary

- The **pose-reference window** consists of overlapping keyframes shared between consecutive processing windows in `GCTStream.inference_windowed`.
- It enables **global trajectory consistency** by providing common anchor frames used to compute similarity transforms between local window coordinate systems.
- The `_pairwise_alignment` method extracts pose encodings from these overlapping frames to estimate relative scale, rotation, and translation.
- The `_warp_predictions` method applies these transforms to align new predictions with the established global map.
- Configuration parameters `overlap_keyframes`, `num_scale_frames`, and `window_size` control the size and quality of the pose-reference window to balance memory efficiency against alignment accuracy.

## Frequently Asked Questions

### What is the difference between `window_size` and `overlap_keyframes`?

**`window_size`** defines the total number of keyframes processed in each temporal window (affecting KV-cache memory), while **`overlap_keyframes`** specifies how many of those frames are shared with the adjacent window to form the pose-reference set. The overlap ensures there are sufficient common poses to compute alignment transforms, whereas the window size determines the computational granularity.

### How does the pose-reference window maintain scale consistency across long videos?

The overlapping frames in the pose-reference window provide depth-ratio information through `_depth_ratio_scale` calculations in `_pairwise_alignment`. By comparing the depth estimates of the same physical points visible in both windows, the system resolves the relative scale factor `s` and applies it via `_warp_predictions`, preventing the scale drift that typically accumulates in monocular SLAM systems.

### Why are `num_scale_frames` important for the pose-reference window?

The `num_scale_frames` parameter establishes a minimum floor for the effective overlap size. Even if `overlap_keyframes` is set to a small value or the keyframe interval is large, requiring at least `num_scale_frames` ensures the pose-reference window always contains enough frames with valid pose predictions to compute a stable similarity transform. This guarantees robust alignment at window boundaries regardless of motion dynamics.

### Where is the pose encoding utility code located in the repository?

The conversion between camera parameters and the compact pose encoding (`pose_enc`) used during alignment is implemented in [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py). This module handles the encoding of camera extrinsics and intrinsics into the format consumed by `_pairwise_alignment` during the window stitching process.