# How to Debug Pose Drift Issues in LingBot-MAP When Running on Custom Video with Unknown Camera Intrinsics

> Debug pose drift in LingBot-MAP on custom video with unknown intrinsics. Learn to fix pseudo-intrinsics, undersampling, and frame ordering issues for accurate pose estimation.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-31

---

**Pose drift in LingBot-MAP when processing custom video without calibration files typically stems from incorrect pseudo-intrinsics, temporal undersampling, or frame ordering mismatches that propagate through the pose-encoding pipeline.**

LingBot-MAP builds dense 3D reconstructions from RGB-D streams, but feeding it a custom video with unknown camera intrinsics forces the pipeline to infer parameters that rarely match real sensor characteristics. This guide examines the specific failure points in the `Robbyant/lingbot-map` codebase—from frame extraction in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) to pose encoding in [`utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/utils/pose_enc.py)—to help you diagnose and eliminate trajectory drift.

## How the Pipeline Handles Unknown Intrinsics

When you supply a video path via `--video_path`, the system executes a multi-stage pipeline where each component assumes accurate camera parameters. If intrinsics are missing, the code synthesizes fallback values that often introduce systematic errors.

**Frame Extraction and Dataset Loading**

The entry point `demo.load_images` uses OpenCV (`cv2.VideoCapture`) to extract frames from your video and caches them in a sibling `*_frames` directory. The `benchmark.datasets.GeneralDataset` class then wraps these frames, providing a uniform `__getitem__` interface and handling the optional `video_fps` parameter for temporal sampling.

**Pseudo-Intrinsics Generation**

If no calibration file is found, the viewer falls back to `lingbot_map.vis.point_cloud_viewer.generate_pseudo_intrinsics`. This function assumes a pinhole camera model with focal length set to `max(H, W)`—a heuristic that is almost always incorrect for real cameras and introduces scale errors immediately.

**Pose Encoding and Decoding**

Camera extrinsics and intrinsics are packed into a 9-dimensional pose encoding (`absT_quaR_FoV`) by `extri_intri_to_pose_encoding` in [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py). During inference, `pose_encoding_to_extri_intri` reverses this conversion. If the `image_size_hw` parameter mismatches the actual frame dimensions, or if the intrinsics matrix contains incorrect `fx, fy, cx, cy` values, the decoded poses drift progressively.

**Geometry Processing**

Functions in [`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py)—including `proj`, `iproj`, and `unproject_depth_map_to_point_map`—rely directly on the intrinsics matrix. Errors here manifest as misaligned point clouds and wobbling camera trajectories in the visualization stage.

## Root Causes of Pose Drift with Unknown Intrinsics

When running on custom video, five specific issues typically trigger drift:

- **Incorrect Intrinsics**: The pseudo-intrinsics assume a focal length of `max(H, W)`, creating systematic scale errors in the back-projected 3D points.
- **Temporal Undersampling**: Setting `--fps` too low yields large inter-frame motion, destabilizing the transformer’s relative-pose estimation.
- **Missing Timestamps**: The [`poses_c2w.txt`](https://github.com/Robbyant/lingbot-map/blob/main/poses_c2w.txt) file expects strict chronological order. If frame extraction re-orders frames (e.g., due to variable GOP lengths), the pose encoder receives temporally jittered input.
- **Inconsistent Image Size**: OpenCV may resize frames during decoding while the intrinsics matrix still reflects the original resolution, causing projection mismatches.
- **Pipeline Mis-alignment**: The dataset loader returns tensors of shape `[B, S, C, H, W]` while the model expects `[B, S, H, W, C]`. This subtle shape mismatch perturbs attention masks and corrupts pose predictions.

## Step-by-Step Debugging Workflow

Follow this sequence to isolate the source of drift in your custom video input.

### 1. Verify Frame Extraction

Run the extraction and check the output directory:

```bash
python demo.py --video_path my_clip.mp4 --fps 10 --mode windowed

```

After execution, inspect the generated `my_clip_frames/` directory. The file count must match the `len(paths)` value printed by [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py). Mismatches indicate dropped frames or decoding errors.

### 2. Inspect the Intrinsics Matrix

Check what parameters the system is using:

```python
from lingbot_map.vis.point_cloud_viewer import generate_pseudo_intrinsics
K = generate_pseudo_intrinsics(h=720, w=1280)
print(K)

```

If you know your sensor’s approximate focal length, replace the pseudo-intrinsics with a calibrated matrix. Store this as [`intrinsics.txt`](https://github.com/Robbyant/lingbot-map/blob/main/intrinsics.txt) in the dataset root to override the fallback generation.

### 3. Validate Pose Encoding Round-Trip

Errors in the encoding/decoding functions corrupt reconstruction. Test the conversion:

```python
import torch
from lingbot_map.utils.pose_enc import extri_intri_to_pose_encoding, pose_encoding_to_extri_intri

# Assume extrinsics and intrinsics are tensors from a single frame

encoding = extri_intri_to_pose_encoding(extrinsics, intrinsics, image_size_hw=(720,1280))

# Round-trip check

e_rec, i_rec = pose_encoding_to_extri_intri(encoding, image_size_hw=(720,1280))
assert torch.allclose(extrinsics, e_rec, atol=1e-5)

```

Any deviation indicates a bug in the pose conversion logic or incorrect image dimensions.

### 4. Visualize the Trajectory

Launch the interactive viewer:

```bash
python -m lingbot_map.vis.point_cloud_viewer --data_path my_clip_frames/

```

Enable the **"Show Current Frame"** checkbox to compare the live video frame against the projected point cloud. Large spatial jumps between consecutive frames correlate with faulty intrinsics or timestamp inconsistencies.

### 5. Check Timestamp Gaps

In [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py), the `MAX_TIME_GAP` constant controls the maximum allowable temporal separation between image-pose pairs. If your source video has variable frame-rate, increase this limit or supply a uniform timestamp list to prevent the dataset from filtering valid frames.

### 6. Calibrate Intrinsics (Optional)

If you have access to a calibration target, run OpenCV calibration on extracted frames and write the result to [`intrinsics.txt`](https://github.com/Robbyant/lingbot-map/blob/main/intrinsics.txt):

```python

# Example: Saving a calibrated intrinsics matrix

import numpy as np
np.savetxt('intrinsics.txt', K, fmt='%.6f')

```

Place this file in the same directory as your frames to bypass pseudo-intrinsic generation.

## Essential Code Snippets for Debugging

Use these utilities to extract frames, generate test intrinsics, and verify pose consistency.

**Extract Frames from Video**

```python
import cv2, pathlib

def extract_frames(video_path, fps=10):
    cap = cv2.VideoCapture(video_path)
    interval = max(1, int(cap.get(cv2.CAP_PROP_FPS) / fps))
    out_dir = pathlib.Path(video_path).with_suffix('_frames')
    out_dir.mkdir(exist_ok=True)
    idx = 0
    while cap.isOpened():
        ret, frame = cap.read()
        if not ret: 
            break
        if idx % interval == 0:
            cv2.imwrite(str(out_dir / f"{idx:06d}.png"), frame)
        idx += 1
    cap.release()
    print(f"Extracted {len(list(out_dir.glob('*.png')))} frames to {out_dir}")

extract_frames("my_clip.mp4", fps=10)

```

**Generate Pseudo-Intrinsics**

```python
import numpy as np
from lingbot_map.vis.point_cloud_viewer import generate_pseudo_intrinsics

K = generate_pseudo_intrinsics(h=720, w=1280)  # Returns (3,3) array

print("Pseudo-intrinsics matrix:\n", K)

```

**Verify Pose Encoding Consistency**

```python
import torch
from lingbot_map.utils.pose_enc import (
    extri_intri_to_pose_encoding, 
    pose_encoding_to_extri_intri
)

# Dummy extrinsic (identity) and intrinsic

extr = torch.eye(4)[None, None, :3, :]           # (B=1, S=1, 3, 4)

intr = torch.from_numpy(K)[None, None, :, :]     # (1,1,3,3)

enc = extri_intri_to_pose_encoding(extr, intr, image_size_hw=(720,1280))
extr_rec, intr_rec = pose_encoding_to_extri_intri(enc, image_size_hw=(720,1280))

print(f"Extrinsic error: {torch.norm(extr - extr_rec).item():.2e}")
print(f"Intrinsic error: {torch.norm(intr - intr_rec).item():.2e}")

```

## Key Files to Inspect

Understanding these specific files accelerates debugging:

- **[`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py)**: Entry point for video loading and argument parsing (`--video_path`, `--fps`).
- **[`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py)**: Handles frame extraction caching and the `MAX_TIME_GAP` logic that filters temporally distant frames.
- **[`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py)**: Contains `extri_intri_to_pose_encoding` and `pose_encoding_to_extri_intri`; central to how intrinsics affect pose estimation.
- **[`lingbot_map/vis/point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/point_cloud_viewer.py)**: Implements `generate_pseudo_intrinsics` and the interactive 3D viewer for trajectory inspection.
- **[`lingbot_map/vis/viser_wrapper.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/viser_wrapper.py)**: Renders camera poses and point clouds; useful for spotting visual drift.
- **[`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py)**: Houses projection functions (`proj`, `iproj`) that consume the intrinsics matrix and directly affect point cloud alignment.

## Summary

- **Pseudo-intrinsics** generated from image size assume `focal_length = max(H, W)`, which is rarely correct for real cameras and causes scale drift.
- Verify frame extraction order matches chronological sequence; check `MAX_TIME_GAP` in [`general.py`](https://github.com/Robbyant/lingbot-map/blob/main/general.py) for variable frame-rate videos.
- Use **round-trip testing** on `extri_intri_to_pose_encoding` to catch dimension mismatches before running full reconstruction.
- Override synthetic intrinsics by placing a calibrated [`intrinsics.txt`](https://github.com/Robbyant/lingbot-map/blob/main/intrinsics.txt) file in the dataset root.
- Visualize trajectories using `point_cloud_viewer` with "Show Current Frame" enabled to correlate video content with pose jumps.

## Frequently Asked Questions

### How does LingBot-MAP generate camera intrinsics when no calibration file is provided?

When [`intrinsics.txt`](https://github.com/Robbyant/lingbot-map/blob/main/intrinsics.txt) is absent, the code calls `generate_pseudo_intrinsics` in [`lingbot_map/vis/point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/point_cloud_viewer.py). This function creates a pinhole camera matrix using the image height and width, setting the focal length to `max(H, W)` and the principal point to `(W/2, H/2)`. These assumptions rarely match real sensor characteristics, leading to systematic reconstruction errors.

### Why does my camera trajectory wobble even when the physical camera moved smoothly?

Wobbling typically indicates **temporal undersampling** or **timestamp misalignment**. If `--fps` is set too low, the large motion between frames destabilizes the transformer's relative pose estimation. Alternatively, if `cv2.VideoCapture` extracts frames out of order (common with variable GOP video), the [`poses_c2w.txt`](https://github.com/Robbyant/lingbot-map/blob/main/poses_c2w.txt) sequence becomes temporally inconsistent, causing the pose encoder to predict erratic motion.

### What is the correct format for providing custom intrinsics to override the defaults?

Create a text file named [`intrinsics.txt`](https://github.com/Robbyant/lingbot-map/blob/main/intrinsics.txt) in the same directory as your extracted frames. The file should contain a 3×3 camera matrix (fx, 0, cx; 0, fy, cy; 0, 0, 1) in space-delimited format. When `GeneralDataset` loads the data, it will read this file instead of calling `generate_pseudo_intrinsics`, ensuring the geometry pipeline uses your calibrated parameters.

### How can I verify that frame extraction preserved the correct temporal order?

After running [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py), check the alphabetical sorting of files in the `*_frames/` directory. OpenCV extracts frames sequentially by default, but if your video has non-standard encoding, inspect the frame indices printed by the extraction loop. Ensure the sequence matches the expected chronological order and that no frames are missing between indices, as gaps trigger the `MAX_TIME_GAP` filter in [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py).