How to Debug Pose Drift Issues in LingBot-MAP When Running on Custom Video with Unknown Camera Intrinsics

Pose drift in LingBot-MAP when processing custom video without calibration files typically stems from incorrect pseudo-intrinsics, temporal undersampling, or frame ordering mismatches that propagate through the pose-encoding pipeline.

LingBot-MAP builds dense 3D reconstructions from RGB-D streams, but feeding it a custom video with unknown camera intrinsics forces the pipeline to infer parameters that rarely match real sensor characteristics. This guide examines the specific failure points in the Robbyant/lingbot-map codebase—from frame extraction in demo.py to pose encoding in utils/pose_enc.py—to help you diagnose and eliminate trajectory drift.

How the Pipeline Handles Unknown Intrinsics

When you supply a video path via --video_path, the system executes a multi-stage pipeline where each component assumes accurate camera parameters. If intrinsics are missing, the code synthesizes fallback values that often introduce systematic errors.

Frame Extraction and Dataset Loading

The entry point demo.load_images uses OpenCV (cv2.VideoCapture) to extract frames from your video and caches them in a sibling *_frames directory. The benchmark.datasets.GeneralDataset class then wraps these frames, providing a uniform __getitem__ interface and handling the optional video_fps parameter for temporal sampling.

Pseudo-Intrinsics Generation

If no calibration file is found, the viewer falls back to lingbot_map.vis.point_cloud_viewer.generate_pseudo_intrinsics. This function assumes a pinhole camera model with focal length set to max(H, W)—a heuristic that is almost always incorrect for real cameras and introduces scale errors immediately.

Pose Encoding and Decoding

Camera extrinsics and intrinsics are packed into a 9-dimensional pose encoding (absT_quaR_FoV) by extri_intri_to_pose_encoding in lingbot_map/utils/pose_enc.py. During inference, pose_encoding_to_extri_intri reverses this conversion. If the image_size_hw parameter mismatches the actual frame dimensions, or if the intrinsics matrix contains incorrect fx, fy, cx, cy values, the decoded poses drift progressively.

Geometry Processing

Functions in lingbot_map/utils/geometry.py—including proj, iproj, and unproject_depth_map_to_point_map—rely directly on the intrinsics matrix. Errors here manifest as misaligned point clouds and wobbling camera trajectories in the visualization stage.

Root Causes of Pose Drift with Unknown Intrinsics

When running on custom video, five specific issues typically trigger drift:

  • Incorrect Intrinsics: The pseudo-intrinsics assume a focal length of max(H, W), creating systematic scale errors in the back-projected 3D points.
  • Temporal Undersampling: Setting --fps too low yields large inter-frame motion, destabilizing the transformer’s relative-pose estimation.
  • Missing Timestamps: The poses_c2w.txt file expects strict chronological order. If frame extraction re-orders frames (e.g., due to variable GOP lengths), the pose encoder receives temporally jittered input.
  • Inconsistent Image Size: OpenCV may resize frames during decoding while the intrinsics matrix still reflects the original resolution, causing projection mismatches.
  • Pipeline Mis-alignment: The dataset loader returns tensors of shape [B, S, C, H, W] while the model expects [B, S, H, W, C]. This subtle shape mismatch perturbs attention masks and corrupts pose predictions.

Step-by-Step Debugging Workflow

Follow this sequence to isolate the source of drift in your custom video input.

1. Verify Frame Extraction

Run the extraction and check the output directory:

python demo.py --video_path my_clip.mp4 --fps 10 --mode windowed

After execution, inspect the generated my_clip_frames/ directory. The file count must match the len(paths) value printed by demo.py. Mismatches indicate dropped frames or decoding errors.

2. Inspect the Intrinsics Matrix

Check what parameters the system is using:

from lingbot_map.vis.point_cloud_viewer import generate_pseudo_intrinsics
K = generate_pseudo_intrinsics(h=720, w=1280)
print(K)

If you know your sensor’s approximate focal length, replace the pseudo-intrinsics with a calibrated matrix. Store this as intrinsics.txt in the dataset root to override the fallback generation.

3. Validate Pose Encoding Round-Trip

Errors in the encoding/decoding functions corrupt reconstruction. Test the conversion:

import torch
from lingbot_map.utils.pose_enc import extri_intri_to_pose_encoding, pose_encoding_to_extri_intri

# Assume extrinsics and intrinsics are tensors from a single frame

encoding = extri_intri_to_pose_encoding(extrinsics, intrinsics, image_size_hw=(720,1280))

# Round-trip check

e_rec, i_rec = pose_encoding_to_extri_intri(encoding, image_size_hw=(720,1280))
assert torch.allclose(extrinsics, e_rec, atol=1e-5)

Any deviation indicates a bug in the pose conversion logic or incorrect image dimensions.

4. Visualize the Trajectory

Launch the interactive viewer:

python -m lingbot_map.vis.point_cloud_viewer --data_path my_clip_frames/

Enable the "Show Current Frame" checkbox to compare the live video frame against the projected point cloud. Large spatial jumps between consecutive frames correlate with faulty intrinsics or timestamp inconsistencies.

5. Check Timestamp Gaps

In benchmark/datasets/general.py, the MAX_TIME_GAP constant controls the maximum allowable temporal separation between image-pose pairs. If your source video has variable frame-rate, increase this limit or supply a uniform timestamp list to prevent the dataset from filtering valid frames.

6. Calibrate Intrinsics (Optional)

If you have access to a calibration target, run OpenCV calibration on extracted frames and write the result to intrinsics.txt:


# Example: Saving a calibrated intrinsics matrix

import numpy as np
np.savetxt('intrinsics.txt', K, fmt='%.6f')

Place this file in the same directory as your frames to bypass pseudo-intrinsic generation.

Essential Code Snippets for Debugging

Use these utilities to extract frames, generate test intrinsics, and verify pose consistency.

Extract Frames from Video

import cv2, pathlib

def extract_frames(video_path, fps=10):
    cap = cv2.VideoCapture(video_path)
    interval = max(1, int(cap.get(cv2.CAP_PROP_FPS) / fps))
    out_dir = pathlib.Path(video_path).with_suffix('_frames')
    out_dir.mkdir(exist_ok=True)
    idx = 0
    while cap.isOpened():
        ret, frame = cap.read()
        if not ret: 
            break
        if idx % interval == 0:
            cv2.imwrite(str(out_dir / f"{idx:06d}.png"), frame)
        idx += 1
    cap.release()
    print(f"Extracted {len(list(out_dir.glob('*.png')))} frames to {out_dir}")

extract_frames("my_clip.mp4", fps=10)

Generate Pseudo-Intrinsics

import numpy as np
from lingbot_map.vis.point_cloud_viewer import generate_pseudo_intrinsics

K = generate_pseudo_intrinsics(h=720, w=1280)  # Returns (3,3) array

print("Pseudo-intrinsics matrix:\n", K)

Verify Pose Encoding Consistency

import torch
from lingbot_map.utils.pose_enc import (
    extri_intri_to_pose_encoding, 
    pose_encoding_to_extri_intri
)

# Dummy extrinsic (identity) and intrinsic

extr = torch.eye(4)[None, None, :3, :]           # (B=1, S=1, 3, 4)

intr = torch.from_numpy(K)[None, None, :, :]     # (1,1,3,3)

enc = extri_intri_to_pose_encoding(extr, intr, image_size_hw=(720,1280))
extr_rec, intr_rec = pose_encoding_to_extri_intri(enc, image_size_hw=(720,1280))

print(f"Extrinsic error: {torch.norm(extr - extr_rec).item():.2e}")
print(f"Intrinsic error: {torch.norm(intr - intr_rec).item():.2e}")

Key Files to Inspect

Understanding these specific files accelerates debugging:

Summary

  • Pseudo-intrinsics generated from image size assume focal_length = max(H, W), which is rarely correct for real cameras and causes scale drift.
  • Verify frame extraction order matches chronological sequence; check MAX_TIME_GAP in general.py for variable frame-rate videos.
  • Use round-trip testing on extri_intri_to_pose_encoding to catch dimension mismatches before running full reconstruction.
  • Override synthetic intrinsics by placing a calibrated intrinsics.txt file in the dataset root.
  • Visualize trajectories using point_cloud_viewer with "Show Current Frame" enabled to correlate video content with pose jumps.

Frequently Asked Questions

How does LingBot-MAP generate camera intrinsics when no calibration file is provided?

When intrinsics.txt is absent, the code calls generate_pseudo_intrinsics in lingbot_map/vis/point_cloud_viewer.py. This function creates a pinhole camera matrix using the image height and width, setting the focal length to max(H, W) and the principal point to (W/2, H/2). These assumptions rarely match real sensor characteristics, leading to systematic reconstruction errors.

Why does my camera trajectory wobble even when the physical camera moved smoothly?

Wobbling typically indicates temporal undersampling or timestamp misalignment. If --fps is set too low, the large motion between frames destabilizes the transformer's relative pose estimation. Alternatively, if cv2.VideoCapture extracts frames out of order (common with variable GOP video), the poses_c2w.txt sequence becomes temporally inconsistent, causing the pose encoder to predict erratic motion.

What is the correct format for providing custom intrinsics to override the defaults?

Create a text file named intrinsics.txt in the same directory as your extracted frames. The file should contain a 3×3 camera matrix (fx, 0, cx; 0, fy, cy; 0, 0, 1) in space-delimited format. When GeneralDataset loads the data, it will read this file instead of calling generate_pseudo_intrinsics, ensuring the geometry pipeline uses your calibrated parameters.

How can I verify that frame extraction preserved the correct temporal order?

After running demo.py, check the alphabetical sorting of files in the *_frames/ directory. OpenCV extracts frames sequentially by default, but if your video has non-standard encoding, inspect the frame indices printed by the extraction loop. Ensure the sequence matches the expected chronological order and that no frames are missing between indices, as gaps trigger the MAX_TIME_GAP filter in benchmark/datasets/general.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →