How to Convert Pose Encodings to Camera Extrinsics and Intrinsics in Ling-Bot-Map
The pose_encoding_to_extri_intri function in lingbot_map/utils/pose_enc.py converts a compact 9‑dimensional pose encoding—containing translation, quaternion rotation, and field‑of‑view angles—into standard OpenCV‑style extrinsic [R|t] and intrinsic K camera matrices by decomposing the tensor, converting quaternions to rotation matrices, and reconstructing focal lengths from FoV angles.
The Ling‑Bot‑Map framework by Robbyant/lingbot‑map represents camera poses as a lightweight 9‑dimensional encoding to reduce memory overhead during training and inference. Understanding how to convert pose encodings to camera extrinsics and intrinsics is essential for projecting depth maps, computing photometric flow, and visualizing camera trajectories. This article breaks down the conversion pipeline implemented in the source code, referencing the actual function signatures and file paths from the repository.
Understanding the Pose Encoding Format
Before conversion, the model stores camera parameters in a compact tensor format that packs three geometric components into a single 9‑dimensional vector.
The 9‑Dimensional Representation
The pose encoding follows the layout defined in lingbot_map/utils/pose_enc.py:
| Component | Meaning | Tensor Slice |
|---|---|---|
| T | Absolute translation vector (3‑D) | [..., :3] |
| Q | Rotation as quaternion (4‑D) | [..., 3:7] |
| FoV | Horizontal and vertical field‑of‑view angles (2‑D) | [..., 7:] |
This encoding allows the network to predict or store camera motions using only nine float values per frame, significantly reducing parameter count compared to full 3×4 extrinsic and 3×3 intrinsic matrices.
The Conversion Pipeline in pose_encoding_to_extri_intri
The core conversion logic resides in pose_encoding_to_extri_intri at line 72 of lingbot_map/utils/pose_enc.py. The function takes a batch of pose encodings and reconstructs the full camera matrices through four distinct steps.
Step 1: Split the Encoding Tensor
The function first slices the input tensor into its constituent parts:
T = pose_encoding[..., :3] # Translation: B×S×3
quat = pose_encoding[..., 3:7] # Quaternion: B×S×4
fov_h = pose_encoding[..., 7] # Vertical FoV: B×S
fov_w = pose_encoding[..., 8] # Horizontal FoV: B×S
Here, B represents batch size and S represents sequence length.
Step 2: Quaternion to Rotation Matrix
The conversion relies on quat_to_mat from lingbot_map/utils/rotation.py to transform the quaternion into a 3×3 rotation matrix:
R = quat_to_mat(quat) # Output shape: B×S×3×3
This rotation matrix R represents the camera orientation in world coordinates.
Step 3: Build the Extrinsic Matrix
The function constructs the OpenCV‑style extrinsic matrix [R|t] by concatenating the rotation matrix with the translation vector:
extrinsics = torch.cat([R, T[..., None]], dim=-1) # Shape: B×S×3×4
This 3×4 matrix transforms homogeneous world coordinates into the camera coordinate system.
Step 4: Reconstruct the Intrinsic Matrix
When build_intrinsics=True, the function calculates focal lengths from the FoV angles and image dimensions provided via image_size_hw=(H, W):
Vertical focal length:
$$ f_y = \frac{H/2}{\tan(\text{fov}_h/2)} $$
Horizontal focal length:
$$ f_x = \frac{W/2}{\tan(\text{fov}_w/2)} $$
The implementation assembles the 3×3 intrinsic matrix K with the principal point at the image center:
fy = (H / 2.0) / torch.tan(fov_h / 2.0)
fx = (W / 2.0) / torch.tan(fov_w / 2.0)
intrinsics = torch.zeros(pose_encoding.shape[:2] + (3, 3), device=pose_encoding.device)
intrinsics[..., 0, 0] = fx # Focal length x
intrinsics[..., 1, 1] = fy # Focal length y
intrinsics[..., 0, 2] = W / 2 # Principal point x
intrinsics[..., 1, 2] = H / 2 # Principal point y
intrinsics[..., 2, 2] = 1.0 # Homogeneous coordinate
The function returns a tuple (extrinsics, intrinsics). If build_intrinsics=False, intrinsics is set to None.
Practical Usage Examples
Basic Conversion
import torch
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri
# Dummy pose encoding: [B, S, 9] (translation, quaternion, fov_h, fov_w)
pose_enc = torch.randn(1, 1, 9)
# Image size (height, width) used for intrinsic reconstruction
img_h, img_w = 256, 512
# Convert to extrinsic & intrinsic matrices
extrinsics, intrinsics = pose_encoding_to_extri_intri(
pose_enc,
image_size_hw=(img_h, img_w),
pose_encoding_type="absT_quaR_FoV",
build_intrinsics=True,
)
print("Extrinsic matrix shape:", extrinsics.shape) # → (1, 1, 3, 4)
print("Intrinsic matrix shape:", intrinsics.shape) # → (1, 1, 3, 3)
Integration in Model Forward Pass
In lingbot_map/models/gct_stream_window_v2.py, the conversion enables depth map projection:
# cur_pose_enc and kf_pose_enc are tensors of shape [B, 1, 9]
cur_ext, cur_int = pose_encoding_to_extri_intri(
cur_pose_enc, image_size_hw=(H, W)
)
kf_ext, kf_int = pose_encoding_to_extri_intri(
kf_pose_enc, image_size_hw=(H, W)
)
# cur_ext and kf_ext provide world-to-camera transforms
# cur_int and kf_int provide focal lengths for projection
Where the Conversion is Used in the Codebase
The pose_encoding_to_extri_intri function serves as the bridge between the network's lightweight representation and geometric operations requiring standard camera matrices:
- GCTBase (
lingbot_map/models/gct_base.py, line 260): Builds extrinsics and intrinsics from predicted pose encodings during training and evaluation. - GCTStream (
lingbot_map/models/gct_stream_window_v2.py, line 19): Reconstructs camera matrices to project depth maps and compute photometric flow between frames. - Demo script (
demo.py, line 280): Converts model outputs to visualizable camera parameters for trajectory rendering.
Summary
Converting pose encodings to camera extrinsics and intrinsics in Ling‑Bot‑Map involves these key operations:
- Decompose the 9‑D pose encoding into translation
T, quaternionQ, and FoV angles using tensor slicing. - Convert the quaternion to a 3×3 rotation matrix
Rusingquat_to_matfromlingbot_map/utils/rotation.py. - Construct the 3×4 extrinsic matrix
[R|t]by concatenating rotation and translation. - Calculate focal lengths $f_x$ and $f_y$ from FoV angles and image dimensions using the tangent formula.
- Assemble the 3×3 intrinsic matrix
Kwith focal lengths and principal point at the image center.
This pipeline enables the model to maintain efficient 9‑dimensional pose representations internally while interfacing with standard computer vision libraries that expect full camera matrices.
Frequently Asked Questions
How is the principal point determined in the reconstructed intrinsics?
The principal point is fixed at the geometric center of the image. Specifically, the function sets cx = W / 2 and cy = H / 2, where H and W are the image height and width passed as the image_size_hw parameter. This assumes a standard pinhole camera model with the optical center aligned to the image center.
Can I use pose_encoding_to_extri_intri without building intrinsics?
Yes. Set build_intrinsics=False when calling the function. In this mode, the function skips the focal length calculations and returns None for the intrinsics tuple element, producing only the 3×4 extrinsic matrix. This is useful when you need only the camera pose for rigid transformation operations.
Why does the encoding use quaternions instead of rotation matrices?
Quaternions provide a compact, 4‑dimensional representation of rotation that avoids the gimbal lock issues of Euler angles and requires significantly less memory than a 3×3 matrix (4 floats vs. 9). The quat_to_mat utility in lingbot_map/utils/rotation.py handles the conversion to rotation matrices whenever geometric operations require them.
What coordinate convention does the extrinsic matrix follow?
The extrinsic matrix follows the OpenCV convention: a 3×4 matrix [R|t] where R is a 3×3 rotation matrix and t is a 3×1 translation vector. This matrix transforms points from world coordinates to camera coordinates using the formula $X_{cam} = R \cdot X_{world} + t$, compatible with standard computer vision pipelines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →