How Pose Encodings Are Converted to 4×4 Extrinsic Matrices in lingbot-map
The lingbot-map repository converts 9-dimensional pose encodings into 4×4 camera extrinsic matrices by decomposing the tensor into translation and quaternion components, converting the quaternion to a 3×3 rotation matrix via quat_to_mat, forming a 3×4 extrinsic, and appending a homogeneous bottom row.
The lingbot-map library stores camera poses in a compact encoding format to minimize memory overhead while preserving geometric accuracy. Understanding how pose encodings are converted to 4×4 extrinsic matrices enables precise camera pose estimation for 3D reconstruction and novel view synthesis pipelines.
The Pose Encoding Format
The pose encoding consists of nine values packed into a single tensor:
- T — Absolute translation vector (3 values: x, y, z)
- q — Rotation expressed as a unit quaternion (4 values: w, x, y, z)
- FoV — Vertical and horizontal field-of-view angles (2 values)
This 9-dimensional representation reduces storage requirements compared to full 4×4 matrices while maintaining lossless reconstruction capability.
Step-by-Step Conversion Pipeline
1. Extract Components from the 9-D Tensor
In lingbot_map/utils/pose_enc.py, the function pose_encoding_to_extri_intri splits the input tensor into geometric primitives. Lines 73–78 parse the encoding into translation T, quaternion quat, fov_h, and fov_w:
# Extraction logic from lines 73-78
T = pose_enc[..., 0:3] # Translation (x, y, z)
quat = pose_enc[..., 3:7] # Quaternion (w, x, y, z)
fov_h = pose_enc[..., 7] # Horizontal field of view
fov_w = pose_enc[..., 8] # Vertical field of view
2. Convert Quaternion to Rotation Matrix
The quaternion undergoes conversion to a 3×3 rotation matrix R via the quat_to_mat function implemented in lingbot_map/utils/rotation.py. Lines 18–22 perform the mathematical transformation from quaternion coefficients to an orthonormal rotation matrix using standard Hamilton quaternion-to-matrix formulae.
3. Form the 3×4 Extrinsic Matrix
At line 79 in pose_enc.py, the rotation matrix R and translation vector T are concatenated along the last dimension to construct the 3×4 extrinsic matrix [R | T]:
# Operation performed at line 79
extrinsic_3x4 = torch.cat([R, T.unsqueeze(-1)], dim=-1) # Shape: (..., 3, 4)
This compact representation stores the camera-to-world transformation.
4. Upgrade to 4×4 Homogeneous Matrix
While pose_encoding_to_extri_intri returns the 3×4 form, converting to a full 4×4 homogeneous matrix requires prepending the bottom row [0, 0, 0, 1]. This standard homogeneous transformation enables matrix multiplication with 4-D homogeneous coordinates:
extrinsic_4x4 = [[R, T], [0, 0, 0, 1]]
Implementation Reference
The conversion pipeline spans three key utility modules:
lingbot_map/utils/pose_enc.py— Implementspose_encoding_to_extri_intrifor encoding-to-matrix conversion and the inverseextri_intri_to_pose_encodinglingbot_map/utils/rotation.py— Providesquat_to_matandmat_to_quatfor bidirectional quaternion-matrix conversionslingbot_map/utils/geometry.py— Containsclosed_form_inverse_se3for relative pose calculations using the resulting extrinsics
Working Code Example
The following implementation demonstrates the complete conversion from pose encoding to 4×4 extrinsic matrix according to the lingbot-map source code:
import torch
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri
# Example pose encoding: B=1, S=1 (9-dim vector)
pose_enc = torch.tensor([[
[1.2, 0.5, -0.3, # translation T
0.7071, 0.0, 0.7071, 0.0, # quaternion q (w, x, y, z)
1.0, 1.2] # FoV_h, FoV_w (radians)
]])
# Reconstruct 3×4 extrinsic (and intrinsics, if needed)
extrinsics_3x4, intrinsics = pose_encoding_to_extri_intri(
pose_enc,
image_size_hw=(480, 640), # required for intrinsics
build_intrinsics=False # we only need the extrinsic here
)
# Convert to 4×4 homogeneous matrix
extrinsics_4x4 = torch.cat(
[extrinsics_3x4,
torch.tensor([[[0.0, 0.0, 0.0, 1.0]]], dtype=extrinsics_3x4.dtype)],
dim=1
)
print(extrinsics_4x4)
Summary
- Pose encodings in lingbot-map store camera poses as compact 9-dimensional tensors containing translation, quaternion rotation, and field-of-view parameters according to the implementation in
pose_enc.py. - Component extraction occurs at lines 73–78, separating the encoding into
T,quat,fov_h, andfov_wvariables. - Quaternion conversion utilizes
quat_to_matfromrotation.py(lines 18–22) to generate orthonormal 3×3 rotation matrices. - Matrix construction follows the homogeneous transformation pattern
extrinsic_4x4 = [[R, T], [0, 0, 0, 1]], upgrading the internal 3×4 representation for 4-D geometric operations.
Frequently Asked Questions
What is the exact structure of the 9-dimensional pose encoding?
The pose encoding arranges nine float values in the specific order: three translation coordinates (x, y, z), four quaternion components (w, x, y, z), and two field-of-view angles (horizontal, vertical). This ordering is strictly enforced in lingbot_map/utils/pose_enc.py to ensure correct tensor slicing at lines 73–78.
Why does lingbot-map use quaternions instead of rotation matrices for storage?
Quaternions provide a compact, non-degenerate representation of 3D rotations using only four parameters versus nine for rotation matrices, while avoiding gimbal lock issues associated with Euler angles. The quat_to_mat function converts these to matrices only when geometric transformations are required for computational efficiency.
How do I obtain the 4×4 extrinsic matrix if the function returns 3×4?
Append a homogeneous bottom row [0, 0, 0, 1] to the 3×4 matrix returned by pose_encoding_to_extri_intri. In PyTorch, use torch.cat to concatenate the bottom row along dimension 1, enabling standard 4×4 matrix multiplication with homogeneous coordinates as shown in the code example above.
What are the field-of-view parameters used for during conversion?
While the FoV parameters (fov_h and fov_w) primarily drive intrinsic matrix computation when build_intrinsics=True, they are preserved in the pose encoding to maintain a complete camera specification. The extrinsic conversion itself depends only on the translation and quaternion components extracted in steps 1–3.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →