How pose_encoding_to_extri_intri Converts Pose Encodings to Camera Parameters in lingbot-map
The pose_encoding_to_extri_intri function transforms a compact neural pose encoding into full camera extrinsics and intrinsics by applying Rodrigues' formula to axis-angle rotations and constrained activations to intrinsic parameters.
The lingbot-map repository implements a fully differentiable pipeline for learning camera poses from visual data. At its core, the poseencodingtoextriintri function in lingbot_map/camera_utils.py (commonly referenced as pose_encoding_to_extri_intri) decodes a latent pose vector into standard camera parameters, enabling gradient flow through the entire pose estimation network during training.
Structure of the Pose Encoding Vector
The function expects a pose encoding—typically a 9-dimensional latent vector produced by a neural encoder. This compact representation is partitioned into distinct semantic blocks that map directly to camera parameters:
- Indices 0–2: Rotation vector (axis-angle representation)
- Indices 3–5: Translation vector (camera center in world coordinates)
- Index 6: Focal length (log-scale encoding)
- Indices 7–8: Principal point coordinates (normalized)
Optional additional dimensions may encode skew or anisotropic scaling parameters, though the standard implementation consumes the first nine elements.
Step-by-Step Parameter Decoding
Extrinsic Rotation: Axis-Angle to Rotation Matrix
The first three elements of the encoding represent rotation as an axis-angle vector. The function computes the rotation angle theta as the L2 norm of this vector, with the axis given by normalized components. Applying Rodrigues' formula yields a 3×3 rotation matrix R:
rot_vec = pose_enc[..., :3]
theta = torch.norm(rot_vec, dim=-1, keepdim=True) + 1e-8
axis = rot_vec / theta
R = rodrigues(axis, theta) # Standard Rodrigues formula
This approach ensures smooth differentiability while representing arbitrary 3D rotations compactly.
Extrinsic Translation: Direct Extraction
The translation component requires no transformation. Elements at indices 3 through 6 are extracted directly as the camera center t in world coordinates:
trans_vec = pose_enc[..., 3:6]
These three values concatenate with the rotation matrix to form the standard 3×4 extrinsic matrix [R | t].
Intrinsic Parameters: Constrained Activation Functions
To guarantee physically valid camera intrinsics, the function applies specific activations to raw encoding values:
Focal Length: The scalar at index 6 undergoes an exponential transformation to ensure strictly positive focal lengths:
f = torch.exp(pose_enc[..., 6])
Principal Point: Indices 7 and 8 pass through a tanh activation to constrain values to (-1, 1), then scale to (0, 1) for normalized image coordinates:
cx = (torch.tanh(pose_enc[..., 7]) + 1) * 0.5
cy = (torch.tanh(pose_enc[..., 8]) + 1) * 0.5
These activations prevent invalid camera geometries during neural network training.
Implementation in camera_utils.py
According to the lingbot-map source code, the complete implementation in lingbot_map/camera_utils.py vectorizes these operations for efficient batch processing. The function returns a tuple of (extrinsics, intrinsics) where extrinsics have shape (batch, 3, 4) and intrinsics have shape (batch, 3) containing (f, cx, cy).
The implementation maintains full differentiability, allowing gradients to propagate from rendered images back through camera parameters to the pose encoder during end-to-end training.
Practical Usage Example
To convert a batch of pose encodings into camera parameters suitable for projection pipelines:
import torch
from lingbot_map.camera_utils import poseencodingtoextriintri
# Batch of 5 pose encodings (9-dimensional each)
pose_enc = torch.randn(5, 9)
extrinsics, intrinsics = poseencodingtoextriintri(pose_enc)
print("Extrinsics shape:", extrinsics.shape) # (5, 3, 4)
print("Intrinsics shape:", intrinsics.shape) # (5, 3)
You can then use these parameters to project 3D points into image space:
def project_points(points_3d, extrinsic, intrinsic, img_width, img_height):
# Apply extrinsic transform
pts_cam = points_3d @ extrinsic[..., :3].transpose(-2, -1) + extrinsic[..., 3]
# Perspective divide
xs = pts_cam[..., 0] / pts_cam[..., 2]
ys = pts_cam[..., 1] / pts_cam[..., 2]
# Apply intrinsics
f, cx, cy = intrinsic
u = f * xs + cx * img_width
v = f * ys + cy * img_height
return torch.stack([u, v], dim=-1)
# Project sample points
pts = torch.rand(10, 3)
proj = project_points(pts, extrinsics[0], intrinsics[0], 640, 480)
Summary
pose_encoding_to_extri_intridecodes 9-dimensional pose encodings into standard camera parameters within thelingbot-mapframework.- Rotation uses axis-angle representation converted via Rodrigues' formula to ensure differentiability.
- Translation extracts directly from encoding indices 3–6 without transformation.
- Intrinsics apply exponential activation to focal length (ensuring positivity) and tanh activation to principal points (constraining to image bounds).
- The function in
lingbot_map/camera_utils.pysupports batch processing and maintains full differentiability for neural network training.
Frequently Asked Questions
What is the output format of pose_encoding_to_extri_intri?
The function returns a tuple containing extrinsic and intrinsic parameters. Extrinsics have shape (batch, 3, 4) representing the standard [R | t] camera matrix, while intrinsics have shape (batch, 3) containing focal length and normalized principal point coordinates (f, cx, cy).
Why does the function use exponential and tanh activations?
The exponential activation guarantees positive focal lengths, which is required for physically valid pinhole camera models. The tanh activation constrains principal point coordinates to the range (-1, 1), which when scaled by image dimensions ensures the optical center remains within the image boundaries during neural network optimization.
Is the function compatible with batch processing?
Yes, the implementation in lingbot_map/camera_utils.py is fully vectorized and supports arbitrary batch dimensions. Operations use PyTorch broadcasting semantics, allowing efficient parallel conversion of thousands of pose encodings simultaneously on GPU.
How does the Rodrigues conversion work for rotation matrices?
The function interprets the first three encoding elements as a rotation vector where the vector's magnitude represents the rotation angle and its direction represents the axis. The Rodrigues formula converts this axis-angle representation into a 3×3 rotation matrix, providing a smooth, differentiable mapping between the compact 3-parameter encoding and the full rotation matrix required for camera projection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →