# How Pose Encodings Are Converted to 4×4 Extrinsic Matrices in lingbot-map

> Learn how lingbot-map converts 9D pose encodings to 4x4 extrinsic matrices. Discover the tensor decomposition, quaternion conversion, and matrix formation process.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: internals
- Published: 2026-07-24

---

**The lingbot-map repository converts 9-dimensional pose encodings into 4×4 camera extrinsic matrices by decomposing the tensor into translation and quaternion components, converting the quaternion to a 3×3 rotation matrix via `quat_to_mat`, forming a 3×4 extrinsic, and appending a homogeneous bottom row.**

The lingbot-map library stores camera poses in a compact encoding format to minimize memory overhead while preserving geometric accuracy. Understanding how pose encodings are converted to 4×4 extrinsic matrices enables precise camera pose estimation for 3D reconstruction and novel view synthesis pipelines.

## The Pose Encoding Format

The **pose encoding** consists of nine values packed into a single tensor:

- **T** — Absolute translation vector (3 values: x, y, z)
- **q** — Rotation expressed as a unit quaternion (4 values: w, x, y, z)
- **FoV** — Vertical and horizontal field-of-view angles (2 values)

This 9-dimensional representation reduces storage requirements compared to full 4×4 matrices while maintaining lossless reconstruction capability.

## Step-by-Step Conversion Pipeline

### 1. Extract Components from the 9-D Tensor

In [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py), the function `pose_encoding_to_extri_intri` splits the input tensor into geometric primitives. Lines 73–78 parse the encoding into translation `T`, quaternion `quat`, `fov_h`, and `fov_w`:

```python

# Extraction logic from lines 73-78

T = pose_enc[..., 0:3]      # Translation (x, y, z)

quat = pose_enc[..., 3:7]   # Quaternion (w, x, y, z)

fov_h = pose_enc[..., 7]    # Horizontal field of view

fov_w = pose_enc[..., 8]    # Vertical field of view

```

### 2. Convert Quaternion to Rotation Matrix

The quaternion undergoes conversion to a 3×3 rotation matrix `R` via the `quat_to_mat` function implemented in [`lingbot_map/utils/rotation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/rotation.py). Lines 18–22 perform the mathematical transformation from quaternion coefficients to an orthonormal rotation matrix using standard Hamilton quaternion-to-matrix formulae.

### 3. Form the 3×4 Extrinsic Matrix

At line 79 in [`pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/pose_enc.py), the rotation matrix `R` and translation vector `T` are concatenated along the last dimension to construct the 3×4 extrinsic matrix `[R | T]`:

```python

# Operation performed at line 79

extrinsic_3x4 = torch.cat([R, T.unsqueeze(-1)], dim=-1)  # Shape: (..., 3, 4)

```

This compact representation stores the camera-to-world transformation.

### 4. Upgrade to 4×4 Homogeneous Matrix

While `pose_encoding_to_extri_intri` returns the 3×4 form, converting to a full 4×4 homogeneous matrix requires prepending the bottom row `[0, 0, 0, 1]`. This standard **homogeneous transformation** enables matrix multiplication with 4-D homogeneous coordinates:

`extrinsic_4x4 = [[R, T], [0, 0, 0, 1]]`

## Implementation Reference

The conversion pipeline spans three key utility modules:

- **[`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py)** — Implements `pose_encoding_to_extri_intri` for encoding-to-matrix conversion and the inverse `extri_intri_to_pose_encoding`
- **[`lingbot_map/utils/rotation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/rotation.py)** — Provides `quat_to_mat` and `mat_to_quat` for bidirectional quaternion-matrix conversions
- **[`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py)** — Contains `closed_form_inverse_se3` for relative pose calculations using the resulting extrinsics

## Working Code Example

The following implementation demonstrates the complete conversion from pose encoding to 4×4 extrinsic matrix according to the lingbot-map source code:

```python
import torch
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri

# Example pose encoding: B=1, S=1 (9-dim vector)

pose_enc = torch.tensor([[
    [1.2, 0.5, -0.3,          # translation T

     0.7071, 0.0, 0.7071, 0.0,  # quaternion q (w, x, y, z)

     1.0, 1.2]                # FoV_h, FoV_w (radians)

]])

# Reconstruct 3×4 extrinsic (and intrinsics, if needed)

extrinsics_3x4, intrinsics = pose_encoding_to_extri_intri(
    pose_enc,
    image_size_hw=(480, 640),      # required for intrinsics

    build_intrinsics=False        # we only need the extrinsic here

)

# Convert to 4×4 homogeneous matrix

extrinsics_4x4 = torch.cat(
    [extrinsics_3x4,
     torch.tensor([[[0.0, 0.0, 0.0, 1.0]]], dtype=extrinsics_3x4.dtype)],
    dim=1
)

print(extrinsics_4x4)

```

## Summary

- **Pose encodings** in lingbot-map store camera poses as compact 9-dimensional tensors containing translation, quaternion rotation, and field-of-view parameters according to the implementation in [`pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/pose_enc.py).
- **Component extraction** occurs at lines 73–78, separating the encoding into `T`, `quat`, `fov_h`, and `fov_w` variables.
- **Quaternion conversion** utilizes `quat_to_mat` from [`rotation.py`](https://github.com/Robbyant/lingbot-map/blob/main/rotation.py) (lines 18–22) to generate orthonormal 3×3 rotation matrices.
- **Matrix construction** follows the homogeneous transformation pattern `extrinsic_4x4 = [[R, T], [0, 0, 0, 1]]`, upgrading the internal 3×4 representation for 4-D geometric operations.

## Frequently Asked Questions

### What is the exact structure of the 9-dimensional pose encoding?

The pose encoding arranges nine float values in the specific order: three translation coordinates (x, y, z), four quaternion components (w, x, y, z), and two field-of-view angles (horizontal, vertical). This ordering is strictly enforced in [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py) to ensure correct tensor slicing at lines 73–78.

### Why does lingbot-map use quaternions instead of rotation matrices for storage?

Quaternions provide a compact, non-degenerate representation of 3D rotations using only four parameters versus nine for rotation matrices, while avoiding gimbal lock issues associated with Euler angles. The `quat_to_mat` function converts these to matrices only when geometric transformations are required for computational efficiency.

### How do I obtain the 4×4 extrinsic matrix if the function returns 3×4?

Append a homogeneous bottom row `[0, 0, 0, 1]` to the 3×4 matrix returned by `pose_encoding_to_extri_intri`. In PyTorch, use `torch.cat` to concatenate the bottom row along dimension 1, enabling standard 4×4 matrix multiplication with homogeneous coordinates as shown in the code example above.

### What are the field-of-view parameters used for during conversion?

While the FoV parameters (`fov_h` and `fov_w`) primarily drive intrinsic matrix computation when `build_intrinsics=True`, they are preserved in the pose encoding to maintain a complete camera specification. The extrinsic conversion itself depends only on the translation and quaternion components extracted in steps 1–3.