# How to Convert Pose Encodings to Camera Extrinsics and Intrinsics in Ling-Bot-Map

> Learn how lingbot-map's pose_encoding_to_extri_intri function converts 9D pose encodings to camera extrinsics and intrinsics by decomposing translation, quaternion rotations, and FoV angles into OpenCV matrices.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-29

---

**The `pose_encoding_to_extri_intri` function in [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py) converts a compact 9‑dimensional pose encoding—containing translation, quaternion rotation, and field‑of‑view angles—into standard OpenCV‑style extrinsic `[R|t]` and intrinsic `K` camera matrices by decomposing the tensor, converting quaternions to rotation matrices, and reconstructing focal lengths from FoV angles.**

The Ling‑Bot‑Map framework by Robbyant/lingbot‑map represents camera poses as a lightweight 9‑dimensional encoding to reduce memory overhead during training and inference. Understanding how to convert pose encodings to camera extrinsics and intrinsics is essential for projecting depth maps, computing photometric flow, and visualizing camera trajectories. This article breaks down the conversion pipeline implemented in the source code, referencing the actual function signatures and file paths from the repository.

## Understanding the Pose Encoding Format

Before conversion, the model stores camera parameters in a compact tensor format that packs three geometric components into a single 9‑dimensional vector.

### The 9‑Dimensional Representation

The pose encoding follows the layout defined in [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py):

| Component | Meaning | Tensor Slice |
|-----------|---------|--------------|
| **T** | Absolute translation vector (3‑D) | `[..., :3]` |
| **Q** | Rotation as quaternion (4‑D) | `[..., 3:7]` |
| **FoV** | Horizontal and vertical field‑of‑view angles (2‑D) | `[..., 7:]` |

This encoding allows the network to predict or store camera motions using only nine float values per frame, significantly reducing parameter count compared to full 3×4 extrinsic and 3×3 intrinsic matrices.

## The Conversion Pipeline in pose_encoding_to_extri_intri

The core conversion logic resides in `pose_encoding_to_extri_intri` at line 72 of [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py). The function takes a batch of pose encodings and reconstructs the full camera matrices through four distinct steps.

### Step 1: Split the Encoding Tensor

The function first slices the input tensor into its constituent parts:

```python
T = pose_encoding[..., :3]          # Translation: B×S×3

quat = pose_encoding[..., 3:7]      # Quaternion: B×S×4

fov_h = pose_encoding[..., 7]       # Vertical FoV: B×S

fov_w = pose_encoding[..., 8]       # Horizontal FoV: B×S

```

Here, `B` represents batch size and `S` represents sequence length.

### Step 2: Quaternion to Rotation Matrix

The conversion relies on `quat_to_mat` from [`lingbot_map/utils/rotation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/rotation.py) to transform the quaternion into a 3×3 rotation matrix:

```python
R = quat_to_mat(quat)               # Output shape: B×S×3×3

```

This rotation matrix `R` represents the camera orientation in world coordinates.

### Step 3: Build the Extrinsic Matrix

The function constructs the OpenCV‑style extrinsic matrix `[R|t]` by concatenating the rotation matrix with the translation vector:

```python
extrinsics = torch.cat([R, T[..., None]], dim=-1)   # Shape: B×S×3×4

```

This 3×4 matrix transforms homogeneous world coordinates into the camera coordinate system.

### Step 4: Reconstruct the Intrinsic Matrix

When `build_intrinsics=True`, the function calculates focal lengths from the FoV angles and image dimensions provided via `image_size_hw=(H, W)`:

**Vertical focal length:**

$$
f_y = \frac{H/2}{\tan(\text{fov}_h/2)}
$$

**Horizontal focal length:**

$$
f_x = \frac{W/2}{\tan(\text{fov}_w/2)}
$$

The implementation assembles the 3×3 intrinsic matrix `K` with the principal point at the image center:

```python
fy = (H / 2.0) / torch.tan(fov_h / 2.0)
fx = (W / 2.0) / torch.tan(fov_w / 2.0)

intrinsics = torch.zeros(pose_encoding.shape[:2] + (3, 3), device=pose_encoding.device)
intrinsics[..., 0, 0] = fx          # Focal length x

intrinsics[..., 1, 1] = fy          # Focal length y

intrinsics[..., 0, 2] = W / 2       # Principal point x

intrinsics[..., 1, 2] = H / 2       # Principal point y

intrinsics[..., 2, 2] = 1.0         # Homogeneous coordinate

```

The function returns a tuple `(extrinsics, intrinsics)`. If `build_intrinsics=False`, intrinsics is set to `None`.

## Practical Usage Examples

### Basic Conversion

```python
import torch
from lingbot_map.utils.pose_enc import pose_encoding_to_extri_intri

# Dummy pose encoding: [B, S, 9]  (translation, quaternion, fov_h, fov_w)

pose_enc = torch.randn(1, 1, 9)

# Image size (height, width) used for intrinsic reconstruction

img_h, img_w = 256, 512

# Convert to extrinsic & intrinsic matrices

extrinsics, intrinsics = pose_encoding_to_extri_intri(
    pose_enc,
    image_size_hw=(img_h, img_w),
    pose_encoding_type="absT_quaR_FoV",
    build_intrinsics=True,
)

print("Extrinsic matrix shape:", extrinsics.shape)   # → (1, 1, 3, 4)

print("Intrinsic matrix shape:", intrinsics.shape)   # → (1, 1, 3, 3)

```

### Integration in Model Forward Pass

In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), the conversion enables depth map projection:

```python

# cur_pose_enc and kf_pose_enc are tensors of shape [B, 1, 9]

cur_ext, cur_int = pose_encoding_to_extri_intri(
    cur_pose_enc, image_size_hw=(H, W)
)
kf_ext, kf_int = pose_encoding_to_extri_intri(
    kf_pose_enc, image_size_hw=(H, W)
)

# cur_ext and kf_ext provide world-to-camera transforms

# cur_int and kf_int provide focal lengths for projection

```

## Where the Conversion is Used in the Codebase

The `pose_encoding_to_extri_intri` function serves as the bridge between the network's lightweight representation and geometric operations requiring standard camera matrices:

- **GCTBase** ([`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py), line 260): Builds extrinsics and intrinsics from predicted pose encodings during training and evaluation.
- **GCTStream** ([`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), line 19): Reconstructs camera matrices to project depth maps and compute photometric flow between frames.
- **Demo script** ([`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py), line 280): Converts model outputs to visualizable camera parameters for trajectory rendering.

## Summary

Converting pose encodings to camera extrinsics and intrinsics in Ling‑Bot‑Map involves these key operations:

- **Decompose** the 9‑D pose encoding into translation `T`, quaternion `Q`, and FoV angles using tensor slicing.
- **Convert** the quaternion to a 3×3 rotation matrix `R` using `quat_to_mat` from [`lingbot_map/utils/rotation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/rotation.py).
- **Construct** the 3×4 extrinsic matrix `[R|t]` by concatenating rotation and translation.
- **Calculate** focal lengths $f_x$ and $f_y$ from FoV angles and image dimensions using the tangent formula.
- **Assemble** the 3×3 intrinsic matrix `K` with focal lengths and principal point at the image center.

This pipeline enables the model to maintain efficient 9‑dimensional pose representations internally while interfacing with standard computer vision libraries that expect full camera matrices.

## Frequently Asked Questions

### How is the principal point determined in the reconstructed intrinsics?

The principal point is fixed at the geometric center of the image. Specifically, the function sets `cx = W / 2` and `cy = H / 2`, where `H` and `W` are the image height and width passed as the `image_size_hw` parameter. This assumes a standard pinhole camera model with the optical center aligned to the image center.

### Can I use pose_encoding_to_extri_intri without building intrinsics?

Yes. Set `build_intrinsics=False` when calling the function. In this mode, the function skips the focal length calculations and returns `None` for the intrinsics tuple element, producing only the 3×4 extrinsic matrix. This is useful when you need only the camera pose for rigid transformation operations.

### Why does the encoding use quaternions instead of rotation matrices?

Quaternions provide a compact, 4‑dimensional representation of rotation that avoids the gimbal lock issues of Euler angles and requires significantly less memory than a 3×3 matrix (4 floats vs. 9). The `quat_to_mat` utility in [`lingbot_map/utils/rotation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/rotation.py) handles the conversion to rotation matrices whenever geometric operations require them.

### What coordinate convention does the extrinsic matrix follow?

The extrinsic matrix follows the OpenCV convention: a 3×4 matrix `[R|t]` where `R` is a 3×3 rotation matrix and `t` is a 3×1 translation vector. This matrix transforms points from world coordinates to camera coordinates using the formula $X_{cam} = R \cdot X_{world} + t$, compatible with standard computer vision pipelines.