# How LingBot-Map Performs Camera Head Iterative Refinement for Pose Estimation

> Learn how LingBot-Map refines camera pose estimation through iterative refinement using transformer trunk. Achieve progressive error reduction for precise pose predictions.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-27

---

**LingBot-Map refines camera pose predictions over multiple iterations by embedding current estimates, generating modulation parameters, and accumulating delta corrections through a transformer trunk, enabling progressive error reduction from coarse to fine precision.**

LingBot-Map is an open-source visual localization system that predicts camera parameters—including translation, rotation (quaternion), and field-of-view—from transformer token features. The **CameraHead** and **CameraCausalHead** classes implement an iterative refinement mechanism that progressively improves pose accuracy over configurable iterations, trading computation time for precision.

## The Six-Step Refinement Pipeline

The core logic resides in `CameraHead.trunk_fn` and `CameraCausalHead.trunk_fn` within [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) (lines 96-112). Each iteration executes a structured update cycle that transforms an initial coarse estimate into a precise camera pose.

### 1. Pose Embedding

The first iteration initializes with a learned "empty" pose token, while subsequent iterations embed the current estimate using `self.embed_pose`. This embedding provides the conditional signal for the refinement step.

### 2. Modulation Parameter Generation

The embedded pose feeds into `self.poseLN_modulation` to generate shift, scale, and gate parameters. These parameters control how the transformer adapts its processing based on the current pose error.

### 3. Adaptive Layer Normalization

The camera token undergoes `self.adaln_norm` combined with modulation operations. This adaptive normalization focuses the network on residual errors rather than absolute values, allowing the model to learn corrections rather than full poses.

### 4. Transformer Context Processing

The modulated token passes through `self.trunk`—a stack of transformer blocks that aggregates contextual information from visual tokens. This shared trunk reuses contextual features across all refinement iterations.

### 5. Delta Prediction and Accumulation

A lightweight MLP (`self.pose_branch`) predicts pose deltas, which are added to the current estimate via `pred_pose_enc = pred_pose_enc + pred_pose_enc_delta`. This accumulation mechanism ensures that each iteration refines the previous result.

### 6. Final Activation

The `activate_pose` function (defined in [`lingbot_map/heads/head_act.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/head_act.py)) converts raw predictions into valid translation, quaternion, and FOV values, constraining outputs to geometrically feasible camera parameters.

## Source Code Architecture

The implementation spans multiple modules that work together to enable iterative refinement:

- [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py): Contains the `trunk_fn` method (lines 99-112) implementing the iteration loop for both `CameraHead` and `CameraCausalHead` classes.
- [`lingbot_map/heads/head_act.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/head_act.py): Implements `activate_pose` for converting network outputs to valid camera parameters.
- [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py): Integrates the `CameraCausalHead` into the full GCT streaming architecture for real-time applications.
- [`benchmark/benchmark/evaluation/trajectory.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/evaluation/trajectory.py): Provides evaluation metrics that quantify refinement gains on TUM, Oxford, and KITTI datasets.

## Configuring the Refinement Process

The `num_iterations` parameter controls the speed-accuracy trade-off. You can configure this during initialization or modify it at runtime:

```python
import torch
from lingbot_map.heads.camera_head import CameraCausalHead

# Create dummy token tensor (B, N, C) representing visual features

dummy_tokens = torch.randn(2, 128, 2048)  # B=2, N=128 tokens, C=2048 features

tokens_list = [dummy_tokens]  # Model expects a list; last entry is used

# Instantiate with default 4 iterations for maximum accuracy

camera_head = CameraCausalHead(
    dim_in=2048,
    trunk_depth=4,
    num_heads=16,
    num_iterations=4,  # Higher values improve accuracy but increase latency

)

# Forward pass returns a list of pose predictions (one per iteration)

pose_predictions = camera_head(tokens_list)

# Each entry is tensor of shape (B, 9) = [Tx,Ty,Tz, Qx,Qy,Qz,Qw, FoVx, FoVy]

final_pose = pose_predictions[-1]
print(final_pose.shape)  # -> torch.Size([2, 9])

```

### Adjusting Iterations at Runtime

For faster inference with coarser estimates, reduce the iteration count:

```python
camera_head.num_iterations = 1  # Fast inference mode

pose_predictions = camera_head(tokens_list)

```

## Why Iterative Refinement Improves Accuracy

This design mirrors classic coarse-to-fine optimization strategies. By modeling non-linear pose manifolds through sequential shallow corrections rather than a single large jump, the network achieves several advantages:

- **Progressive Error Correction**: Each iteration reduces residual errors from the previous pass, allowing the model to converge on precise solutions.
- **Contextual Feature Reuse**: The transformer trunk processes rich visual context once, while lightweight MLPs handle pose-specific updates efficiently.
- **Controllable Complexity**: The `num_iterations` parameter allows deployment-time decisions between real-time performance (1 iteration) and offline precision (4 iterations).

Empirical results on TUM, Oxford, and KITTI benchmarks demonstrate that the default 4 iterations yield optimal translation and rotation error metrics compared to single-pass prediction.

## Summary

- The refinement loop lives in `trunk_fn` methods within [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) (lines 96-112).
- Each iteration embeds the current pose, generates modulation parameters via `poseLN_modulation`, and predicts a delta correction through `pose_branch`.
- Accumulation happens via `pred_pose_enc = pred_pose_enc + pred_pose_enc_delta`, enabling progressive refinement.
- The `num_iterations` parameter (default 4) controls the explicit speed versus accuracy trade-off.
- Final activations convert raw outputs to valid camera parameters via `activate_pose` in [`lingbot_map/heads/head_act.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/head_act.py).

## Frequently Asked Questions

### What is the default number of iterations in LingBot-Map's CameraHead?

The default value is **4 iterations**, which provides the best pose quality on benchmarks such as TUM, Oxford, and KITTI. You can reduce this to 1 for faster inference at the cost of accuracy, or increase it for offline processing where latency is not critical.

### How does the iterative refinement differ between CameraHead and CameraCausalHead?

Both classes implement identical `trunk_fn` refinement logic. The `CameraCausalHead` variant is optimized for streaming applications within the GCT architecture ([`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py)), while `CameraHead` handles standard batch processing. The iterative mechanism, embedding strategy, and delta accumulation remain consistent between both implementations.

### What pose parameters does the camera head output?

The head outputs a 9-dimensional vector per batch item containing translation `[Tx, Ty, Tz]`, rotation quaternion `[Qx, Qy, Qz, Qw]`, and field-of-view `[FoVx, FoVy]`. These values pass through `activate_pose` to ensure geometric validity before returning to the calling application.

### Can I modify the number of iterations after model initialization?

Yes, you can adjust the `num_iterations` attribute at runtime without rebuilding the model. For example, setting `camera_head.num_iterations = 1` immediately switches to fast inference mode, while reverting to `4` restores maximum accuracy. This flexibility allows the same trained weights to serve both real-time and high-precision use cases.