# How CameraCausalHead Implements Iterative Pose Refinement in lingbot-map

> Discover how CameraCausalHead uses iterative pose refinement for lingbot-map. Learn about progressive accuracy improvements enabled by its transformer trunk and adaptive layer normalization.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: internals
- Published: 2026-07-25

---

**CameraCausalHead refines camera pose estimates over configurable iterations by recursively feeding pose encodings through a shared transformer trunk with adaptive layer normalization, enabling progressive accuracy improvements without increasing model parameters.**

The `CameraCausalHead` class in the `Robbyant/lingbot-map` repository performs camera pose prediction through an iterative refinement strategy rather than single-shot estimation. Located in [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py), this architectural component progressively updates translation, rotation (quaternion), and focal length estimates over multiple processing steps using the same network weights.

## The Iterative Refinement Architecture

The core mechanism resides in the `trunk_fn` method (lines 40-78 of [`camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/camera_head.py)), which implements a configurable iteration loop controlled by the `num_iterations` parameter (default 4). Each iteration reuses the same network modules—the transformer trunk, modulation MLP, and pose branch—making the approach computationally efficient while allowing progressive refinement of the 9-dimensional pose vector `[Tx, Ty, Tz, qx, qy, qz, qw, fov]`.

### Initialization and Empty Pose Tokens

On the first iteration, when `pred_pose_enc` is `None`, the model initializes estimation using learnable **empty pose tokens** (`self.empty_pose_tokens`). These tokens are projected to the model dimension via `self.embed_pose` (lines 43-48), providing a learned prior for the refinement process rather than starting from zero initialization.

### Adaptive Layer Normalization Modulation

Each iteration applies **AdaLN-style modulation** through the `self.poseLN_modulation` MLP. This network predicts three modulation vectors—**shift**, **scale**, and **gate**—by splitting the output with `chunk(3, dim=-1)` (lines 52-54). The camera tokens undergo affine-free LayerNorm (`self.adaln_norm`) before modulation via the `modulate` function (lines 49-55), allowing the network to conditionally normalize features based on the current pose estimate.

### Transformer Trunk Processing

The modulated tokens flow through `self.trunk`, an `nn.Sequential` stack of `CameraBlock` transformer blocks defined in [`lingbot_map/layers/block.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/block.py) (lines 22-28). In causal mode, each block receives 3-D RoPE positions, video masks, and KV-cache entries (lines 64-68). This shared trunk processes the conditioned features identically across all iterations, preserving parameter efficiency.

### Delta Prediction and Residual Updates

After trunk processing, a lightweight MLP (`self.pose_branch`) predicts a pose encoding delta (`pred_pose_enc_delta`). The update follows a residual formulation: `pred_pose_enc = pred_pose_enc + pred_pose_enc_delta` (lines 68-74). This incremental correction strategy enables fine-grained adjustments rather than full re-estimation, with gradients flowing through the accumulated encoding across iterations.

### Pose Activation and Output Generation

The accumulated encoding converts to interpretable parameters via `activate_pose` from [`head_act.py`](https://github.com/Robbyant/lingbot-map/blob/main/head_act.py) (lines 76-79). This function applies activation functions—typically linear for translation, exponential for focal length, and normalization for quaternions—to produce the final pose. The loop appends each iteration's result to `pred_pose_enc_list`, returning the complete refinement trajectory to the caller.

## Causal Inference and Streaming Support

For streaming applications, `CameraCausalHead` supports **causal inference mode** via the `causal_inference` flag. When enabled, the method initializes a KV-cache (`self.kv_cache`) in `forward` (lines 84-92) and passes cache entries to each transformer block during trunk processing (line 57). The cache indexes by iteration (`self.kv_cache[i]`), preserving attention keys and values across frames while maintaining the iterative refinement structure.

Optional **3-D Rotary Position Embeddings (RoPE)** enhance temporal awareness in streaming scenarios. When `enable_3d_rope=True`, `WanRotaryPosEmbed` from [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) generates frame-specific positional encodings (lines 106-118 and 16-38), injected into each `CameraBlock` during the trunk forward pass.

## Code Example: Configuring Iterative Refinement

```python
import torch
from lingbot_map.heads.camera_head import CameraCausalHead

# Dummy token stream: (B, N, C) where N includes a camera token at index 0

B, N, C = 2, 10, 2048
tokens = torch.randn(B, N, C)

# Build a list as expected by the head (e.g., outputs from previous layers)

aggregated_tokens = [torch.randn(B, N, C), tokens]   # last entry is used

# Instantiate the head with 4 refinement iterations

cam_head = CameraCausalHead(
    dim_in=C,
    trunk_depth=4,
    num_iterations=4,          # Configurable iteration count

    enable_3d_rope=False,     # Disable for non-streaming inference

)

# Forward pass returns list of poses, one per iteration

poses_per_iter = cam_head(aggregated_tokens)

# Each pose tensor shape: (B, 9) → [Tx,Ty,Tz, qx,qy,qz,qw, fov]

for i, pose in enumerate(poses_per_iter, start=1):
    print(f"Iteration {i}: pose = {pose.shape}")

```

This example demonstrates feeding aggregated token tensors into the head, configuring the refinement depth via `num_iterations`, and retrieving progressive pose estimates. The final list contains intermediate predictions, allowing applications to trade speed for accuracy by selecting earlier iterations or leverage the full sequence for maximum precision.

## Summary

- **CameraCausalHead** implements iterative pose refinement in [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) through a configurable loop in `trunk_fn`.
- Each iteration applies **adaptive layer normalization modulation** (shift, scale, gate) conditioned on the current pose estimate via `self.poseLN_modulation`.
- **Residual delta prediction** enables incremental updates: `pred_pose_enc = pred_pose_enc + pred_pose_enc_delta` (lines 68-74).
- The **shared transformer trunk** (`self.trunk`) processes modulated tokens through `CameraBlock` layers without increasing parameter count across iterations.
- **Causal inference support** includes KV-caching and optional 3-D RoPE for streaming video applications.
- The `num_iterations` parameter (default 4) controls the speed-accuracy trade-off.

## Frequently Asked Questions

### How many iterations does CameraCausalHead use by default?

The default configuration uses **4 iterations**, specified by the `num_iterations=4` parameter in `CameraCausalHead.__init__`. You can reduce this to 1-2 iterations for faster inference or increase to 6-8 for higher accuracy at the cost of computational overhead, as each iteration reuses the same trunk weights defined in [`lingbot_map/layers/block.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/block.py).

### What is the purpose of the empty pose tokens in the first iteration?

When `pred_pose_enc` is `None` during the first iteration, **empty pose tokens** (`self.empty_pose_tokens`) provide a learned initial state. These learnable parameters project to the model dimension via `self.embed_pose` (lines 43-48), giving the network a data-driven starting point rather than zeros or random initialization.

### How does the modulation mechanism work in CameraCausalHead?

The **modulation MLP** (`self.poseLN_modulation`) generates three vectors—shift, scale, and gate—chunked from the MLP output (lines 52-54). These modulate the camera tokens through adaptive LayerNorm: the shift and scale adjust normalized features, while the gate controls residual flow. This **AdaLN-style conditioning** adapts normalization statistics based on the current pose estimate, similar to techniques used in modern diffusion models.

### Can CameraCausalHead process streaming video inputs?

Yes, when `causal_inference=True`, the head initializes a **KV-cache** (lines 84-92) that preserves attention keys and values across frames. Combined with optional **3-D RoPE** positional encodings from [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py), this enables efficient streaming inference where the model processes video frames sequentially while maintaining temporal context through the iterative refinement loop.