How CameraCausalHead Implements Iterative Pose Refinement in lingbot-map

CameraCausalHead refines camera pose estimates over configurable iterations by recursively feeding pose encodings through a shared transformer trunk with adaptive layer normalization, enabling progressive accuracy improvements without increasing model parameters.

The CameraCausalHead class in the Robbyant/lingbot-map repository performs camera pose prediction through an iterative refinement strategy rather than single-shot estimation. Located in lingbot_map/heads/camera_head.py, this architectural component progressively updates translation, rotation (quaternion), and focal length estimates over multiple processing steps using the same network weights.

The Iterative Refinement Architecture

The core mechanism resides in the trunk_fn method (lines 40-78 of camera_head.py), which implements a configurable iteration loop controlled by the num_iterations parameter (default 4). Each iteration reuses the same network modules—the transformer trunk, modulation MLP, and pose branch—making the approach computationally efficient while allowing progressive refinement of the 9-dimensional pose vector [Tx, Ty, Tz, qx, qy, qz, qw, fov].

Initialization and Empty Pose Tokens

On the first iteration, when pred_pose_enc is None, the model initializes estimation using learnable empty pose tokens (self.empty_pose_tokens). These tokens are projected to the model dimension via self.embed_pose (lines 43-48), providing a learned prior for the refinement process rather than starting from zero initialization.

Adaptive Layer Normalization Modulation

Each iteration applies AdaLN-style modulation through the self.poseLN_modulation MLP. This network predicts three modulation vectors—shift, scale, and gate—by splitting the output with chunk(3, dim=-1) (lines 52-54). The camera tokens undergo affine-free LayerNorm (self.adaln_norm) before modulation via the modulate function (lines 49-55), allowing the network to conditionally normalize features based on the current pose estimate.

Transformer Trunk Processing

The modulated tokens flow through self.trunk, an nn.Sequential stack of CameraBlock transformer blocks defined in lingbot_map/layers/block.py (lines 22-28). In causal mode, each block receives 3-D RoPE positions, video masks, and KV-cache entries (lines 64-68). This shared trunk processes the conditioned features identically across all iterations, preserving parameter efficiency.

Delta Prediction and Residual Updates

After trunk processing, a lightweight MLP (self.pose_branch) predicts a pose encoding delta (pred_pose_enc_delta). The update follows a residual formulation: pred_pose_enc = pred_pose_enc + pred_pose_enc_delta (lines 68-74). This incremental correction strategy enables fine-grained adjustments rather than full re-estimation, with gradients flowing through the accumulated encoding across iterations.

Pose Activation and Output Generation

The accumulated encoding converts to interpretable parameters via activate_pose from head_act.py (lines 76-79). This function applies activation functions—typically linear for translation, exponential for focal length, and normalization for quaternions—to produce the final pose. The loop appends each iteration's result to pred_pose_enc_list, returning the complete refinement trajectory to the caller.

Causal Inference and Streaming Support

For streaming applications, CameraCausalHead supports causal inference mode via the causal_inference flag. When enabled, the method initializes a KV-cache (self.kv_cache) in forward (lines 84-92) and passes cache entries to each transformer block during trunk processing (line 57). The cache indexes by iteration (self.kv_cache[i]), preserving attention keys and values across frames while maintaining the iterative refinement structure.

Optional 3-D Rotary Position Embeddings (RoPE) enhance temporal awareness in streaming scenarios. When enable_3d_rope=True, WanRotaryPosEmbed from lingbot_map/layers/rope.py generates frame-specific positional encodings (lines 106-118 and 16-38), injected into each CameraBlock during the trunk forward pass.

Code Example: Configuring Iterative Refinement

import torch
from lingbot_map.heads.camera_head import CameraCausalHead

# Dummy token stream: (B, N, C) where N includes a camera token at index 0

B, N, C = 2, 10, 2048
tokens = torch.randn(B, N, C)

# Build a list as expected by the head (e.g., outputs from previous layers)

aggregated_tokens = [torch.randn(B, N, C), tokens]   # last entry is used

# Instantiate the head with 4 refinement iterations

cam_head = CameraCausalHead(
    dim_in=C,
    trunk_depth=4,
    num_iterations=4,          # Configurable iteration count

    enable_3d_rope=False,     # Disable for non-streaming inference

)

# Forward pass returns list of poses, one per iteration

poses_per_iter = cam_head(aggregated_tokens)

# Each pose tensor shape: (B, 9) → [Tx,Ty,Tz, qx,qy,qz,qw, fov]

for i, pose in enumerate(poses_per_iter, start=1):
    print(f"Iteration {i}: pose = {pose.shape}")

This example demonstrates feeding aggregated token tensors into the head, configuring the refinement depth via num_iterations, and retrieving progressive pose estimates. The final list contains intermediate predictions, allowing applications to trade speed for accuracy by selecting earlier iterations or leverage the full sequence for maximum precision.

Summary

  • CameraCausalHead implements iterative pose refinement in lingbot_map/heads/camera_head.py through a configurable loop in trunk_fn.
  • Each iteration applies adaptive layer normalization modulation (shift, scale, gate) conditioned on the current pose estimate via self.poseLN_modulation.
  • Residual delta prediction enables incremental updates: pred_pose_enc = pred_pose_enc + pred_pose_enc_delta (lines 68-74).
  • The shared transformer trunk (self.trunk) processes modulated tokens through CameraBlock layers without increasing parameter count across iterations.
  • Causal inference support includes KV-caching and optional 3-D RoPE for streaming video applications.
  • The num_iterations parameter (default 4) controls the speed-accuracy trade-off.

Frequently Asked Questions

How many iterations does CameraCausalHead use by default?

The default configuration uses 4 iterations, specified by the num_iterations=4 parameter in CameraCausalHead.__init__. You can reduce this to 1-2 iterations for faster inference or increase to 6-8 for higher accuracy at the cost of computational overhead, as each iteration reuses the same trunk weights defined in lingbot_map/layers/block.py.

What is the purpose of the empty pose tokens in the first iteration?

When pred_pose_enc is None during the first iteration, empty pose tokens (self.empty_pose_tokens) provide a learned initial state. These learnable parameters project to the model dimension via self.embed_pose (lines 43-48), giving the network a data-driven starting point rather than zeros or random initialization.

How does the modulation mechanism work in CameraCausalHead?

The modulation MLP (self.poseLN_modulation) generates three vectors—shift, scale, and gate—chunked from the MLP output (lines 52-54). These modulate the camera tokens through adaptive LayerNorm: the shift and scale adjust normalized features, while the gate controls residual flow. This AdaLN-style conditioning adapts normalization statistics based on the current pose estimate, similar to techniques used in modern diffusion models.

Can CameraCausalHead process streaming video inputs?

Yes, when causal_inference=True, the head initializes a KV-cache (lines 84-92) that preserves attention keys and values across frames. Combined with optional 3-D RoPE positional encodings from lingbot_map/layers/rope.py, this enables efficient streaming inference where the model processes video frames sequentially while maintaining temporal context through the iterative refinement loop.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →