How to Reduce Camera Head Iterations for Faster Inference in LingBot-Map

Reduce the camera_num_iterations parameter from its default of 4 to a lower integer (e.g., 2) when instantiating the GCTStream model or via the --camera_num_iterations CLI flag to decrease computational overhead per frame.

The LingBot-Map repository implements an iterative camera pose refinement mechanism inside its camera heads. By default, the model performs four refinement steps to converge on accurate camera poses, but this loop can be shortened for latency-sensitive applications. Lowering the iteration count directly reduces the number of transformer trunk passes, accelerating inference while trading off pose precision.

Understanding the Camera Head Architecture

The camera head in LingBot-Map uses an iterative refinement loop to progressively improve pose estimates. Both the non-causal CameraHead and the streaming-optimized CameraCausalHead implement this pattern in lingbot_map/heads/camera_head.py.

How the Iterative Refinement Loop Works

Inside the trunk_fn method, the model embeds the current pose prediction, modulates the input token, and passes it through a transformer trunk to compute a delta update. This process repeats for num_iterations steps:

  • CameraHead (lines 114-144): Iterates over num_iterations in trunk_fn, applying learned deltas with activation functions after each transformer pass.
  • CameraCausalHead (lines 341-371): Implements identical logic with additional KV-cache support and 3-D RoPE for streaming scenarios.

Each iteration invokes the full transformer trunk, making this parameter a primary lever for controlling computational cost.

How to Reduce Camera Head Iterations

You can adjust the iteration count through three interfaces depending on your integration needs.

Via the Python API

When constructing a GCTStream model, pass the camera_num_iterations argument to the constructor. The value is stored in __init__ (lines 28-30 of lingbot_map/models/gct_stream.py) and forwarded to _build_camera_head (lines 38-44).

import torch
from lingbot_map.models.gct_stream import GCTStream

# Initialize model with only 2 refinement steps instead of default 4

model = GCTStream(
    img_size=518,
    embed_dim=1024,
    enable_camera=True,
    camera_num_iterations=2,  # Reduce iterations for faster inference

)

# Run streaming inference

dummy_frames = torch.randn(1, 10, 3, 518, 518)  # [B, S, C, H, W]

preds = model.inference_streaming(dummy_frames, keyframe_interval=1)
print(preds["pose_enc"].shape)  # Output: [1, 10, 9]

Via Command Line Arguments

The profiling and demo scripts expose the iteration count as a CLI flag. In gct_profile.py and demo.py (lines 233-236), the --camera_num_iterations argument allows runtime adjustment without code changes:

python demo.py --camera_num_iterations 2 --input video.mp4

Direct Head Manipulation (Advanced)

For specialized use cases, instantiate CameraHead directly and override the iteration count when calling trunk_fn:

from lingbot_map.heads.camera_head import CameraHead
import torch

head = CameraHead(
    dim_in=2048,
    trunk_depth=4,
    num_heads=16,
    mlp_ratio=4,
    init_values=0.01,
    trans_act="linear",
    quat_act="linear",
    fl_act="relu",
)

tokens = torch.randn(1, 1, 2048)

# Manually specify fewer iterations during the forward pass

poses = head.trunk_fn(tokens, num_iterations=2)
print(poses[0].shape)  # [1, 9]

Performance vs. Accuracy Trade-offs

Reducing camera_num_iterations from 4 to 2 cuts the camera head computation roughly in half, significantly improving frames-per-second throughput. However, fewer refinement steps mean less opportunity for the model to converge on precise pose estimates. Applications requiring high-precision SLAM or photogrammetry should retain the default value, while real-time preview or low-latency robotics scenarios benefit from reduced iterations.

Summary

  • Default behavior: The camera head performs 4 refinement iterations as defined in lingbot_map/models/gct_stream.py.
  • Primary control: Adjust the camera_num_iterations parameter when constructing GCTStream or via the --camera_num_iterations CLI flag.
  • Implementation location: The iteration loop resides in trunk_fn within lingbot_map/heads/camera_head.py for both CameraHead and CameraCausalHead.
  • Trade-off: Lower values increase inference speed but may degrade pose estimation accuracy.

Frequently Asked Questions

What is the default number of camera head iterations?

The default value is 4, set in the GCTStream constructor (lingbot_map/models/gct_stream.py, lines 28-30) and passed to the camera head during initialization.

How does reducing iterations affect pose estimation accuracy?

Fewer iterations mean less refinement of the pose estimate. While the model still produces valid poses, the accuracy may decrease compared to the default 4-step refinement, particularly for complex camera motions or challenging visual conditions.

Can I adjust iterations for streaming inference?

Yes. The CameraCausalHead class supports the same num_iterations parameter as the standard CameraHead. When using inference_streaming(), the iteration count set during model construction applies to all frames in the stream.

Which component controls the iteration count in the source code?

The camera_num_iterations attribute is defined in GCTStream.__init__ and wired to the camera head in _build_camera_head. The actual loop execution occurs in trunk_fn within lingbot_map/heads/camera_head.py, where the method iterates num_iterations times over the transformer trunk and delta computation logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →