Speed vs Pose Accuracy Trade-offs for camera_num_iterations in LingBot-Map

The camera_num_iterations parameter controls how many iterative refinement passes the CameraHead performs, trading pose estimation accuracy for faster inference and lower memory usage.

LingBot-Map is an open-source visual localization framework that refines camera pose estimates through multiple transformer iterations. The camera_num_iterations configuration is the primary mechanism for balancing real-time performance requirements against the precision of camera trajectory estimation.

How camera_num_iterations Works

The camera_num_iterations parameter (default: 4) determines the number of refinement loops executed inside the CameraHead.trunk_fn method. During each iteration, the model performs a full transformer forward pass over the camera token to progressively refine the pose prediction.

Implementation Location

In lingbot_map/heads/camera_head.py, the iterative loop begins at line 77:

for _ in range(num_iterations):
    ...

This loop receives its count from the model constructor in lingbot_map/models/gct_stream.py (lines 129 and 242):

self.camera_head = CameraHead(..., num_iterations=self.camera_num_iterations)

The parameter is exposed to users via the command-line interface in demo.py at line 379:

parser.add_argument("--camera_num_iterations", type=int, default=4,
                    help="Number of refinement passes in the camera head")

The Accuracy vs Speed Trade-off

Changing this value creates a direct tension between three system resources: localization precision, inference latency, and GPU memory consumption.

4 Iterations (Default)

  • Pose Accuracy: Highest quality estimates, particularly beneficial for challenging viewpoints, rapid motion, or noisy input sequences
  • Inference Speed: Slowest, as each iteration adds a full transformer forward pass over the camera token
  • Memory Usage: Largest KV-cache, storing key/value pairs for all four camera token iterations simultaneously

1 Iteration

  • Pose Accuracy: Slightly lower accuracy, using only the initial prediction without refinement passes
  • Inference Speed: Fastest, reducing inference time by approximately a factor of 4 by eliminating extra transformer passes
  • Memory Usage: Smallest KV-cache, shrinking cache size by roughly 4× and freeing significant GPU memory

Memory and KV-Cache Implications

The KV-cache scales linearly with the iteration count because each refinement pass adds a new set of key/value pairs for the camera token. Reducing camera_num_iterations from 4 to 1 shrinks the cache allocation by approximately 75%, making single-iteration mode viable for edge devices and consumer GPUs with limited VRAM.

Practical Configuration Guidance

Use the default (4 iterations) when:

  • Processing datasets with fast camera motion or large viewpoint changes
  • GPU memory is sufficient (typically ≥ 8 GB)
  • Maximum pose accuracy is prioritized over latency

Reduce to 1 iteration when:

  • Running real-time or low-latency applications where wall-clock speed dominates
  • Operating in memory-constrained environments (consumer GPUs or edge devices)
  • Processing scenes with modest camera motion where initial pose estimates are close to ground truth

In most standard scenarios, single-iteration mode produces qualitatively good trajectories while dramatically reducing runtime and memory footprint.

Configuration Examples

Command-line usage with single iteration:

python demo.py \
    --model_path /path/to/checkpoint.pt \
    --image_folder /path/to/images/ \
    --camera_num_iterations 1

Python API configuration:

from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    backend="flashinfer",          # or "sdpa"

    img_size=480,
    sliding_window=False,
    max_frame_num=1024,
    camera_num_iterations=1,      # Set to 1 for speed

)

Performance benchmarking comparison:

import time
import torch

def benchmark(iterations):
    model = GCTStream(camera_num_iterations=iterations)
    dummy_imgs = torch.randn(1, 3, 480, 640).cuda()
    start = time.time()
    with torch.no_grad():
        model(dummy_imgs)  # forward pass

    return time.time() - start

print("4-iter time:", benchmark(4))
print("1-iter time:", benchmark(1))

Summary

  • camera_num_iterations controls the depth of iterative pose refinement in the CameraHead, defaulting to 4 passes
  • Higher values (4) maximize pose accuracy for challenging sequences but require ~4× more compute and GPU memory
  • Lower values (1) sacrifice marginal accuracy for near-linear speedups and significant memory savings
  • The parameter is implemented in camera_head.py (line 77), configured in gct_stream.py (lines 129, 242), and exposed via CLI in demo.py (line 379)
  • For real-time applications or memory-constrained devices, setting this to 1 is recommended

Frequently Asked Questions

What is the default value for camera_num_iterations in LingBot-Map?

The default value is 4, as defined in demo.py line 379 and propagated through gct_stream.py. This setting performs three refinement passes after the initial pose prediction to maximize localization accuracy.

How much faster is inference with camera_num_iterations set to 1?

Inference speed increases by approximately a factor of 4× (proportional to the iteration count reduction), because each eliminated iteration removes one full transformer forward pass over the camera token.

Does reducing camera_num_iterations affect the quality of the trajectory?

Yes, but typically only modestly. While single-iteration mode skips refinement passes that help with challenging viewpoints and rapid motion, it still produces qualitatively good trajectories for scenes with modest camera movement. The trade-off favors speed and memory efficiency over marginal gains in pose precision.

Where is the camera_num_iterations parameter actually used in the code?

The parameter is defined in demo.py (line 379), stored in GCTStream classes like gct_stream.py (lines 129, 242), and consumed by the CameraHead in lingbot_map/heads/camera_head.py (line 77) where it controls the refinement loop inside trunk_fn.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →