# Speed vs Pose Accuracy Trade-offs for camera_num_iterations in LingBot-Map

> Explore the speed vs. pose accuracy trade-offs for camera_num_iterations in LingBot-Map. Optimize your system by understanding this key parameter's impact.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-27

---

**The `camera_num_iterations` parameter controls how many iterative refinement passes the CameraHead performs, trading pose estimation accuracy for faster inference and lower memory usage.**

LingBot-Map is an open-source visual localization framework that refines camera pose estimates through multiple transformer iterations. The `camera_num_iterations` configuration is the primary mechanism for balancing real-time performance requirements against the precision of camera trajectory estimation.

## How camera_num_iterations Works

The `camera_num_iterations` parameter (default: **4**) determines the number of refinement loops executed inside the `CameraHead.trunk_fn` method. During each iteration, the model performs a full transformer forward pass over the camera token to progressively refine the pose prediction.

### Implementation Location

In [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py), the iterative loop begins at line 77:

```python
for _ in range(num_iterations):
    ...

```

This loop receives its count from the model constructor in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) (lines 129 and 242):

```python
self.camera_head = CameraHead(..., num_iterations=self.camera_num_iterations)

```

The parameter is exposed to users via the command-line interface in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) at line 379:

```python
parser.add_argument("--camera_num_iterations", type=int, default=4,
                    help="Number of refinement passes in the camera head")

```

## The Accuracy vs Speed Trade-off

Changing this value creates a direct tension between three system resources: localization precision, inference latency, and GPU memory consumption.

**4 Iterations (Default)**
- **Pose Accuracy**: Highest quality estimates, particularly beneficial for challenging viewpoints, rapid motion, or noisy input sequences
- **Inference Speed**: Slowest, as each iteration adds a full transformer forward pass over the camera token
- **Memory Usage**: Largest KV-cache, storing key/value pairs for all four camera token iterations simultaneously

**1 Iteration**
- **Pose Accuracy**: Slightly lower accuracy, using only the initial prediction without refinement passes
- **Inference Speed**: Fastest, reducing inference time by approximately a factor of 4 by eliminating extra transformer passes
- **Memory Usage**: Smallest KV-cache, shrinking cache size by roughly 4× and freeing significant GPU memory

## Memory and KV-Cache Implications

The KV-cache scales linearly with the iteration count because each refinement pass adds a new set of key/value pairs for the camera token. Reducing `camera_num_iterations` from 4 to 1 shrinks the cache allocation by approximately 75%, making single-iteration mode viable for edge devices and consumer GPUs with limited VRAM.

## Practical Configuration Guidance

**Use the default (4 iterations) when:**
- Processing datasets with fast camera motion or large viewpoint changes
- GPU memory is sufficient (typically ≥ 8 GB)
- Maximum pose accuracy is prioritized over latency

**Reduce to 1 iteration when:**
- Running real-time or low-latency applications where wall-clock speed dominates
- Operating in memory-constrained environments (consumer GPUs or edge devices)
- Processing scenes with modest camera motion where initial pose estimates are close to ground truth

In most standard scenarios, single-iteration mode produces qualitatively good trajectories while dramatically reducing runtime and memory footprint.

## Configuration Examples

**Command-line usage** with single iteration:

```bash
python demo.py \
    --model_path /path/to/checkpoint.pt \
    --image_folder /path/to/images/ \
    --camera_num_iterations 1

```

**Python API** configuration:

```python
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    backend="flashinfer",          # or "sdpa"

    img_size=480,
    sliding_window=False,
    max_frame_num=1024,
    camera_num_iterations=1,      # Set to 1 for speed

)

```

**Performance benchmarking** comparison:

```python
import time
import torch

def benchmark(iterations):
    model = GCTStream(camera_num_iterations=iterations)
    dummy_imgs = torch.randn(1, 3, 480, 640).cuda()
    start = time.time()
    with torch.no_grad():
        model(dummy_imgs)  # forward pass

    return time.time() - start

print("4-iter time:", benchmark(4))
print("1-iter time:", benchmark(1))

```

## Summary

- **`camera_num_iterations`** controls the depth of iterative pose refinement in the CameraHead, defaulting to 4 passes
- **Higher values** (4) maximize pose accuracy for challenging sequences but require ~4× more compute and GPU memory
- **Lower values** (1) sacrifice marginal accuracy for near-linear speedups and significant memory savings
- The parameter is implemented in [`camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/camera_head.py) (line 77), configured in [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py) (lines 129, 242), and exposed via CLI in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (line 379)
- For real-time applications or memory-constrained devices, setting this to 1 is recommended

## Frequently Asked Questions

### What is the default value for camera_num_iterations in LingBot-Map?

The default value is **4**, as defined in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) line 379 and propagated through [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py). This setting performs three refinement passes after the initial pose prediction to maximize localization accuracy.

### How much faster is inference with camera_num_iterations set to 1?

Inference speed increases by approximately a factor of **4×** (proportional to the iteration count reduction), because each eliminated iteration removes one full transformer forward pass over the camera token.

### Does reducing camera_num_iterations affect the quality of the trajectory?

Yes, but typically only modestly. While single-iteration mode skips refinement passes that help with challenging viewpoints and rapid motion, it still produces qualitatively good trajectories for scenes with modest camera movement. The trade-off favors speed and memory efficiency over marginal gains in pose precision.

### Where is the camera_num_iterations parameter actually used in the code?

The parameter is defined in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (line 379), stored in `GCTStream` classes like [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py) (lines 129, 242), and consumed by the `CameraHead` in [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) (line 77) where it controls the refinement loop inside `trunk_fn`.