# How to Speed Up LingBot-Map Inference: 9 Optimization Techniques

> Speed up LingBot-Map inference up to 30% using techniques like keyframe intervals, mixed precision, and FlashInfer. Optimize your pose estimation performance now.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-28

---

**You can speed up LingBot-Map inference by increasing the keyframe interval to reduce KV-cache writes, lowering camera-head iteration counts, enabling mixed-precision dtype (BF16/FP16) with torch.compile, and switching to the FlashInfer backend, yielding up to 30% speed improvements while maintaining pose estimation accuracy.**

LingBot-Map is an open-source visual mapping system that employs a transformer-style GCT (Geometric Correspondence Transformer) model for real-time camera pose estimation and depth prediction. The repository provides several tunable parameters that directly control the speed-versus-accuracy trade-off during inference. This guide explains how to optimize LingBot-Map inference speed by adjusting keyframe caching, refinement iterations, and backend configurations in the core source files.

## Understanding the Inference Architecture

LingBot-Map processes video through two primary modes: **streaming** (frame-by-frame) and **windowed** (overlapping chunks). Both rely on a shared architecture consisting of the `GCTStream` model, an `AggregatorStream` for KV-cache management, and a `CameraCausalHead` for iterative pose refinement.

### The Streaming Pipeline

In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), the `GCTStream.inference_streaming` method processes frames sequentially using a KV-cache mechanism. The aggregator stores key-value pairs only at specific **keyframes**, while non-keyframes read from cache without writing back. This behavior is governed by the `keyframe_interval` parameter passed from [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) through to the inference functions.

### The Windowed Pipeline

For long sequences, `GCTStream.inference_windowed` splits video into overlapping windows with independent KV caches. The implementation stitches windows using a `_pairwise_alignment` similarity transform. This mode is essential for videos exceeding 5,000 frames, as it bounds memory usage while preserving scale consistency across window boundaries.

## Key Parameters to Optimize Inference Speed

The following knobs, controlled via CLI arguments in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and constructor parameters in `CameraCausalHead` and `AggregatorStream`, directly impact latency and throughput.

### Adjust the Keyframe Interval

The `--keyframe_interval` argument in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) determines how often the model stores KV-cache entries. Setting this to 12 (instead of the default 1) reduces memory traffic significantly because non-keyframes skip the cache write operation via the `_skip_append` flag in the aggregator.

- **Speed impact**: Higher values reduce GPU memory bandwidth usage.
- **Accuracy impact**: Fewer keyframes reduce temporal context, causing minor pose estimation drift.
- **Location**: [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) argparse → `GCTStream.inference_*` methods.

### Reduce Camera-Head Iterations

Inside [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py), the `CameraCausalHead.__init__` method sets `self.num_iterations`, which controls refinement passes in `trunk_fn`. Each iteration costs a full transformer block per frame.

- **Speed impact**: Reducing from 4 to 1 iteration yields linear speedup per frame.
- **Accuracy impact**: Each removed iteration increases pose error by approximately 0.5%.
- **CLI flag**: `--camera_num_iterations` (passed to the constructor).

### Enable Mixed-Precision and Torch-Compile

The model supports BF16 on Ampere+ GPUs and FP16 on older hardware. In [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py), the `dtype` selection triggers `model.aggregator.to(dtype)`, while the `--compile` flag invokes `torch.compile` with `reduce-overhead` mode and CUDA-graph warm-up via `_warm_streaming`.

- **Speed impact**: Mixed-precision cuts tensor bandwidth for a 2× speedup; compilation adds ~5 FPS on 518×378 inputs (≈30% boost).
- **Accuracy impact**: Negligible; prediction heads (`_predict_*`) remain in FP32 internally to preserve geometric stability.

### Optimize KV-Cache Backend and Sliding Windows

In [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), the `AggregatorStream.__init__` accepts `--use_sdpa` to toggle between FlashInfer (paged KV cache with custom CUDA kernels) and a pure-Python SDPA dict cache.

- **Speed impact**: FlashInfer is approximately 30% faster for long sequences.
- **Accuracy impact**: None; results are identical between backends.

The `sliding_window_size` parameter limits attention to the most recent *W* blocks, eviction older KV entries. Smaller windows keep compute constant but may cause depth drift in long sequences.

## Practical Configuration Recipes

Use these specific flag combinations in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) to target different performance goals.

### Maximum FPS Configuration (Real-Time Streaming)

For RTX 4090 or similar hardware targeting highest throughput:

```bash
python demo.py \
    --model_path checkpoints/gct.pt \
    --image_folder /data/seq/ \
    --mode streaming \
    --keyframe_interval 12 \
    --camera_num_iterations 1 \
    --offload_to_cpu \
    --compile

```

This configuration minimizes KV-cache writes, performs single-pass pose refinement, frees GPU memory after each frame, and leverages CUDA graphs for kernel launch optimization.

### Balanced Speed and Accuracy

For approximately 2× speedup with less than 1% pose error increase:

```bash
python demo.py \
    --model_path checkpoints/gct.pt \
    --image_folder /data/seq/ \
    --mode streaming \
    --keyframe_interval 6 \
    --camera_num_iterations 4 \
    --no-offload_to_cpu

```

This retains the default four refinement iterations while halving the keyframe frequency, keeping sufficient temporal context for accurate reconstructions.

### Maximum Accuracy for Offline Processing

For evaluation datasets where latency is irrelevant:

```bash
python demo.py \
    --model_path checkpoints/gct.pt \
    --image_folder /data/seq/ \
    --mode streaming \
    --keyframe_interval 1 \
    --camera_num_iterations 8 \
    --enable_3d_rope

```

Storing every frame and increasing iterations to 8 provides the best pose and depth estimates, with 3-D RoPE (temporal sinusoidal bias) improving continuity.

### Windowed Inference for Long Videos

To prevent out-of-memory errors on videos longer than 5,000 frames:

```bash
python demo.py \
    --model_path checkpoints/gct.pt \
    --video_path movie.mp4 \
    --mode windowed \
    --window_size 64 \
    --overlap_keyframes 4 \
    --keyframe_interval 1 \
    --camera_num_iterations 4

```

The `window_size` controls keyframes per chunk, while `overlap_keyframes` ensures scale tokens transfer across windows via `_pairwise_alignment`.

## Code Implementation Details

When modifying the source directly rather than using CLI flags, target these specific locations:

In [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py), adjust iteration count:

```python

# Inside CameraCausalHead.__init__

self.num_iterations = 2  # Default is 4

```

In [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), configure the attention backend:

```python

# Inside AggregatorStream.__init__

self.use_sdpa = False  # False enables FlashInfer (faster)

```

In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), modify sliding window behavior:

```python

# Inside GCTStream.__init__

sliding_window_size = 32  # Limit attention to recent 32 blocks

```

## Summary

- **Increase `--keyframe_interval`** from 1 to 6-12 to slash KV-cache memory traffic and boost FPS, at the cost of minor temporal context.
- **Reduce `--camera_num_iterations`** in `CameraCausalHead` to lower per-frame transformer block computations; each removed iteration saves ~0.5% accuracy.
- **Enable `--compile`** in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) to activate torch.compile with CUDA graphs, delivering ~30% speedup on modern GPUs.
- **Use FlashInfer** (default) instead of `--use_sdpa` to leverage paged KV-cache CUDA kernels, improving long-sequence throughput by 30%.
- **Select `--mode windowed`** with overlap parameters for videos exceeding 5,000 frames to maintain linear memory scaling without accuracy collapse.

## Frequently Asked Questions

### What is the fastest configuration for real-time streaming on an RTX 3090?

Set `--keyframe_interval 12`, `--camera_num_iterations 1`, `--offload_to_cpu`, and `--compile` in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py). This minimizes KV-cache writes to every 12th frame, uses single-pass pose refinement, and leverages CUDA graphs, achieving approximately 8 FPS on 518×378 inputs according to the benchmark suite.

### How does the keyframe interval affect mapping accuracy?

Increasing `--keyframe_interval` reduces the frequency of KV-cache storage in `AggregatorStream`, which slightly degrades temporal consistency because non-keyframes cannot update the cache. Empirically, raising the interval from 1 to 6 causes less than 1% pose error increase, while setting it to 12 is suitable for real-time applications where minor drift is acceptable.

### Can I run LingBot-Map without FlashInfer installed?

Yes, pass `--use_sdpa` to [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) to fall back to the pure-Python SDPA backend implemented in `AggregatorStream.__init__`. While this uses a standard dictionary-based cache instead of FlashInfer's paged memory management, it respects the same KV-cache controls and produces identical numerical results, albeit approximately 30% slower on long sequences.

### When should I use windowed mode instead of streaming?

Use `--mode windowed` when processing videos longer than 5,000 frames or when GPU memory is constrained. The `inference_windowed` method in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py) processes overlapping chunks with independent caches, stitching them via `_pairwise_alignment`. This bounds memory usage to the window size rather than growing linearly with sequence length, though it requires setting `--overlap_keyframes` (typically 4) to maintain scale consistency across boundaries.