How to Speed Up LingBot-Map Inference: 9 Optimization Techniques

You can speed up LingBot-Map inference by increasing the keyframe interval to reduce KV-cache writes, lowering camera-head iteration counts, enabling mixed-precision dtype (BF16/FP16) with torch.compile, and switching to the FlashInfer backend, yielding up to 30% speed improvements while maintaining pose estimation accuracy.

LingBot-Map is an open-source visual mapping system that employs a transformer-style GCT (Geometric Correspondence Transformer) model for real-time camera pose estimation and depth prediction. The repository provides several tunable parameters that directly control the speed-versus-accuracy trade-off during inference. This guide explains how to optimize LingBot-Map inference speed by adjusting keyframe caching, refinement iterations, and backend configurations in the core source files.

Understanding the Inference Architecture

LingBot-Map processes video through two primary modes: streaming (frame-by-frame) and windowed (overlapping chunks). Both rely on a shared architecture consisting of the GCTStream model, an AggregatorStream for KV-cache management, and a CameraCausalHead for iterative pose refinement.

The Streaming Pipeline

In lingbot_map/models/gct_stream_window_v2.py, the GCTStream.inference_streaming method processes frames sequentially using a KV-cache mechanism. The aggregator stores key-value pairs only at specific keyframes, while non-keyframes read from cache without writing back. This behavior is governed by the keyframe_interval parameter passed from demo.py through to the inference functions.

The Windowed Pipeline

For long sequences, GCTStream.inference_windowed splits video into overlapping windows with independent KV caches. The implementation stitches windows using a _pairwise_alignment similarity transform. This mode is essential for videos exceeding 5,000 frames, as it bounds memory usage while preserving scale consistency across window boundaries.

Key Parameters to Optimize Inference Speed

The following knobs, controlled via CLI arguments in demo.py and constructor parameters in CameraCausalHead and AggregatorStream, directly impact latency and throughput.

Adjust the Keyframe Interval

The --keyframe_interval argument in demo.py determines how often the model stores KV-cache entries. Setting this to 12 (instead of the default 1) reduces memory traffic significantly because non-keyframes skip the cache write operation via the _skip_append flag in the aggregator.

  • Speed impact: Higher values reduce GPU memory bandwidth usage.
  • Accuracy impact: Fewer keyframes reduce temporal context, causing minor pose estimation drift.
  • Location: demo.py argparse → GCTStream.inference_* methods.

Reduce Camera-Head Iterations

Inside lingbot_map/heads/camera_head.py, the CameraCausalHead.__init__ method sets self.num_iterations, which controls refinement passes in trunk_fn. Each iteration costs a full transformer block per frame.

  • Speed impact: Reducing from 4 to 1 iteration yields linear speedup per frame.
  • Accuracy impact: Each removed iteration increases pose error by approximately 0.5%.
  • CLI flag: --camera_num_iterations (passed to the constructor).

Enable Mixed-Precision and Torch-Compile

The model supports BF16 on Ampere+ GPUs and FP16 on older hardware. In demo.py, the dtype selection triggers model.aggregator.to(dtype), while the --compile flag invokes torch.compile with reduce-overhead mode and CUDA-graph warm-up via _warm_streaming.

  • Speed impact: Mixed-precision cuts tensor bandwidth for a 2× speedup; compilation adds ~5 FPS on 518×378 inputs (≈30% boost).
  • Accuracy impact: Negligible; prediction heads (_predict_*) remain in FP32 internally to preserve geometric stability.

Optimize KV-Cache Backend and Sliding Windows

In lingbot_map/aggregator/stream.py, the AggregatorStream.__init__ accepts --use_sdpa to toggle between FlashInfer (paged KV cache with custom CUDA kernels) and a pure-Python SDPA dict cache.

  • Speed impact: FlashInfer is approximately 30% faster for long sequences.
  • Accuracy impact: None; results are identical between backends.

The sliding_window_size parameter limits attention to the most recent W blocks, eviction older KV entries. Smaller windows keep compute constant but may cause depth drift in long sequences.

Practical Configuration Recipes

Use these specific flag combinations in demo.py to target different performance goals.

Maximum FPS Configuration (Real-Time Streaming)

For RTX 4090 or similar hardware targeting highest throughput:

python demo.py \
    --model_path checkpoints/gct.pt \
    --image_folder /data/seq/ \
    --mode streaming \
    --keyframe_interval 12 \
    --camera_num_iterations 1 \
    --offload_to_cpu \
    --compile

This configuration minimizes KV-cache writes, performs single-pass pose refinement, frees GPU memory after each frame, and leverages CUDA graphs for kernel launch optimization.

Balanced Speed and Accuracy

For approximately 2× speedup with less than 1% pose error increase:

python demo.py \
    --model_path checkpoints/gct.pt \
    --image_folder /data/seq/ \
    --mode streaming \
    --keyframe_interval 6 \
    --camera_num_iterations 4 \
    --no-offload_to_cpu

This retains the default four refinement iterations while halving the keyframe frequency, keeping sufficient temporal context for accurate reconstructions.

Maximum Accuracy for Offline Processing

For evaluation datasets where latency is irrelevant:

python demo.py \
    --model_path checkpoints/gct.pt \
    --image_folder /data/seq/ \
    --mode streaming \
    --keyframe_interval 1 \
    --camera_num_iterations 8 \
    --enable_3d_rope

Storing every frame and increasing iterations to 8 provides the best pose and depth estimates, with 3-D RoPE (temporal sinusoidal bias) improving continuity.

Windowed Inference for Long Videos

To prevent out-of-memory errors on videos longer than 5,000 frames:

python demo.py \
    --model_path checkpoints/gct.pt \
    --video_path movie.mp4 \
    --mode windowed \
    --window_size 64 \
    --overlap_keyframes 4 \
    --keyframe_interval 1 \
    --camera_num_iterations 4

The window_size controls keyframes per chunk, while overlap_keyframes ensures scale tokens transfer across windows via _pairwise_alignment.

Code Implementation Details

When modifying the source directly rather than using CLI flags, target these specific locations:

In lingbot_map/heads/camera_head.py, adjust iteration count:


# Inside CameraCausalHead.__init__

self.num_iterations = 2  # Default is 4

In lingbot_map/aggregator/stream.py, configure the attention backend:


# Inside AggregatorStream.__init__

self.use_sdpa = False  # False enables FlashInfer (faster)

In lingbot_map/models/gct_stream_window_v2.py, modify sliding window behavior:


# Inside GCTStream.__init__

sliding_window_size = 32  # Limit attention to recent 32 blocks

Summary

  • Increase --keyframe_interval from 1 to 6-12 to slash KV-cache memory traffic and boost FPS, at the cost of minor temporal context.
  • Reduce --camera_num_iterations in CameraCausalHead to lower per-frame transformer block computations; each removed iteration saves ~0.5% accuracy.
  • Enable --compile in demo.py to activate torch.compile with CUDA graphs, delivering ~30% speedup on modern GPUs.
  • Use FlashInfer (default) instead of --use_sdpa to leverage paged KV-cache CUDA kernels, improving long-sequence throughput by 30%.
  • Select --mode windowed with overlap parameters for videos exceeding 5,000 frames to maintain linear memory scaling without accuracy collapse.

Frequently Asked Questions

What is the fastest configuration for real-time streaming on an RTX 3090?

Set --keyframe_interval 12, --camera_num_iterations 1, --offload_to_cpu, and --compile in demo.py. This minimizes KV-cache writes to every 12th frame, uses single-pass pose refinement, and leverages CUDA graphs, achieving approximately 8 FPS on 518×378 inputs according to the benchmark suite.

How does the keyframe interval affect mapping accuracy?

Increasing --keyframe_interval reduces the frequency of KV-cache storage in AggregatorStream, which slightly degrades temporal consistency because non-keyframes cannot update the cache. Empirically, raising the interval from 1 to 6 causes less than 1% pose error increase, while setting it to 12 is suitable for real-time applications where minor drift is acceptable.

Can I run LingBot-Map without FlashInfer installed?

Yes, pass --use_sdpa to demo.py to fall back to the pure-Python SDPA backend implemented in AggregatorStream.__init__. While this uses a standard dictionary-based cache instead of FlashInfer's paged memory management, it respects the same KV-cache controls and produces identical numerical results, albeit approximately 30% slower on long sequences.

When should I use windowed mode instead of streaming?

Use --mode windowed when processing videos longer than 5,000 frames or when GPU memory is constrained. The inference_windowed method in gct_stream_window_v2.py processes overlapping chunks with independent caches, stitching them via _pairwise_alignment. This bounds memory usage to the window size rather than growing linearly with sequence length, though it requires setting --overlap_keyframes (typically 4) to maintain scale consistency across boundaries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →