How to Evaluate LingBot-Map on Standard Benchmarks Like KITTI and Oxford Spires

To evaluate LingBot-Map on KITTI or Oxford Spires, execute the three-phase benchmark pipeline: prepare raw data into BSS format with prepare.py, run streaming inference via run.py, and compute trajectory and geometry metrics with evaluate.py, configuring dataset paths and model checkpoints through YAML files.

Evaluating LingBot-Map on standard benchmarks such as KITTI and Oxford Spires requires converting raw sequences into the Benchmark Storage Structure (BSS), executing the model in an isolated conda environment, and aggregating trajectory and depth metrics. The lingbot-map repository provides a deterministic evaluation framework that automates data preparation, inference dispatch, and metric computation through modular Python scripts.

Prerequisites: Install the Benchmark Environments

The evaluation framework requires two separate conda environments to ensure isolation between the benchmark orchestration and the LingBot-Map inference backend.

  • Method environment (lingbot_map): Contains PyTorch, CUDA dependencies, evo, Open3D, and the LingBot-Map model itself.
  • Benchmark environment (bench): Runs the framework scripts that dispatch jobs and aggregate results.

Install both using the provided shell scripts:


# Install the LingBot-Map inference environment

bash envs/install_lingbot_map.sh

# Install the benchmark orchestration environment (optional but recommended)

bash envs/install_bench.sh

These scripts create the lingbot_map and bench environments according to the specifications in envs/install_lingbot_map.sh and envs/install_bench.sh.

Step 1: Configure Dataset and Model Paths

The framework uses YAML configuration files to locate raw datasets and model checkpoints. Edit the following files before running the pipeline:

  1. Dataset configuration: Update configs/kitti.yaml or configs/oxford.yaml to set raw_data_root to the local folder containing the raw KITTI odometry or Oxford Spires sequences.
  2. Method configuration: Update configs/methods/lingbot_map.yaml to set _checkpoint to the path of your downloaded lingbot-map.pt weights.

Additional parameters such as image_size, area_budget for down-sampling, and key-frame intervals can be adjusted in these YAML files to control memory usage and inference behavior.

Step 2: Prepare the Benchmark Storage Structure (BSS)

The preparation phase parses raw images, depth maps (if available), and poses, converting them into the BSS format required by the evaluation core. This step writes ground-truth data to workspace/<dataset>/<scene>/gt/.

Run the preparation script for your target benchmark:

python prepare.py --config configs/kitti.yaml

For Oxford Spires, use the corresponding configuration:

python prepare.py --config configs/oxford.yaml

The script handles optional image down-sampling based on the area_budget parameter specified in the configuration, reducing storage and memory footprint for high-resolution sequences.

Step 3: Execute Streaming Inference

The inference dispatcher (benchmark/run.py) spawns a subprocess in the lingbot_map conda environment to execute run_worker.py, which loads the LingbotMapMethod class from benchmark/methods/lingbot_map.py and processes the dataset in a streaming fashion.

To run LingBot-Map in default streaming mode:

python run.py --config configs/kitti.yaml

To run in windowed mode instead:

python run.py --config configs/oxford.yaml --method lingbot_map --mode windowed

During execution, the method wrapper initializes the upstream model by importing lingbot_map.models.gct_stream.GCTStream (streaming) or gct_stream_window.GCTStream (windowed), loads the checkpoint via torch.load(self.checkpoint, map_location=self.device), and processes images through the _prepare_images method. The wrapper automatically resolves key-frame intervals using _resolve_keyframe_interval (lines 24-38) and decodes pose encodings into camera matrices via pose_encoding_to_extri_intri.

Predictions including RGB, depth, pose, intrinsics, and optional confidence maps are written to workspace/<dataset>/<scene>/<method>/.

Step 4: Evaluate Trajectory and Geometry Metrics

The evaluation phase computes per-scene metrics including Absolute Trajectory Error (ATE), Relative Pose Error (RPE), Area Under the Curve (AUC) for relative poses, depth errors, and point-cloud Chamfer distances.

Run the evaluation script:

python evaluate.py --config configs/kitti.yaml

The benchmark/evaluate.py script uses the Evaluator class (defined in benchmark/core/evaluator.py) to process each scene independently, saving results as JSON files (eval/traj.json, eval/auc.json, etc.) in the workspace. The aggregate_eval function (lines 25-88) then merges per-scene results into dataset-level summaries such as auc_micro.json, auc_macro.json, and aggregated trajectory reports.

Optional: Generate Reports and Visualize Results

After evaluation, generate a printable summary table:

python report.py --workspace /path/to/workspace

For interactive inspection of trajectories and point clouds, launch the browser-based 3-D viewer built on the Viser backend:

python viewer.py /path/to/workspace

The viewer displays ground-truth and predicted trajectories, point clouds, and per-frame RGB/depth overlays for qualitative analysis.

Architectural Deep Dive

Method Wrapper (benchmark/methods/lingbot_map.py)

The LingbotMapMethod class serves as the adapter between the upstream LingBot-Map model and the benchmark framework. It handles:

  • Model initialization: Dynamically imports streaming or windowed model classes based on the configuration mode.
  • Checkpoint loading: Uses torch.load to restore model weights to the specified device.
  • Preprocessing: The _prepare_images method converts lists of uint8 RGB arrays into [S,3,H,W] tensors.
  • Inference dispatch: Calls model.inference_streaming or model.inference_windowed depending on the selected mode.
  • Output formatting: Converts raw network outputs into the benchmark schema using pose_encoding_to_extri_intri and returns dictionaries with keys rgb, depth, pose, intrinsics, and optional confidence.

Evaluation Engine (benchmark/evaluate.py)

This script orchestrates the metric computation pipeline. It loads configurations via ConfigManager, instantiates dataset and method objects, and iterates through scenes using the Evaluator class. The engine computes trajectory metrics (ATE/RPE), geometric consistency (AUC), depth accuracy, and point-cloud distances. Results are aggregated across scenes using the aggregate_eval function to produce statistically robust dataset-level reports.

Benchmark Storage Structure (BSS) Layout

The BSS defines a strict directory hierarchy to ensure reproducibility:

  • workspace/<dataset>/<scene>/gt/: Contains ground-truth poses, images, and depth maps.
  • workspace/<dataset>/<scene>/<method>/: Stores inference outputs from LingBot-Map.
  • workspace/<dataset>/<scene>/<method>/eval/: Holds per-scene JSON metric files.
  • Dataset-level summaries reside in the workspace root under eval/.

This layout is documented in benchmark/README.md under the BSS Storage System section.

Debug Mode for Rapid Iteration

To evaluate a single scene without processing the full dataset, append the --debug flag to both the run and evaluate commands:

python run.py --config configs/oxford.yaml --debug
python evaluate.py --config configs/oxford.yaml --debug

When --debug is detected, the framework truncates the scene list to the first entry (see lines 84-92 in run.py and evaluate.py), enabling rapid validation of configuration changes or model modifications.

Summary

Frequently Asked Questions

What metrics does the LingBot-Map benchmark compute?

The benchmark computes Absolute Trajectory Error (ATE), Relative Pose Error (RPE), Area Under the Curve (AUC) for relative pose accuracy, depth estimation errors, and point-cloud Chamfer distances. These are calculated per-scene and aggregated across the dataset using the aggregate_eval function in benchmark/evaluate.py.

Can I run LingBot-Map in windowed mode instead of streaming?

Yes. Append --mode windowed to the run.py command. The LingbotMapMethod wrapper will import gct_stream_window.GCTStream instead of the streaming variant and execute model.inference_windowed rather than model.inference_streaming, allowing evaluation of temporal window-based reconstruction.

How do I evaluate only a single scene for debugging?

Add the --debug flag to both run.py and evaluate.py. This truncates the scene list to the first entry (implemented around lines 84-92 in both scripts), enabling rapid iteration without processing the full KITTI or Oxford Spires dataset.

Where are the evaluation results stored?

Per-scene metrics are stored as JSON files in workspace/<dataset>/<scene>/<method>/eval/ (e.g., traj.json, auc.json). Dataset-level aggregates such as auc_micro.json and traj.json are saved in the workspace root under eval/, providing macro and micro averages across all sequences.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →