Datasets Supported by the Lingbot-Map Benchmark Evaluation Pipeline: Complete Technical Reference

The lingbot-map benchmark evaluation pipeline supports 12 distinct dataset loaders—including KITTI Odometry, TUM RGB-D, ETH3D, Tanks & Temples, Seven Scenes, and generic image folders—via a unified interface defined in benchmark/dataset/base.py that standardizes scene enumeration, frame listing, and data retrieval across all sources.

The lingbot-map repository provides a modular evaluation framework for SLAM and visual odometry systems. Its benchmark evaluation pipeline can ingest diverse real-world and synthetic datasets through specialized Python classes that inherit from a common abstract base. This architecture enables seamless switching between academic benchmarks and custom data sources without modifying downstream evaluation code.

Core Dataset Architecture

All dataset loaders derive from BaseDataset defined in benchmark/dataset/base.py. Every concrete implementation must override three core methods to satisfy the pipeline contract:

  • get_scenes() – Returns a list of scene identifiers available in the dataset.
  • get_frame_list(scene) – Returns an ordered list of frame indices for a given scene.
  • load_frame_data(scene, frame_id) – Returns a dictionary containing at least an RGB image ('rgb'), and when available, camera intrinsics ('intrinsics') and pose ('pose').

The benchmark resolves dataset types through a registry pattern implemented in benchmark/benchmark/core/registry.py, which maps YAML configuration strings to concrete loader classes. The high-level entry point in benchmark/benchmark/core/loader.py instantiates these classes and orchestrates the iteration loop:

for scene in dataset.get_scenes():
    frame_ids = dataset.get_frame_list(scene)
    for fid in frame_ids:
        frame = dataset.load_frame_data(scene, fid)
        # frame['rgb'], frame['intrinsics'], frame.get('pose')

Supported Datasets and Directory Layouts

The lingbot-map benchmark ships with native support for 10+ SLAM datasets, each encapsulated in a dedicated loader class under the benchmark/datasets/ directory.

General Image Folders and Video (GeneralDataset)

Loader: benchmark/datasets/general.py

This loader handles arbitrary image collections or video files without requiring a pre-defined calibration or pose file. For image folders, place PNG, JPG, BMP, or TIFF files directly in the root directory. For video inputs, place the .mp4 file in the root; the loader automatically extracts frames to <root>/<video_name>_frames/.

To obtain camera parameters for image-only sources, set _use_colmap: true in the YAML configuration. The GeneralDataset._run_colmap method executes COLMAP feature extraction, sequential matching, and mapping, caching results under <image_dir>/colmap_workspace/. The COLMAP binary must be available on the system PATH or specified via the colmap_binary parameter.

KITTI Odometry (KittiDataset)

Loader: benchmark/datasets/kitti.py

Expected Layout:

<root>/poses/00.txt … 10.txt
<root>/sequences/00/image_2/xxxxx.png
<root>/sequences/00/calib.txt

The loader requires ground-truth pose files in poses/ and camera calibration in sequences/<seq>/calib.txt. It supports optional target_size parameters for patch-aligned resizing, automatically rescaling intrinsics to match the resized images using the _cover_fit_center_crop logic.

VBR (Vision Benchmark in Rome) (VbrDataset)

Loader: benchmark/datasets/vbr.py

Expected Layout:

<root>/<scene>_processed_aligned/rgb/*.png
<root>/<scene>_processed_aligned/camera_pose.txt
<root>/<scene>_processed_aligned/intrinsics.txt
<root>/processed_gt/<scene>_gt.txt

Intrinsics are stored as a 3×3 K matrix in intrinsics.txt. Ground-truth poses follow the TUM format in the processed_gt/ directory.

TUM RGB-D (TumDataset)

Loader: benchmark/datasets/tum.py

Expected Layout:

<root>/<scene_name>/rgb/*.png
<root>/<scene_name>/rgb.txt
<root>/<scene_name>/groundtruth.txt

Ground-truth poses are provided in TUM format within groundtruth.txt. Intrinsics are selected automatically based on the Freiburg camera identifier (freiburg1, freiburg2, or freiburg3) parsed from the scene metadata.

Tanks & Temples (TntDataset)

Loader: benchmark/datasets/tnt.py

Expected Layout:

<root>/<scene>/000001.jpg …
<root>/<scene>/<scene>_COLMAP_SfM.log
<root>/<scene>/<scene>.ply
<root>/<scene>/<scene>.json
<root>/<scene>/<scene>_trans.txt

The loader parses COLMAP SfM logs to obtain per-frame camera-to-world poses. Intrinsics are approximated with fx = fy ≈ 1.2 * width when not explicitly provided.

ETH3D (Eth3dDataset)

Loader: benchmark/datasets/eth3d.py

Expected Layout:

<root>/undistorted/<scene>/images/*.png
<root>/undisturbed/<scene>/poses.txt

Poses follow the ETH3D format, and intrinsics are read from the dataset’s calibration files.

NeuralRGB-D (NeuralRgbdDataset)

Loader: benchmark/datasets/neural_rgbd.py

Expected Layout:

<root>/data/<scene>/rgb/*.png
<root>/data/<scene>/pose.txt

Designed specifically for the NeuralRGB-D benchmark, this loader expects per-frame poses aggregated in pose.txt.

Oxford Spires (OxfordSpiresDataset)

Loader: benchmark/datasets/oxford_spires.py

Expected Layout:

<root>/oxford_spires/<scene>/rgb/*.png
<root>/oxford_spires/<scene>/poses.txt

Similar to the General loader but with a fixed scene list and predefined directory conventions under oxford_spires/.

Droid-W (DroidWDataset)

Loader: benchmark/datasets/droid_w.py

Expected Layout:

<root>/droid_w/<scene>/rgb/*.png
<root>/droid_w/<scene>/colmap.txt

Uses COLMAP-generated trajectories where intrinsics are read directly from the colmap.txt file.

Seven Scenes (SevenScenesDataset)

Loader: benchmark/datasets/seven_scenes.py

Expected Layout:

<root>/7scenes/<scene>/rgb/*.png
<root>/7scenes/<scene>/groundtruth.txt

A classic indoor benchmark where intrinsics are hard-coded per scene within the loader implementation.

Configuration and Data Preparation

To prepare data for the lingbot-map benchmark evaluation pipeline, follow these validation steps:

  1. Download raw data from official sources (KITTI, TUM, ETH3D, etc.).
  2. Create the directory tree exactly as specified for each dataset loader. The pipeline does not reshape or move files; it only reads from expected paths.
  3. (Optional) Run COLMAP for datasets lacking calibration by setting _use_colmap: true in the YAML config.
  4. (Optional) Resize images using the target_size or load_img_size parameters. Patch-based transformer models require both dimensions to be multiples of 14; the loader raises a ValueError if this constraint is violated.
  5. Validate the layout using the demo script:
python demo.py --config configs/kitti.yaml

YAML Configuration Examples

KITTI Odometry:

datasets:
  kitti_odometry:
    dataset: kitti
    raw_data_root: /data/kitti_odometry
    sequences: ["00", "02"]
    target_size: [640, 480]
    _use_colmap: false

General Image Folder with COLMAP:

datasets:
  my_images:
    dataset: general
    raw_data_root: /data/my_image_folder
    _use_colmap: true
    load_img_size: 720

Programmatic Usage Examples

Loading a Single Frame

from benchmark.dataset.base import BaseDataset
from benchmark.datasets.kitti import KittiDataset

# Initialize the loader

kitti = KittiDataset(
    raw_data_root="/data/kitti_odometry",
    sequences=["00"],
    target_size=[640, 480],
)

scene = kitti.get_scenes()[0]      # "00"

frame_id = 10
frame = kitti.load_frame_data(scene, frame_id)

print(frame["rgb"].shape)          # (480, 640, 3)

print(frame["intrinsics"])         # [fx, fy, cx, cy]

print(frame["pose"].shape)         # (4, 4) or None

Iterating Over Scenes

from benchmark.datasets.vbr import VbrDataset

vbr = VbrDataset(
    raw_data_root="/data/vbr",
    target_size=[512, 512],
)

for scene in vbr.get_scenes():
    for fid in vbr.get_frame_list(scene):
        data = vbr.load_frame_data(scene, fid)
        # Access data["rgb"], data["pose"], data["intrinsics"]

Summary

  • The lingbot-map benchmark evaluation pipeline provides unified dataset loaders for 12+ SLAM benchmarks including KITTI, TUM RGB-D, ETH3D, Tanks & Temples, Seven Scenes, and custom image folders.
  • All loaders inherit from BaseDataset (benchmark/dataset/base.py) and implement get_scenes(), get_frame_list(), and load_frame_data().
  • Dataset-specific logic resides in benchmark/datasets/<dataset>.py, while registration and orchestration occur in benchmark/benchmark/core/registry.py and benchmark/benchmark/core/loader.py.
  • Image resizing supports patch-aligned constraints (multiples of 14), and optional COLMAP integration enables benchmarking on uncalibrated image sequences.
  • Configuration is managed through YAML files specifying raw_data_root, scene selection, and preprocessing parameters.

Frequently Asked Questions

How do I add a custom dataset to the lingbot-map benchmark?

Create a new class inheriting from BaseDataset in benchmark/datasets/, implementing the three required abstract methods (get_scenes, get_frame_list, load_frame_data). Register the class in benchmark/benchmark/core/registry.py by mapping a string key to your class constructor. The pipeline will then instantiate your loader when that key appears in the YAML configuration under the dataset: field.

Why does the benchmark require image dimensions to be multiples of 14?

Patch-based transformer vision models (such as those used in modern SLAM systems) process images by dividing them into fixed-size patches. Requiring both width and height to be multiples of 14 ensures perfect alignment between image pixels and model patch boundaries without padding artifacts. The loaders enforce this constraint and raise a ValueError if violated, though you can bypass resizing by ensuring your source data meets this requirement natively.

Can I run the benchmark on video files without extracting frames manually?

Yes. When using the GeneralDataset loader (benchmark/datasets/general.py), you can point raw_data_root to a video file (e.g., .mp4). The loader automatically extracts frames to a subdirectory named <video_name>_frames/ on first initialization. If COLMAP reconstruction is enabled via _use_colmap: true, the pipeline will process these extracted frames to generate camera poses and intrinsics.

Where are dataset poses and intrinsics cached when using COLMAP?

When _use_colmap: true is set for the GeneralDataset, the benchmark creates a colmap_workspace/ directory inside the image folder. This workspace contains the COLMAP database, sparse reconstruction files, and cached camera parameters. Subsequent runs reuse this cache to avoid reprocessing, unless the directory is deleted or the source images change.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →