LingBot-Map Benchmark Suite Datasets: Supported SLAM and Visual Odometry Loaders

LingBot-Map supports 10 distinct SLAM and visual odometry datasets including KITTI, TUM RGB-D, ETH3D, Seven Scenes, and generic image folders, each implemented as a Python class inheriting from BaseDataset with standardized get_scenes(), get_frame_list(), and load_frame_data() methods.

The LingBot-Map repository (Robbyant/lingbot-map) provides a modular benchmark infrastructure for visual SLAM and odometry research. This guide details every dataset supported in the LingBot-Map benchmark suite, including directory layout requirements, configuration parameters, and programmatic interfaces.

Dataset Loader Architecture

All dataset loaders in LingBot-Map derive from the abstract BaseDataset class defined in benchmark/dataset/base.py. Each concrete implementation must provide three core methods:

  • get_scenes() – Returns the list of scene identifiers available in the dataset.
  • get_frame_list(scene) – Returns an ordered list of frame indices for a specific scene.
  • load_frame_data(scene, frame_id) – Returns a dictionary containing at least an RGB image ('rgb') and, when available, camera intrinsics ('intrinsics') and pose ('pose').

The registry in benchmark/benchmark/core/registry.py maps YAML configuration keys to these concrete classes, while benchmark/benchmark/core/loader.py serves as the high-level entry point for instantiation.

Supported Datasets

LingBot-Map provides dedicated loaders for major academic benchmarks and generic image sources. Each requires a specific on-disk layout that the benchmark reads directly without file reorganization.

Autonomous Driving and Outdoor

KITTI Odometry (KittiDataset – benchmark/datasets/kitti.py)

The KITTI loader expects the standard odometry devkit structure:

<root>/poses/00.txt … 10.txt
<root>/sequences/00/image_2/xxxxx.png

Ground-truth poses reside in poses/, while camera calibration files are located at sequences/<seq>/calib.txt. The loader supports optional target_size resizing with automatic intrinsic rescaling for patch-aligned dimensions.

Oxford Spires (OxfordSpiresDataset – benchmark/datasets/oxford_spires.py)

Designed for architectural mapping, this loader expects:

<root>/oxford_spires/<scene>/rgb/*.png
<root>/oxford_spires/<scene>/poses.txt

The layout mirrors the General loader but enforces a fixed scene list and specific subfolder naming.

Indoor and RGB-D

TUM RGB-D (TumDataset – benchmark/datasets/tum.py)

TUM sequences require timestamp-synchronized data:

<root>/<scene_name>/rgb/*.png
<root>/<scene_name>/rgb.txt
<root>/<scene_name>/groundtruth.txt```

Intrinsics are automatically selected based on the Freiburg camera identifier (`freiburg1`, `freiburg2`, or `freiburg3`) extracted from the scene path.

**Seven Scenes** (`SevenScenesDataset` – [`benchmark/datasets/seven_scenes.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/seven_scenes.py))

This classic indoor benchmark uses:

```text
<root>/7scenes/<scene>/rgb/*.png
<root>/7scenes/<scene>/groundtruth.txt```

Intrinsics are hard-coded per scene within the loader implementation.

**NeuralRGB-D** (`NeuralRgbdDataset` – [`benchmark/datasets/neural_rgbd.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/neural_rgbd.py))

Designed for the NeuralRGB-D benchmark:

```text
<root>/data/<scene>/rgb/*.png
<root>/data/<scene>/pose.txt```

Poses are stored per frame in a specialized format used by neural rendering evaluations.

**ETH3D** (`Eth3dDataset` – [`benchmark/datasets/eth3d.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/eth3d.py))

The ETH3D loader expects undistorted imagery:

```text
<root>/undistorted/<scene>/images/*.png
<root>/undistorted/<scene>/poses.txt```

Intrinsics are read from the dataset's native calibration files.

### Cultural Heritage and Large-Scale

**Tanks & Temples** (`TntDataset` – [`benchmark/datasets/tnt.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tnt.py))

For large-scale reconstruction benchmarking:

```text
<root>/<scene>/000001.jpg …
<root>/<scene>/<scene>_COLMAP_SfM.log
<root>/<scene>/<scene>.ply
<root>/<scene>/<scene>.json
<root>/<scene>/<scene>_trans.txt```

Per-frame camera-to-world poses are parsed from the COLMAP SfM log. Intrinsics are approximated using `fx = fy ≈ 1.2 × width` when exact calibration is unavailable.

**Vision Benchmark in Rome (VBR)** (`VbrDataset` – [`benchmark/datasets/vbr.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/vbr.py))

VBR uses a processed aligned format:

```text
<root>/<scene>_processed_aligned/rgb/*.png
<root>/<scene>_processed_aligned/camera_pose.txt
<root>/<scene>_processed_aligned/intrinsics.txt
<root>/processed_gt/<scene>_gt.txt```

The loader expects intrinsics as a 3×3 `K` matrix and ground-truth poses in TUM format under `processed_gt/`.

### Specialized and Custom

**Droid-W** (`DroidWDataset` – [`benchmark/datasets/droid_w.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/droid_w.py))

This specialized loader uses COLMAP-generated trajectories:

```text
<root>/droid_w/<scene>/rgb/*.png
<root>/droid_w/<scene>/colmap.txt```

Intrinsics are extracted directly from the COLMAP reconstruction file.

**General Image Folders and Video** (`GeneralDataset` – [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py))

For custom data, the general loader accepts arbitrary image folders containing PNG, JPG, BMP, or TIFF files, or single video files:

```text
<root>/                    # Image folder mode

<root>/my_video.mp4       # Video mode (extracts to <root>/my_video_frames/)

When _use_colmap: true is set in the YAML configuration, the loader automatically runs COLMAP feature extraction, sequential matching, and mapping. Results are cached under <image_dir>/colmap_workspace/. The COLMAP binary must be available in the system PATH or specified via the colmap_binary parameter.

Data Preparation and Configuration

Preparing data for the LingBot-Map benchmark suite requires strict adherence to directory layouts. The benchmark does not reshape file hierarchies; it only reads existing structures.

Preparation Steps

  1. Download official datasets from their respective sources (KITTI, TUM, etc.).
  2. Create the exact directory tree shown in the loader specifications above.
  3. (Optional) Run COLMAP for pose-less datasets by enabling _use_colmap: true in the configuration.
  4. Configure resizing using target_size: [width, height] or load_img_size. Note that patch-based transformer models require dimensions to be multiples of 14; the loader raises a ValueError if this constraint is violated while intrinsics are automatically rescaled to match resized images (see _cover_fit_center_crop implementations in relevant loaders).

YAML Configuration Examples

KITTI Odometry:

datasets:
  kitti_odometry:
    dataset: kitti
    raw_data_root: /data/kitti_odometry
    sequences: ["00", "02"]
    target_size: [640, 480]
    _use_colmap: false

General Dataset with COLMAP Reconstruction:

datasets:
  my_images:
    dataset: general
    raw_data_root: /data/my_image_folder
    _use_colmap: true
    load_img_size: 720

Validate your configuration using the demonstration script:

python demo.py --config configs/kitti.yaml

Programmatic Usage

All datasets expose a unified Python interface regardless of underlying format.

Loading Individual Frames

from benchmark.datasets.kitti import KittiDataset

kitti = KittiDataset(
    raw_data_root="/data/kitti_odometry",
    sequences=["00"],
    target_size=[640, 480],
)

scene = kitti.get_scenes()[0]
frame_id = 10
frame = kitti.load_frame_data(scene, frame_id)

print(frame["rgb"].shape)          # (480, 640, 3)

print(frame["intrinsics"])         # [fx, fy, cx, cy]

print(frame["pose"].shape)         # (4, 4) or None

Iterating Over Scenes

from benchmark.datasets.vbr import VbrDataset

vbr = VbrDataset(
    raw_data_root="/data/vbr",
    target_size=[512, 512],
)

for scene in vbr.get_scenes():
    for fid in vbr.get_frame_list(scene):
        data = vbr.load_frame_data(scene, fid)
        # Access data["rgb"], data["pose"], data["intrinsics"]

Summary

  • LingBot-Map provides 10 specialized dataset loaders located in benchmark/datasets/, ranging from autonomous driving (KITTI) to indoor SLAM (TUM, Seven Scenes).
  • Every loader inherits from BaseDataset in benchmark/dataset/base.py and implements get_scenes(), get_frame_list(), and load_frame_data().
  • Strict directory layouts are mandatory; the benchmark reads data without file reorganization.
  • COLMAP integration via _use_colmap: true enables pose estimation for arbitrary image collections using the GeneralDataset loader.
  • Image resizing with automatic intrinsic calibration is supported, requiring dimensions to be multiples of 14 for compatible vision models.

Frequently Asked Questions

How do I add a custom dataset to the LingBot-Map benchmark suite?

Create a new class inheriting from BaseDataset in benchmark/dataset/base.py, implement the three required abstract methods (get_scenes(), get_frame_list(), load_frame_data()), and register the class in benchmark/benchmark/core/registry.py with a unique YAML key. Follow the existing implementations in benchmark/datasets/kitti.py or benchmark/datasets/tum.py as reference templates.

Does LingBot-Map automatically download dataset files?

No. You must manually download raw data from official sources (KITTI Odometry, TUM RGB-D, ETH3D, etc.) and organize them into the exact directory structures expected by each loader. The benchmark strictly reads existing files and does not perform downloads, decompression, or file reorganization.

Why must target_size dimensions be multiples of 14?

Patch-based vision transformers (such as those used in modern SLAM models integrated with LingBot-Map) typically use a patch size of 14×14 pixels. The loaders enforce this constraint through validation checks that raise a ValueError if the requested dimensions violate the requirement, ensuring proper alignment for feature extraction while automatically rescaling camera intrinsics to match the new image resolution.

Can I process video files directly without extracting frames manually?

Yes. The GeneralDataset loader accepts a video file path as raw_data_root. When configured with a video file, it automatically extracts frames to a subdirectory (e.g., <root>/my_video_frames/) on first load. You can also enable COLMAP reconstruction (_use_colmap: true) to generate poses and intrinsics for video sequences that lack calibration data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →