# What Benchmarking Datasets Are Used for LingBot-Map? A Complete Guide to SLAM Evaluation

> Explore the benchmarking datasets powering LingBot-Map, including KITTI, TUM RGB-D, and ETH3D. This guide details SLAM evaluation with Robbyant/lingbot-map.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: tutorial
- Published: 2026-07-27

---

**LingBot-Map supports ten major computer vision datasets—including KITTI, TUM RGB-D, ETH3D, and Seven-Scenes—through a unified `BaseDataset` interface defined in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py) that standardizes scene traversal, frame loading, and calibration handling.**

The Robbyant/lingbot-map repository provides a modular benchmarking framework designed to evaluate SLAM and visual odometry algorithms against industry-standard datasets. Each dataset is encapsulated in a Python class implementing three core methods: `get_scenes()`, `get_frame_list(scene)`, and `load_frame_data(scene, frame_id)`. This architecture allows researchers to seamlessly switch between KITTI Odometry, TUM RGB-D, and other benchmarks without modifying evaluation code.

## Core Dataset Architecture

All loaders inherit from `BaseDataset` located in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py). The abstract base class enforces a consistent contract for data retrieval:

- `get_scenes()` returns available scene identifiers
- `get_frame_list(scene)` returns ordered frame indices
- `load_frame_data(scene, frame_id)` returns a dictionary with `'rgb'` (image), `'intrinsics'` (camera parameters), and `'pose'` (ground-truth transformation when available)

The high-level entry point in [`benchmark/benchmark/core/loader.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/loader.py) instantiates specific loaders via a registry defined in [`benchmark/benchmark/core/registry.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/registry.py), mapping YAML configuration strings to concrete classes.

## Supported Benchmarking Datasets

LingBot-Map includes dedicated loaders for ten distinct datasets, each handling unique directory structures and calibration formats.

### KITTI Odometry

The `KittiDataset` class in [`benchmark/datasets/kitti.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/kitti.py) ingests the KITTI Vision Benchmark Suite. It expects the canonical layout with `poses/` containing ground-truth trajectory files ([`00.txt`](https://github.com/Robbyant/lingbot-map/blob/main/00.txt) through [`10.txt`](https://github.com/Robbyant/lingbot-map/blob/main/10.txt)) and `sequences/` with stereo imagery.

Required structure:

```text
<root>/poses/00.txt ... 10.txt
<root>/sequences/00/image_2/xxxxx.png

```

The loader reads calibration from `sequences/<seq>/calib.txt` and supports optional `target_size` resizing for patch-aligned processing.

### TUM RGB-D

`TumDataset` in [`benchmark/datasets/tum.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tum.py) handles the TUM RGB-D SLAM dataset. It automatically associates timestamps between RGB images and ground-truth poses.

Required structure:

```text
<root>/<scene_name>/rgb/
<root>/<scene_name>/rgb.txt
<root>/<scene_name>/groundtruth.txt

```

Intrinsics are selected automatically based on the Freiburg camera identifier (`freiburg1`, `freiburg2`, or `freiburg3`).

### ETH3D

The `Eth3dDataset` class in [`benchmark/datasets/eth3d.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/eth3d.py) loads the ETH3D SLAM benchmark, supporting both training and test sequences with undistorted images.

Required structure:

```text
<root>/undistorted/<scene>/images/
<root>/undistorted/<scene>/poses.txt

```

Poses follow the ETH3D format, with intrinsics read from the dataset's calibration file.

### Seven-Scenes

`SevenScenesDataset` in [`benchmark/datasets/seven_scenes.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/seven_scenes.py) manages the Microsoft 7-Scenes indoor dataset.

Required structure:

```text
<root>/7scenes/<scene>/rgb/
<root>/7scenes/<scene>/groundtruth.txt

```

Camera intrinsics are hard-coded per scene according to the official specifications.

### Tanks & Temples (TNT)

The `TntDataset` class in [`benchmark/datasets/tnt.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tnt.py) processes the Tanks & Temples benchmark using COLMAP SfM logs.

Required structure:

```text
<root>/<scene>/000001.jpg
<root>/<scene>/<scene>_COLMAP_SfM.log
<root>/<scene>/<scene>.ply
<root>/<scene>/<scene>.json
<root>/<scene>/<scene>_trans.txt

```

Per-frame camera-to-world poses are extracted from the COLMAP log, with intrinsics approximated as `fx = fy ≈ 1.2 * width`.

### Vision Benchmark in Rome (VBR)

`VbrDataset` in [`benchmark/datasets/vbr.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/vbr.py) supports the Vision Benchmark in Rome dataset with processed and aligned sequences.

Required structure:

```text
<root>/<scene>_processed_aligned/rgb/
<root>/<scene>_processed_aligned/camera_pose.txt
<root>/<scene>_processed_aligned/intrinsics.txt
<root>/processed_gt/<scene>_gt.txt

```

The loader expects a 3×3 `K` intrinsics matrix and TUM-format ground-truth poses.

### NeuralRGB-D

The `NeuralRgbdDataset` in [`benchmark/datasets/neural_rgbd.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/neural_rgbd.py) interfaces with the NeuralRGB-D benchmark data.

Required structure:

```text
<root>/data/<scene>/rgb/
<root>/data/<scene>/pose.txt

```

Poses are stored per frame in the specific format required by neural rendering evaluations.

### Oxford Spires

`OxfordSpiresDataset` in [`benchmark/datasets/oxford_spires.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/oxford_spires.py) loads the Oxford Spires dataset with a fixed scene list.

Required structure:

```text
<root>/oxford_spires/<scene>/rgb/
<root>/oxford_spires/<scene>/poses.txt

```

### Droid-W

The `DroidWDataset` class in [`benchmark/datasets/droid_w.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/droid_w.py) handles the Droid-W dataset using COLMAP-generated trajectories.

Required structure:

```text
<root>/droid_w/<scene>/rgb/
<root>/droid_w/<scene>/colmap.txt

```

Intrinsics are parsed directly from the COLMAP reconstruction file.

### General (Arbitrary Images or Video)

`GeneralDataset` in [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py) serves as a flexible fallback for custom image folders or video files. It supports PNG, JPG, BMP, and TIFF formats, with optional COLMAP reconstruction for pose estimation.

When `_use_colmap: true` is set in the YAML configuration, the loader invokes `GeneralDataset._run_colmap` to perform feature extraction, sequential matching, and mapping, caching results under `<image_dir>/colmap_workspace/`.

## Dataset Preparation Requirements

### Directory Layout Validation

LingBot-Map does not modify or relocate files; it strictly reads existing structures. You must organize downloaded data exactly as specified for each loader class. Verify your layout by running the demo script:

```bash
python demo.py --config configs/kitti.yaml

```

### Image Resizing Constraints

Many loaders accept `target_size` or `load_img_size` parameters. When resizing, intrinsics are automatically scaled to match the new dimensions. **Critical constraint**: Patch-based transformer models require both width and height to be multiples of 14; the loader raises a `ValueError` if this requirement is violated.

### COLMAP Integration

For datasets lacking ground-truth poses, enable `_use_colmap: true` in the configuration. The system requires the COLMAP binary to be available on the system `PATH` or specified via the `colmap_binary` parameter.

## YAML Configuration Examples

Configure datasets in your experiment YAML using the registry keys:

**KITTI Odometry:**

```yaml
datasets:
  kitti_odometry:
    dataset: kitti
    raw_data_root: /data/kitti_odometry
    sequences: ["00", "02"]
    target_size: [640, 480]
    _use_colmap: false

```

**Custom image folder with COLMAP reconstruction:**

```yaml
datasets:
  my_images:
    dataset: general
    raw_data_root: /data/my_image_folder
    _use_colmap: true
    load_img_size: 720

```

## Loading Data Programmatically

Iterate through any dataset using the unified interface:

```python
from benchmark.datasets.kitti import KittiDataset

dataset = KittiDataset(
    raw_data_root="/data/kitti_odometry",
    sequences=["00"],
    target_size=[640, 480],
)

for scene in dataset.get_scenes():
    for frame_id in dataset.get_frame_list(scene):
        data = dataset.load_frame_data(scene, frame_id)
        rgb = data["rgb"]           # Shape: (480, 640, 3)

        intrinsics = data["intrinsics"]  # [fx, fy, cx, cy]

        pose = data.get("pose")     # 4×4 transformation or None

```

The returned dictionary standardizes access across all ten supported LingBot-Map benchmarking datasets.

## Summary

- LingBot-Map provides dedicated loaders for **ten major SLAM datasets**: KITTI, TUM RGB-D, ETH3D, Seven-Scenes, Tanks & Temples, VBR, NeuralRGB-D, Oxford Spires, Droid-W, and general image folders.
- All loaders inherit from `BaseDataset` in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py) and implement `get_scenes()`, `get_frame_list()`, and `load_frame_data()`.
- Dataset classes are registered in [`benchmark/benchmark/core/registry.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/registry.py) and instantiated via [`benchmark/benchmark/core/loader.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/loader.py).
- Directory structures must match exact expectations; no automatic file reorganization occurs.
- Image dimensions must be multiples of 14 when using patch-based transformers.
- COLMAP integration is available for pose estimation in custom datasets through `GeneralDataset`.

## Frequently Asked Questions

### How do I add a new custom dataset to LingBot-Map?

Create a Python class inheriting from `BaseDataset` in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py), implementing the three required methods: `get_scenes()`, `get_frame_list(scene)`, and `load_frame_data(scene, frame_id)`. Register the class in [`benchmark/benchmark/core/registry.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/registry.py) with a unique string key, then reference this key in your YAML configuration under the `dataset:` field.

### Why does LingBot-Map require image dimensions to be multiples of 14?

This constraint supports patch-based transformer architectures that process images using 14×14 pixel patches. When you specify a `target_size` or `load_img_size`, the loaders in `KittiDataset` and `VbrDataset` automatically handle resizing, but you must ensure the final dimensions satisfy this requirement or the loader will raise a `ValueError` to prevent runtime errors in the vision models.

### Can I use LingBot-Map without ground-truth pose files?

Yes. Enable `_use_colmap: true` in your YAML configuration for the `GeneralDataset` loader. This triggers `GeneralDataset._run_colmap` to execute COLMAP's feature extraction, sequential matching, and mapper pipelines, generating camera poses and intrinsics from the images alone. The results are cached in `<image_dir>/colmap_workspace/` to avoid recomputation on subsequent runs.

### Where are the intrinsics stored for the TUM RGB-D dataset?

The `TumDataset` class in [`benchmark/datasets/tum.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tum.py) does not read intrinsics from a separate file. Instead, it automatically selects the appropriate camera parameters based on the Freiburg camera identifier (`freiburg1`, `freiburg2`, or `freiburg3`) detected in the scene path. These hard-coded values match the official TUM RGB-D specifications for each camera type.