# Datasets Supported by the Lingbot-Map Benchmark Evaluation Pipeline: Complete Technical Reference

> Explore the 12 datasets supported by the Lingbot-Map benchmark evaluation pipeline. Discover loaders for KITTI, TUM RGB-D, ETH3D, and more, all managed by a unified interface.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: api-reference
- Published: 2026-07-22

---

**The lingbot-map benchmark evaluation pipeline supports 12 distinct dataset loaders—including KITTI Odometry, TUM RGB-D, ETH3D, Tanks & Temples, Seven Scenes, and generic image folders—via a unified interface defined in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py) that standardizes scene enumeration, frame listing, and data retrieval across all sources.**

The lingbot-map repository provides a modular evaluation framework for SLAM and visual odometry systems. Its **benchmark evaluation pipeline** can ingest diverse real-world and synthetic datasets through specialized Python classes that inherit from a common abstract base. This architecture enables seamless switching between academic benchmarks and custom data sources without modifying downstream evaluation code.

## Core Dataset Architecture

All dataset loaders derive from `BaseDataset` defined in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py). Every concrete implementation must override three core methods to satisfy the pipeline contract:

- `get_scenes()` – Returns a list of scene identifiers available in the dataset.
- `get_frame_list(scene)` – Returns an ordered list of frame indices for a given scene.
- `load_frame_data(scene, frame_id)` – Returns a dictionary containing at least an RGB image (`'rgb'`), and when available, camera intrinsics (`'intrinsics'`) and pose (`'pose'`).

The benchmark resolves dataset types through a registry pattern implemented in [`benchmark/benchmark/core/registry.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/registry.py), which maps YAML configuration strings to concrete loader classes. The high-level entry point in [`benchmark/benchmark/core/loader.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/loader.py) instantiates these classes and orchestrates the iteration loop:

```python
for scene in dataset.get_scenes():
    frame_ids = dataset.get_frame_list(scene)
    for fid in frame_ids:
        frame = dataset.load_frame_data(scene, fid)
        # frame['rgb'], frame['intrinsics'], frame.get('pose')

```

## Supported Datasets and Directory Layouts

The lingbot-map benchmark ships with native support for 10+ SLAM datasets, each encapsulated in a dedicated loader class under the `benchmark/datasets/` directory.

### General Image Folders and Video (GeneralDataset)

**Loader:** [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py)

This loader handles arbitrary image collections or video files without requiring a pre-defined calibration or pose file. For image folders, place PNG, JPG, BMP, or TIFF files directly in the root directory. For video inputs, place the `.mp4` file in the root; the loader automatically extracts frames to `<root>/<video_name>_frames/`.

To obtain camera parameters for image-only sources, set `_use_colmap: true` in the YAML configuration. The `GeneralDataset._run_colmap` method executes COLMAP feature extraction, sequential matching, and mapping, caching results under `<image_dir>/colmap_workspace/`. The COLMAP binary must be available on the system `PATH` or specified via the `colmap_binary` parameter.

### KITTI Odometry (KittiDataset)

**Loader:** [`benchmark/datasets/kitti.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/kitti.py)

**Expected Layout:**

```text
<root>/poses/00.txt … 10.txt
<root>/sequences/00/image_2/xxxxx.png
<root>/sequences/00/calib.txt

```

The loader requires ground-truth pose files in `poses/` and camera calibration in `sequences/<seq>/calib.txt`. It supports optional `target_size` parameters for patch-aligned resizing, automatically rescaling intrinsics to match the resized images using the `_cover_fit_center_crop` logic.

### VBR (Vision Benchmark in Rome) (VbrDataset)

**Loader:** [`benchmark/datasets/vbr.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/vbr.py)

**Expected Layout:**

```text
<root>/<scene>_processed_aligned/rgb/*.png
<root>/<scene>_processed_aligned/camera_pose.txt
<root>/<scene>_processed_aligned/intrinsics.txt
<root>/processed_gt/<scene>_gt.txt

```

Intrinsics are stored as a 3×3 `K` matrix in [`intrinsics.txt`](https://github.com/Robbyant/lingbot-map/blob/main/intrinsics.txt). Ground-truth poses follow the TUM format in the `processed_gt/` directory.

### TUM RGB-D (TumDataset)

**Loader:** [`benchmark/datasets/tum.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tum.py)

**Expected Layout:**

```text
<root>/<scene_name>/rgb/*.png
<root>/<scene_name>/rgb.txt
<root>/<scene_name>/groundtruth.txt

```

Ground-truth poses are provided in TUM format within [`groundtruth.txt`](https://github.com/Robbyant/lingbot-map/blob/main/groundtruth.txt). Intrinsics are selected automatically based on the Freiburg camera identifier (freiburg1, freiburg2, or freiburg3) parsed from the scene metadata.

### Tanks & Temples (TntDataset)

**Loader:** [`benchmark/datasets/tnt.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tnt.py)

**Expected Layout:**

```text
<root>/<scene>/000001.jpg …
<root>/<scene>/<scene>_COLMAP_SfM.log
<root>/<scene>/<scene>.ply
<root>/<scene>/<scene>.json
<root>/<scene>/<scene>_trans.txt

```

The loader parses COLMAP SfM logs to obtain per-frame camera-to-world poses. Intrinsics are approximated with `fx = fy ≈ 1.2 * width` when not explicitly provided.

### ETH3D (Eth3dDataset)

**Loader:** [`benchmark/datasets/eth3d.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/eth3d.py)

**Expected Layout:**

```text
<root>/undistorted/<scene>/images/*.png
<root>/undisturbed/<scene>/poses.txt

```

Poses follow the ETH3D format, and intrinsics are read from the dataset’s calibration files.

### NeuralRGB-D (NeuralRgbdDataset)

**Loader:** [`benchmark/datasets/neural_rgbd.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/neural_rgbd.py)

**Expected Layout:**

```text
<root>/data/<scene>/rgb/*.png
<root>/data/<scene>/pose.txt

```

Designed specifically for the NeuralRGB-D benchmark, this loader expects per-frame poses aggregated in [`pose.txt`](https://github.com/Robbyant/lingbot-map/blob/main/pose.txt).

### Oxford Spires (OxfordSpiresDataset)

**Loader:** [`benchmark/datasets/oxford_spires.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/oxford_spires.py)

**Expected Layout:**

```text
<root>/oxford_spires/<scene>/rgb/*.png
<root>/oxford_spires/<scene>/poses.txt

```

Similar to the General loader but with a fixed scene list and predefined directory conventions under `oxford_spires/`.

### Droid-W (DroidWDataset)

**Loader:** [`benchmark/datasets/droid_w.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/droid_w.py)

**Expected Layout:**

```text
<root>/droid_w/<scene>/rgb/*.png
<root>/droid_w/<scene>/colmap.txt

```

Uses COLMAP-generated trajectories where intrinsics are read directly from the [`colmap.txt`](https://github.com/Robbyant/lingbot-map/blob/main/colmap.txt) file.

### Seven Scenes (SevenScenesDataset)

**Loader:** [`benchmark/datasets/seven_scenes.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/seven_scenes.py)

**Expected Layout:**

```text
<root>/7scenes/<scene>/rgb/*.png
<root>/7scenes/<scene>/groundtruth.txt

```

A classic indoor benchmark where intrinsics are hard-coded per scene within the loader implementation.

## Configuration and Data Preparation

To prepare data for the **lingbot-map benchmark evaluation pipeline**, follow these validation steps:

1. **Download** raw data from official sources (KITTI, TUM, ETH3D, etc.).
2. **Create the directory tree** exactly as specified for each dataset loader. The pipeline does not reshape or move files; it only reads from expected paths.
3. **(Optional) Run COLMAP** for datasets lacking calibration by setting `_use_colmap: true` in the YAML config.
4. **(Optional) Resize images** using the `target_size` or `load_img_size` parameters. Patch-based transformer models require both dimensions to be multiples of 14; the loader raises a `ValueError` if this constraint is violated.
5. **Validate** the layout using the demo script:

```bash
python demo.py --config configs/kitti.yaml

```

### YAML Configuration Examples

**KITTI Odometry:**

```yaml
datasets:
  kitti_odometry:
    dataset: kitti
    raw_data_root: /data/kitti_odometry
    sequences: ["00", "02"]
    target_size: [640, 480]
    _use_colmap: false

```

**General Image Folder with COLMAP:**

```yaml
datasets:
  my_images:
    dataset: general
    raw_data_root: /data/my_image_folder
    _use_colmap: true
    load_img_size: 720

```

## Programmatic Usage Examples

### Loading a Single Frame

```python
from benchmark.dataset.base import BaseDataset
from benchmark.datasets.kitti import KittiDataset

# Initialize the loader

kitti = KittiDataset(
    raw_data_root="/data/kitti_odometry",
    sequences=["00"],
    target_size=[640, 480],
)

scene = kitti.get_scenes()[0]      # "00"

frame_id = 10
frame = kitti.load_frame_data(scene, frame_id)

print(frame["rgb"].shape)          # (480, 640, 3)

print(frame["intrinsics"])         # [fx, fy, cx, cy]

print(frame["pose"].shape)         # (4, 4) or None

```

### Iterating Over Scenes

```python
from benchmark.datasets.vbr import VbrDataset

vbr = VbrDataset(
    raw_data_root="/data/vbr",
    target_size=[512, 512],
)

for scene in vbr.get_scenes():
    for fid in vbr.get_frame_list(scene):
        data = vbr.load_frame_data(scene, fid)
        # Access data["rgb"], data["pose"], data["intrinsics"]

```

## Summary

- The **lingbot-map benchmark evaluation pipeline** provides unified dataset loaders for 12+ SLAM benchmarks including KITTI, TUM RGB-D, ETH3D, Tanks & Temples, Seven Scenes, and custom image folders.
- All loaders inherit from `BaseDataset` ([`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py)) and implement `get_scenes()`, `get_frame_list()`, and `load_frame_data()`.
- Dataset-specific logic resides in `benchmark/datasets/<dataset>.py`, while registration and orchestration occur in [`benchmark/benchmark/core/registry.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/registry.py) and [`benchmark/benchmark/core/loader.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/loader.py).
- Image resizing supports patch-aligned constraints (multiples of 14), and optional COLMAP integration enables benchmarking on uncalibrated image sequences.
- Configuration is managed through YAML files specifying `raw_data_root`, scene selection, and preprocessing parameters.

## Frequently Asked Questions

### How do I add a custom dataset to the lingbot-map benchmark?

Create a new class inheriting from `BaseDataset` in `benchmark/datasets/`, implementing the three required abstract methods (`get_scenes`, `get_frame_list`, `load_frame_data`). Register the class in [`benchmark/benchmark/core/registry.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/registry.py) by mapping a string key to your class constructor. The pipeline will then instantiate your loader when that key appears in the YAML configuration under the `dataset:` field.

### Why does the benchmark require image dimensions to be multiples of 14?

Patch-based transformer vision models (such as those used in modern SLAM systems) process images by dividing them into fixed-size patches. Requiring both width and height to be multiples of 14 ensures perfect alignment between image pixels and model patch boundaries without padding artifacts. The loaders enforce this constraint and raise a `ValueError` if violated, though you can bypass resizing by ensuring your source data meets this requirement natively.

### Can I run the benchmark on video files without extracting frames manually?

Yes. When using the `GeneralDataset` loader ([`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py)), you can point `raw_data_root` to a video file (e.g., `.mp4`). The loader automatically extracts frames to a subdirectory named `<video_name>_frames/` on first initialization. If COLMAP reconstruction is enabled via `_use_colmap: true`, the pipeline will process these extracted frames to generate camera poses and intrinsics.

### Where are dataset poses and intrinsics cached when using COLMAP?

When `_use_colmap: true` is set for the `GeneralDataset`, the benchmark creates a `colmap_workspace/` directory inside the image folder. This workspace contains the COLMAP database, sparse reconstruction files, and cached camera parameters. Subsequent runs reuse this cache to avoid reprocessing, unless the directory is deleted or the source images change.