# LingBot-Map Benchmark Suite Datasets: Supported SLAM and Visual Odometry Loaders

> Explore LingBot-Map benchmark suite datasets. Discover support for KITTI, TUM RGB-D, ETH3D, Seven Scenes & more with standardized Python loaders for seamless SLAM and visual odometry integration.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: api-reference
- Published: 2026-07-25

---

**LingBot-Map supports 10 distinct SLAM and visual odometry datasets including KITTI, TUM RGB-D, ETH3D, Seven Scenes, and generic image folders, each implemented as a Python class inheriting from `BaseDataset` with standardized `get_scenes()`, `get_frame_list()`, and `load_frame_data()` methods.**

The **LingBot-Map** repository (`Robbyant/lingbot-map`) provides a modular benchmark infrastructure for visual SLAM and odometry research. This guide details every dataset supported in the **LingBot-Map benchmark suite**, including directory layout requirements, configuration parameters, and programmatic interfaces.

## Dataset Loader Architecture

All dataset loaders in LingBot-Map derive from the abstract `BaseDataset` class defined in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py). Each concrete implementation must provide three core methods:

- `get_scenes()` – Returns the list of scene identifiers available in the dataset.
- `get_frame_list(scene)` – Returns an ordered list of frame indices for a specific scene.
- `load_frame_data(scene, frame_id)` – Returns a dictionary containing at least an RGB image (`'rgb'`) and, when available, camera intrinsics (`'intrinsics'`) and pose (`'pose'`).

The registry in [`benchmark/benchmark/core/registry.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/registry.py) maps YAML configuration keys to these concrete classes, while [`benchmark/benchmark/core/loader.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/loader.py) serves as the high-level entry point for instantiation.

## Supported Datasets

LingBot-Map provides dedicated loaders for major academic benchmarks and generic image sources. Each requires a specific on-disk layout that the benchmark reads directly without file reorganization.

### Autonomous Driving and Outdoor

**KITTI Odometry** (`KittiDataset` – [`benchmark/datasets/kitti.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/kitti.py))

The KITTI loader expects the standard odometry devkit structure:

```text
<root>/poses/00.txt … 10.txt
<root>/sequences/00/image_2/xxxxx.png

```

Ground-truth poses reside in `poses/`, while camera calibration files are located at `sequences/<seq>/calib.txt`. The loader supports optional `target_size` resizing with automatic intrinsic rescaling for patch-aligned dimensions.

**Oxford Spires** (`OxfordSpiresDataset` – [`benchmark/datasets/oxford_spires.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/oxford_spires.py))

Designed for architectural mapping, this loader expects:

```text
<root>/oxford_spires/<scene>/rgb/*.png
<root>/oxford_spires/<scene>/poses.txt

```

The layout mirrors the General loader but enforces a fixed scene list and specific subfolder naming.

### Indoor and RGB-D

**TUM RGB-D** (`TumDataset` – [`benchmark/datasets/tum.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tum.py))

TUM sequences require timestamp-synchronized data:

```text
<root>/<scene_name>/rgb/*.png
<root>/<scene_name>/rgb.txt
<root>/<scene_name>/groundtruth.txt```

Intrinsics are automatically selected based on the Freiburg camera identifier (`freiburg1`, `freiburg2`, or `freiburg3`) extracted from the scene path.

**Seven Scenes** (`SevenScenesDataset` – [`benchmark/datasets/seven_scenes.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/seven_scenes.py))

This classic indoor benchmark uses:

```text
<root>/7scenes/<scene>/rgb/*.png
<root>/7scenes/<scene>/groundtruth.txt```

Intrinsics are hard-coded per scene within the loader implementation.

**NeuralRGB-D** (`NeuralRgbdDataset` – [`benchmark/datasets/neural_rgbd.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/neural_rgbd.py))

Designed for the NeuralRGB-D benchmark:

```text
<root>/data/<scene>/rgb/*.png
<root>/data/<scene>/pose.txt```

Poses are stored per frame in a specialized format used by neural rendering evaluations.

**ETH3D** (`Eth3dDataset` – [`benchmark/datasets/eth3d.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/eth3d.py))

The ETH3D loader expects undistorted imagery:

```text
<root>/undistorted/<scene>/images/*.png
<root>/undistorted/<scene>/poses.txt```

Intrinsics are read from the dataset's native calibration files.

### Cultural Heritage and Large-Scale

**Tanks & Temples** (`TntDataset` – [`benchmark/datasets/tnt.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tnt.py))

For large-scale reconstruction benchmarking:

```text
<root>/<scene>/000001.jpg …
<root>/<scene>/<scene>_COLMAP_SfM.log
<root>/<scene>/<scene>.ply
<root>/<scene>/<scene>.json
<root>/<scene>/<scene>_trans.txt```

Per-frame camera-to-world poses are parsed from the COLMAP SfM log. Intrinsics are approximated using `fx = fy ≈ 1.2 × width` when exact calibration is unavailable.

**Vision Benchmark in Rome (VBR)** (`VbrDataset` – [`benchmark/datasets/vbr.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/vbr.py))

VBR uses a processed aligned format:

```text
<root>/<scene>_processed_aligned/rgb/*.png
<root>/<scene>_processed_aligned/camera_pose.txt
<root>/<scene>_processed_aligned/intrinsics.txt
<root>/processed_gt/<scene>_gt.txt```

The loader expects intrinsics as a 3×3 `K` matrix and ground-truth poses in TUM format under `processed_gt/`.

### Specialized and Custom

**Droid-W** (`DroidWDataset` – [`benchmark/datasets/droid_w.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/droid_w.py))

This specialized loader uses COLMAP-generated trajectories:

```text
<root>/droid_w/<scene>/rgb/*.png
<root>/droid_w/<scene>/colmap.txt```

Intrinsics are extracted directly from the COLMAP reconstruction file.

**General Image Folders and Video** (`GeneralDataset` – [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py))

For custom data, the general loader accepts arbitrary image folders containing PNG, JPG, BMP, or TIFF files, or single video files:

```text
<root>/                    # Image folder mode

<root>/my_video.mp4       # Video mode (extracts to <root>/my_video_frames/)

```

When `_use_colmap: true` is set in the YAML configuration, the loader automatically runs COLMAP feature extraction, sequential matching, and mapping. Results are cached under `<image_dir>/colmap_workspace/`. The COLMAP binary must be available in the system `PATH` or specified via the `colmap_binary` parameter.

## Data Preparation and Configuration

Preparing data for the LingBot-Map benchmark suite requires strict adherence to directory layouts. The benchmark does not reshape file hierarchies; it only reads existing structures.

### Preparation Steps

1. **Download official datasets** from their respective sources (KITTI, TUM, etc.).
2. **Create the exact directory tree** shown in the loader specifications above.
3. **(Optional) Run COLMAP** for pose-less datasets by enabling `_use_colmap: true` in the configuration.
4. **Configure resizing** using `target_size: [width, height]` or `load_img_size`. Note that patch-based transformer models require dimensions to be multiples of 14; the loader raises a `ValueError` if this constraint is violated while intrinsics are automatically rescaled to match resized images (see `_cover_fit_center_crop` implementations in relevant loaders).

### YAML Configuration Examples

**KITTI Odometry:**

```yaml
datasets:
  kitti_odometry:
    dataset: kitti
    raw_data_root: /data/kitti_odometry
    sequences: ["00", "02"]
    target_size: [640, 480]
    _use_colmap: false

```

**General Dataset with COLMAP Reconstruction:**

```yaml
datasets:
  my_images:
    dataset: general
    raw_data_root: /data/my_image_folder
    _use_colmap: true
    load_img_size: 720

```

Validate your configuration using the demonstration script:

```bash
python demo.py --config configs/kitti.yaml

```

## Programmatic Usage

All datasets expose a unified Python interface regardless of underlying format.

### Loading Individual Frames

```python
from benchmark.datasets.kitti import KittiDataset

kitti = KittiDataset(
    raw_data_root="/data/kitti_odometry",
    sequences=["00"],
    target_size=[640, 480],
)

scene = kitti.get_scenes()[0]
frame_id = 10
frame = kitti.load_frame_data(scene, frame_id)

print(frame["rgb"].shape)          # (480, 640, 3)

print(frame["intrinsics"])         # [fx, fy, cx, cy]

print(frame["pose"].shape)         # (4, 4) or None

```

### Iterating Over Scenes

```python
from benchmark.datasets.vbr import VbrDataset

vbr = VbrDataset(
    raw_data_root="/data/vbr",
    target_size=[512, 512],
)

for scene in vbr.get_scenes():
    for fid in vbr.get_frame_list(scene):
        data = vbr.load_frame_data(scene, fid)
        # Access data["rgb"], data["pose"], data["intrinsics"]

```

## Summary

- **LingBot-Map** provides 10 specialized dataset loaders located in `benchmark/datasets/`, ranging from autonomous driving (KITTI) to indoor SLAM (TUM, Seven Scenes).
- Every loader inherits from `BaseDataset` in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py) and implements `get_scenes()`, `get_frame_list()`, and `load_frame_data()`.
- Strict directory layouts are mandatory; the benchmark reads data without file reorganization.
- **COLMAP integration** via `_use_colmap: true` enables pose estimation for arbitrary image collections using the `GeneralDataset` loader.
- Image resizing with automatic intrinsic calibration is supported, requiring dimensions to be multiples of 14 for compatible vision models.

## Frequently Asked Questions

### How do I add a custom dataset to the LingBot-Map benchmark suite?

Create a new class inheriting from `BaseDataset` in [`benchmark/dataset/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/dataset/base.py), implement the three required abstract methods (`get_scenes()`, `get_frame_list()`, `load_frame_data()`), and register the class in [`benchmark/benchmark/core/registry.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/registry.py) with a unique YAML key. Follow the existing implementations in [`benchmark/datasets/kitti.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/kitti.py) or [`benchmark/datasets/tum.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tum.py) as reference templates.

### Does LingBot-Map automatically download dataset files?

No. You must manually download raw data from official sources (KITTI Odometry, TUM RGB-D, ETH3D, etc.) and organize them into the exact directory structures expected by each loader. The benchmark strictly reads existing files and does not perform downloads, decompression, or file reorganization.

### Why must target_size dimensions be multiples of 14?

Patch-based vision transformers (such as those used in modern SLAM models integrated with LingBot-Map) typically use a patch size of 14×14 pixels. The loaders enforce this constraint through validation checks that raise a `ValueError` if the requested dimensions violate the requirement, ensuring proper alignment for feature extraction while automatically rescaling camera intrinsics to match the new image resolution.

### Can I process video files directly without extracting frames manually?

Yes. The `GeneralDataset` loader accepts a video file path as `raw_data_root`. When configured with a video file, it automatically extracts frames to a subdirectory (e.g., `<root>/my_video_frames/`) on first load. You can also enable COLMAP reconstruction (`_use_colmap: true`) to generate poses and intrinsics for video sequences that lack calibration data.