LingBot-Map Benchmark Suite Datasets: Supported SLAM and Visual Odometry Loaders
LingBot-Map supports 10 distinct SLAM and visual odometry datasets including KITTI, TUM RGB-D, ETH3D, Seven Scenes, and generic image folders, each implemented as a Python class inheriting from BaseDataset with standardized get_scenes(), get_frame_list(), and load_frame_data() methods.
The LingBot-Map repository (Robbyant/lingbot-map) provides a modular benchmark infrastructure for visual SLAM and odometry research. This guide details every dataset supported in the LingBot-Map benchmark suite, including directory layout requirements, configuration parameters, and programmatic interfaces.
Dataset Loader Architecture
All dataset loaders in LingBot-Map derive from the abstract BaseDataset class defined in benchmark/dataset/base.py. Each concrete implementation must provide three core methods:
get_scenes()– Returns the list of scene identifiers available in the dataset.get_frame_list(scene)– Returns an ordered list of frame indices for a specific scene.load_frame_data(scene, frame_id)– Returns a dictionary containing at least an RGB image ('rgb') and, when available, camera intrinsics ('intrinsics') and pose ('pose').
The registry in benchmark/benchmark/core/registry.py maps YAML configuration keys to these concrete classes, while benchmark/benchmark/core/loader.py serves as the high-level entry point for instantiation.
Supported Datasets
LingBot-Map provides dedicated loaders for major academic benchmarks and generic image sources. Each requires a specific on-disk layout that the benchmark reads directly without file reorganization.
Autonomous Driving and Outdoor
KITTI Odometry (KittiDataset – benchmark/datasets/kitti.py)
The KITTI loader expects the standard odometry devkit structure:
<root>/poses/00.txt … 10.txt
<root>/sequences/00/image_2/xxxxx.png
Ground-truth poses reside in poses/, while camera calibration files are located at sequences/<seq>/calib.txt. The loader supports optional target_size resizing with automatic intrinsic rescaling for patch-aligned dimensions.
Oxford Spires (OxfordSpiresDataset – benchmark/datasets/oxford_spires.py)
Designed for architectural mapping, this loader expects:
<root>/oxford_spires/<scene>/rgb/*.png
<root>/oxford_spires/<scene>/poses.txt
The layout mirrors the General loader but enforces a fixed scene list and specific subfolder naming.
Indoor and RGB-D
TUM RGB-D (TumDataset – benchmark/datasets/tum.py)
TUM sequences require timestamp-synchronized data:
<root>/<scene_name>/rgb/*.png
<root>/<scene_name>/rgb.txt
<root>/<scene_name>/groundtruth.txt```
Intrinsics are automatically selected based on the Freiburg camera identifier (`freiburg1`, `freiburg2`, or `freiburg3`) extracted from the scene path.
**Seven Scenes** (`SevenScenesDataset` – [`benchmark/datasets/seven_scenes.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/seven_scenes.py))
This classic indoor benchmark uses:
```text
<root>/7scenes/<scene>/rgb/*.png
<root>/7scenes/<scene>/groundtruth.txt```
Intrinsics are hard-coded per scene within the loader implementation.
**NeuralRGB-D** (`NeuralRgbdDataset` – [`benchmark/datasets/neural_rgbd.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/neural_rgbd.py))
Designed for the NeuralRGB-D benchmark:
```text
<root>/data/<scene>/rgb/*.png
<root>/data/<scene>/pose.txt```
Poses are stored per frame in a specialized format used by neural rendering evaluations.
**ETH3D** (`Eth3dDataset` – [`benchmark/datasets/eth3d.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/eth3d.py))
The ETH3D loader expects undistorted imagery:
```text
<root>/undistorted/<scene>/images/*.png
<root>/undistorted/<scene>/poses.txt```
Intrinsics are read from the dataset's native calibration files.
### Cultural Heritage and Large-Scale
**Tanks & Temples** (`TntDataset` – [`benchmark/datasets/tnt.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/tnt.py))
For large-scale reconstruction benchmarking:
```text
<root>/<scene>/000001.jpg …
<root>/<scene>/<scene>_COLMAP_SfM.log
<root>/<scene>/<scene>.ply
<root>/<scene>/<scene>.json
<root>/<scene>/<scene>_trans.txt```
Per-frame camera-to-world poses are parsed from the COLMAP SfM log. Intrinsics are approximated using `fx = fy ≈ 1.2 × width` when exact calibration is unavailable.
**Vision Benchmark in Rome (VBR)** (`VbrDataset` – [`benchmark/datasets/vbr.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/vbr.py))
VBR uses a processed aligned format:
```text
<root>/<scene>_processed_aligned/rgb/*.png
<root>/<scene>_processed_aligned/camera_pose.txt
<root>/<scene>_processed_aligned/intrinsics.txt
<root>/processed_gt/<scene>_gt.txt```
The loader expects intrinsics as a 3×3 `K` matrix and ground-truth poses in TUM format under `processed_gt/`.
### Specialized and Custom
**Droid-W** (`DroidWDataset` – [`benchmark/datasets/droid_w.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/droid_w.py))
This specialized loader uses COLMAP-generated trajectories:
```text
<root>/droid_w/<scene>/rgb/*.png
<root>/droid_w/<scene>/colmap.txt```
Intrinsics are extracted directly from the COLMAP reconstruction file.
**General Image Folders and Video** (`GeneralDataset` – [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py))
For custom data, the general loader accepts arbitrary image folders containing PNG, JPG, BMP, or TIFF files, or single video files:
```text
<root>/ # Image folder mode
<root>/my_video.mp4 # Video mode (extracts to <root>/my_video_frames/)
When _use_colmap: true is set in the YAML configuration, the loader automatically runs COLMAP feature extraction, sequential matching, and mapping. Results are cached under <image_dir>/colmap_workspace/. The COLMAP binary must be available in the system PATH or specified via the colmap_binary parameter.
Data Preparation and Configuration
Preparing data for the LingBot-Map benchmark suite requires strict adherence to directory layouts. The benchmark does not reshape file hierarchies; it only reads existing structures.
Preparation Steps
- Download official datasets from their respective sources (KITTI, TUM, etc.).
- Create the exact directory tree shown in the loader specifications above.
- (Optional) Run COLMAP for pose-less datasets by enabling
_use_colmap: truein the configuration. - Configure resizing using
target_size: [width, height]orload_img_size. Note that patch-based transformer models require dimensions to be multiples of 14; the loader raises aValueErrorif this constraint is violated while intrinsics are automatically rescaled to match resized images (see_cover_fit_center_cropimplementations in relevant loaders).
YAML Configuration Examples
KITTI Odometry:
datasets:
kitti_odometry:
dataset: kitti
raw_data_root: /data/kitti_odometry
sequences: ["00", "02"]
target_size: [640, 480]
_use_colmap: false
General Dataset with COLMAP Reconstruction:
datasets:
my_images:
dataset: general
raw_data_root: /data/my_image_folder
_use_colmap: true
load_img_size: 720
Validate your configuration using the demonstration script:
python demo.py --config configs/kitti.yaml
Programmatic Usage
All datasets expose a unified Python interface regardless of underlying format.
Loading Individual Frames
from benchmark.datasets.kitti import KittiDataset
kitti = KittiDataset(
raw_data_root="/data/kitti_odometry",
sequences=["00"],
target_size=[640, 480],
)
scene = kitti.get_scenes()[0]
frame_id = 10
frame = kitti.load_frame_data(scene, frame_id)
print(frame["rgb"].shape) # (480, 640, 3)
print(frame["intrinsics"]) # [fx, fy, cx, cy]
print(frame["pose"].shape) # (4, 4) or None
Iterating Over Scenes
from benchmark.datasets.vbr import VbrDataset
vbr = VbrDataset(
raw_data_root="/data/vbr",
target_size=[512, 512],
)
for scene in vbr.get_scenes():
for fid in vbr.get_frame_list(scene):
data = vbr.load_frame_data(scene, fid)
# Access data["rgb"], data["pose"], data["intrinsics"]
Summary
- LingBot-Map provides 10 specialized dataset loaders located in
benchmark/datasets/, ranging from autonomous driving (KITTI) to indoor SLAM (TUM, Seven Scenes). - Every loader inherits from
BaseDatasetinbenchmark/dataset/base.pyand implementsget_scenes(),get_frame_list(), andload_frame_data(). - Strict directory layouts are mandatory; the benchmark reads data without file reorganization.
- COLMAP integration via
_use_colmap: trueenables pose estimation for arbitrary image collections using theGeneralDatasetloader. - Image resizing with automatic intrinsic calibration is supported, requiring dimensions to be multiples of 14 for compatible vision models.
Frequently Asked Questions
How do I add a custom dataset to the LingBot-Map benchmark suite?
Create a new class inheriting from BaseDataset in benchmark/dataset/base.py, implement the three required abstract methods (get_scenes(), get_frame_list(), load_frame_data()), and register the class in benchmark/benchmark/core/registry.py with a unique YAML key. Follow the existing implementations in benchmark/datasets/kitti.py or benchmark/datasets/tum.py as reference templates.
Does LingBot-Map automatically download dataset files?
No. You must manually download raw data from official sources (KITTI Odometry, TUM RGB-D, ETH3D, etc.) and organize them into the exact directory structures expected by each loader. The benchmark strictly reads existing files and does not perform downloads, decompression, or file reorganization.
Why must target_size dimensions be multiples of 14?
Patch-based vision transformers (such as those used in modern SLAM models integrated with LingBot-Map) typically use a patch size of 14×14 pixels. The loaders enforce this constraint through validation checks that raise a ValueError if the requested dimensions violate the requirement, ensuring proper alignment for feature extraction while automatically rescaling camera intrinsics to match the new image resolution.
Can I process video files directly without extracting frames manually?
Yes. The GeneralDataset loader accepts a video file path as raw_data_root. When configured with a video file, it automatically extracts frames to a subdirectory (e.g., <root>/my_video_frames/) on first load. You can also enable COLMAP reconstruction (_use_colmap: true) to generate poses and intrinsics for video sequences that lack calibration data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →