What Benchmarking Datasets Are Used for LingBot-Map? A Complete Guide to SLAM Evaluation
LingBot-Map supports ten major computer vision datasets—including KITTI, TUM RGB-D, ETH3D, and Seven-Scenes—through a unified BaseDataset interface defined in benchmark/dataset/base.py that standardizes scene traversal, frame loading, and calibration handling.
The Robbyant/lingbot-map repository provides a modular benchmarking framework designed to evaluate SLAM and visual odometry algorithms against industry-standard datasets. Each dataset is encapsulated in a Python class implementing three core methods: get_scenes(), get_frame_list(scene), and load_frame_data(scene, frame_id). This architecture allows researchers to seamlessly switch between KITTI Odometry, TUM RGB-D, and other benchmarks without modifying evaluation code.
Core Dataset Architecture
All loaders inherit from BaseDataset located in benchmark/dataset/base.py. The abstract base class enforces a consistent contract for data retrieval:
get_scenes()returns available scene identifiersget_frame_list(scene)returns ordered frame indicesload_frame_data(scene, frame_id)returns a dictionary with'rgb'(image),'intrinsics'(camera parameters), and'pose'(ground-truth transformation when available)
The high-level entry point in benchmark/benchmark/core/loader.py instantiates specific loaders via a registry defined in benchmark/benchmark/core/registry.py, mapping YAML configuration strings to concrete classes.
Supported Benchmarking Datasets
LingBot-Map includes dedicated loaders for ten distinct datasets, each handling unique directory structures and calibration formats.
KITTI Odometry
The KittiDataset class in benchmark/datasets/kitti.py ingests the KITTI Vision Benchmark Suite. It expects the canonical layout with poses/ containing ground-truth trajectory files (00.txt through 10.txt) and sequences/ with stereo imagery.
Required structure:
<root>/poses/00.txt ... 10.txt
<root>/sequences/00/image_2/xxxxx.png
The loader reads calibration from sequences/<seq>/calib.txt and supports optional target_size resizing for patch-aligned processing.
TUM RGB-D
TumDataset in benchmark/datasets/tum.py handles the TUM RGB-D SLAM dataset. It automatically associates timestamps between RGB images and ground-truth poses.
Required structure:
<root>/<scene_name>/rgb/
<root>/<scene_name>/rgb.txt
<root>/<scene_name>/groundtruth.txt
Intrinsics are selected automatically based on the Freiburg camera identifier (freiburg1, freiburg2, or freiburg3).
ETH3D
The Eth3dDataset class in benchmark/datasets/eth3d.py loads the ETH3D SLAM benchmark, supporting both training and test sequences with undistorted images.
Required structure:
<root>/undistorted/<scene>/images/
<root>/undistorted/<scene>/poses.txt
Poses follow the ETH3D format, with intrinsics read from the dataset's calibration file.
Seven-Scenes
SevenScenesDataset in benchmark/datasets/seven_scenes.py manages the Microsoft 7-Scenes indoor dataset.
Required structure:
<root>/7scenes/<scene>/rgb/
<root>/7scenes/<scene>/groundtruth.txt
Camera intrinsics are hard-coded per scene according to the official specifications.
Tanks & Temples (TNT)
The TntDataset class in benchmark/datasets/tnt.py processes the Tanks & Temples benchmark using COLMAP SfM logs.
Required structure:
<root>/<scene>/000001.jpg
<root>/<scene>/<scene>_COLMAP_SfM.log
<root>/<scene>/<scene>.ply
<root>/<scene>/<scene>.json
<root>/<scene>/<scene>_trans.txt
Per-frame camera-to-world poses are extracted from the COLMAP log, with intrinsics approximated as fx = fy ≈ 1.2 * width.
Vision Benchmark in Rome (VBR)
VbrDataset in benchmark/datasets/vbr.py supports the Vision Benchmark in Rome dataset with processed and aligned sequences.
Required structure:
<root>/<scene>_processed_aligned/rgb/
<root>/<scene>_processed_aligned/camera_pose.txt
<root>/<scene>_processed_aligned/intrinsics.txt
<root>/processed_gt/<scene>_gt.txt
The loader expects a 3×3 K intrinsics matrix and TUM-format ground-truth poses.
NeuralRGB-D
The NeuralRgbdDataset in benchmark/datasets/neural_rgbd.py interfaces with the NeuralRGB-D benchmark data.
Required structure:
<root>/data/<scene>/rgb/
<root>/data/<scene>/pose.txt
Poses are stored per frame in the specific format required by neural rendering evaluations.
Oxford Spires
OxfordSpiresDataset in benchmark/datasets/oxford_spires.py loads the Oxford Spires dataset with a fixed scene list.
Required structure:
<root>/oxford_spires/<scene>/rgb/
<root>/oxford_spires/<scene>/poses.txt
Droid-W
The DroidWDataset class in benchmark/datasets/droid_w.py handles the Droid-W dataset using COLMAP-generated trajectories.
Required structure:
<root>/droid_w/<scene>/rgb/
<root>/droid_w/<scene>/colmap.txt
Intrinsics are parsed directly from the COLMAP reconstruction file.
General (Arbitrary Images or Video)
GeneralDataset in benchmark/datasets/general.py serves as a flexible fallback for custom image folders or video files. It supports PNG, JPG, BMP, and TIFF formats, with optional COLMAP reconstruction for pose estimation.
When _use_colmap: true is set in the YAML configuration, the loader invokes GeneralDataset._run_colmap to perform feature extraction, sequential matching, and mapping, caching results under <image_dir>/colmap_workspace/.
Dataset Preparation Requirements
Directory Layout Validation
LingBot-Map does not modify or relocate files; it strictly reads existing structures. You must organize downloaded data exactly as specified for each loader class. Verify your layout by running the demo script:
python demo.py --config configs/kitti.yaml
Image Resizing Constraints
Many loaders accept target_size or load_img_size parameters. When resizing, intrinsics are automatically scaled to match the new dimensions. Critical constraint: Patch-based transformer models require both width and height to be multiples of 14; the loader raises a ValueError if this requirement is violated.
COLMAP Integration
For datasets lacking ground-truth poses, enable _use_colmap: true in the configuration. The system requires the COLMAP binary to be available on the system PATH or specified via the colmap_binary parameter.
YAML Configuration Examples
Configure datasets in your experiment YAML using the registry keys:
KITTI Odometry:
datasets:
kitti_odometry:
dataset: kitti
raw_data_root: /data/kitti_odometry
sequences: ["00", "02"]
target_size: [640, 480]
_use_colmap: false
Custom image folder with COLMAP reconstruction:
datasets:
my_images:
dataset: general
raw_data_root: /data/my_image_folder
_use_colmap: true
load_img_size: 720
Loading Data Programmatically
Iterate through any dataset using the unified interface:
from benchmark.datasets.kitti import KittiDataset
dataset = KittiDataset(
raw_data_root="/data/kitti_odometry",
sequences=["00"],
target_size=[640, 480],
)
for scene in dataset.get_scenes():
for frame_id in dataset.get_frame_list(scene):
data = dataset.load_frame_data(scene, frame_id)
rgb = data["rgb"] # Shape: (480, 640, 3)
intrinsics = data["intrinsics"] # [fx, fy, cx, cy]
pose = data.get("pose") # 4×4 transformation or None
The returned dictionary standardizes access across all ten supported LingBot-Map benchmarking datasets.
Summary
- LingBot-Map provides dedicated loaders for ten major SLAM datasets: KITTI, TUM RGB-D, ETH3D, Seven-Scenes, Tanks & Temples, VBR, NeuralRGB-D, Oxford Spires, Droid-W, and general image folders.
- All loaders inherit from
BaseDatasetinbenchmark/dataset/base.pyand implementget_scenes(),get_frame_list(), andload_frame_data(). - Dataset classes are registered in
benchmark/benchmark/core/registry.pyand instantiated viabenchmark/benchmark/core/loader.py. - Directory structures must match exact expectations; no automatic file reorganization occurs.
- Image dimensions must be multiples of 14 when using patch-based transformers.
- COLMAP integration is available for pose estimation in custom datasets through
GeneralDataset.
Frequently Asked Questions
How do I add a new custom dataset to LingBot-Map?
Create a Python class inheriting from BaseDataset in benchmark/dataset/base.py, implementing the three required methods: get_scenes(), get_frame_list(scene), and load_frame_data(scene, frame_id). Register the class in benchmark/benchmark/core/registry.py with a unique string key, then reference this key in your YAML configuration under the dataset: field.
Why does LingBot-Map require image dimensions to be multiples of 14?
This constraint supports patch-based transformer architectures that process images using 14×14 pixel patches. When you specify a target_size or load_img_size, the loaders in KittiDataset and VbrDataset automatically handle resizing, but you must ensure the final dimensions satisfy this requirement or the loader will raise a ValueError to prevent runtime errors in the vision models.
Can I use LingBot-Map without ground-truth pose files?
Yes. Enable _use_colmap: true in your YAML configuration for the GeneralDataset loader. This triggers GeneralDataset._run_colmap to execute COLMAP's feature extraction, sequential matching, and mapper pipelines, generating camera poses and intrinsics from the images alone. The results are cached in <image_dir>/colmap_workspace/ to avoid recomputation on subsequent runs.
Where are the intrinsics stored for the TUM RGB-D dataset?
The TumDataset class in benchmark/datasets/tum.py does not read intrinsics from a separate file. Instead, it automatically selects the appropriate camera parameters based on the Freiburg camera identifier (freiburg1, freiburg2, or freiburg3) detected in the scene path. These hard-coded values match the official TUM RGB-D specifications for each camera type.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →