How to Use LingBot-Map with KITTI, Oxford Spires, and 7-Scenes Datasets

LingBot-Map supports KITTI, Oxford Spires, and 7-Scenes through modular dataset adapters that expose a unified API, allowing you to switch datasets by changing only the --dataset flag in demo.py without modifying any model code.

LingBot-Map is a modular benchmark framework for visual localization and mapping built around the Geometric Context Transformer (GCT). As implemented in the Robbyant/lingbot-map repository, the system abstracts dataset I/O through a common interface defined in benchmark/dataset/base.py, making it straightforward to evaluate the streaming LingbotMapMethod across diverse outdoor and indoor environments.

The Modular Dataset Architecture

The framework decouples model inference from data loading through two core abstractions: the BaseDataset abstract class and the BSSLoader streaming wrapper.

All dataset-specific logic is encapsulated in adapter classes that inherit from BaseDataset (benchmark/dataset/base.py). Each adapter—such as KittiDataset, OxfordSpiresDataset, and SevenScenesDataset—implements four critical methods: get_scenes(), get_frame_list(), load_frame_data(), and load_global_data(). This standardized interface ensures that the LingbotMapMethod (benchmark/methods/lingbot_map.py) receives consistent input regardless of the underlying dataset format.

Dataset registration occurs in benchmark/benchmark/core/registry.py, which maps short names like 'kitti', 'oxford_spires', and 'seven_scenes' to their respective Python classes. When you invoke demo.py with --dataset kitti, the benchmark looks up the registry and instantiates the correct adapter with your supplied raw_data_root.

The BSSLoader (benchmark/core/loader.py) acts as a thin streaming wrapper around any registered dataset. It handles frame-by-frame retrieval via load_frame_data(), optional resizing (e.g., target_size for KITTI), and patch alignment using a fixed 14-pixel stride that matches the GCT transformer's internal patch size.

Dataset-Specific Setup Requirements

Each dataset adapter expects a specific folder hierarchy. Preparing your data according to these layouts is the only prerequisite before running inference.

KITTI Odometry

The KITTI adapter (benchmark/datasets/kitti.py) expects the standard KITTI Odometry directory structure:

  • Pose files: <root>/poses/<seq>.txt
  • Image sequences: <root>/sequences/<seq>/image_2/…png

Run inference using the convenience script:

python demo.py --dataset kitti --raw_data_root /path/to/kitti

Oxford Spires

For large-scale outdoor scenes, the Oxford Spires adapter (benchmark/datasets/oxford_spires.py) reads images, intrinsics, and camera-to-world poses from scene-specific subdirectories:

  • Images: <root>/<scene>/images/…png
  • Intrinsics: <root>/<scene>/intrinsics.txt
  • Poses: <root>/<scene>/poses_c2w.txt

The adapter automatically handles on-the-fly resizing and intrinsics scaling. Launch the pipeline with:

python demo.py --dataset oxford_spires --raw_data_root /path/to/oxford_spires

7-Scenes

The indoor 7-Scenes adapter (benchmark/datasets/seven_scenes.py) organizes data by scene and sequence, utilizing the official train/test splits:

  • Color images: <root>/<scene>/seq-XX/frame-XX.color.png
  • Depth maps: <root>/<scene>/seq-XX/frame-XX.depth.png
  • Pose files: <root>/<scene>/seq-XX/frame-XX.pose.txt
  • Split definitions: TrainSplit.txt and TestSplit.txt

To evaluate on the test split:

python demo.py --dataset seven_scenes --raw_data_root /path/to/7scenes --split test

Programmatic Inference Examples

While demo.py provides a command-line interface, you can also instantiate the components directly in Python for custom workflows.

Running on KITTI Sequence 00

import argparse
from benchmark.core.loader import BSSLoader
from benchmark.methods.lingbot_map import LingbotMapMethod
from benchmark.datasets.kitti import KittiDataset

parser = argparse.ArgumentParser()
parser.add_argument('--raw_data_root', required=True)  # e.g. "/data/kitti_odometry"

parser.add_argument('--seq', default='00')
args = parser.parse_args()

# Build the dataset adapter

kitti = KittiDataset(raw_data_root=args.raw_data_root, sequences=[args.seq])

# Wrap it in a streaming loader

loader = BSSLoader(kitti)

# Initialise the model (default mode = "stream")

model = LingbotMapMethod(mode='stream')

# Run inference on the whole sequence

rgb_frames = [loader.load_frame_data(scene=args.seq, frame_id=i)['rgb']
              for i in range(loader.get_frame_count(scene=args.seq))]
outputs = model.run_inference(rgb_frames)

print("Finished KITTI sequence", args.seq, "- produced", len(outputs), "predictions")

Running on Oxford Spires

import argparse
from benchmark.core.loader import BSSLoader
from benchmark.methods.lingbot_map import LingbotMapMethod
from benchmark.datasets.oxford_spires import OxfordSpiresDataset

parser = argparse.ArgumentParser()
parser.add_argument('--raw_data_root', required=True)  # e.g. "/data/oxford_spires"

parser.add_argument('--scene', required=True)         # e.g. "keble-college-02"

args = parser.parse_args()

oxford = OxfordSpiresDataset(raw_data_root=args.raw_data_root)
loader = BSSLoader(oxford)

# Get the list of frame ids for the chosen scene

frame_ids = oxford.get_frame_list(args.scene)

rgb_frames = [oxford.load_frame_data(args.scene, fid)['rgb'] for fid in frame_ids]
model = LingbotMapMethod()
outputs = model.run_inference(rgb_frames)

print(f"Oxford scene {args.scene}: {len(outputs)} frames processed")

Running on 7-Scenes Test Split

import argparse
from benchmark.core.loader import BSSLoader
from benchmark.methods.lingbot_map import LingbotMapMethod
from benchmark.datasets.seven_scenes import SevenScenesDataset

parser = argparse.ArgumentParser()
parser.add_argument('--raw_data_root', required=True)  # e.g. "/data/7scenes"

parser.add_argument('--scene', required=True)         # e.g. "chess"

args = parser.parse_args()

seven = SevenScenesDataset(raw_data_root=args.raw_data_root, split='test')
loader = BSSLoader(seven)

frame_ids = seven.get_frame_list(args.scene)
rgb_frames = [seven.load_frame_data(args.scene, fid)['rgb'] for fid in frame_ids]

model = LingbotMapMethod()
outputs = model.run_inference(rgb_frames)

print(f"7-Scenes scene {args.scene}: {len(outputs)} frames processed")

Summary

  • Unified Interface: All dataset adapters inherit from BaseDataset (benchmark/dataset/base.py) and implement get_scenes(), get_frame_list(), and load_frame_data(), enabling seamless swapping between KITTI, Oxford Spires, and 7-Scenes.
  • Registry Pattern: The benchmark/benchmark/core/registry.py file maps short names (kitti, oxford_spires, seven_scenes) to adapter classes, allowing dataset selection via a single CLI flag.
  • Streaming Pipeline: BSSLoader (benchmark/core/loader.py) handles frame streaming, resizing, and 14-pixel patch alignment before feeding data to LingbotMapMethod (benchmark/methods/lingbot_map.py).
  • Zero-Code Switching: Changing datasets requires only pointing --raw_data_root to the correct directory and setting --dataset to the appropriate key; no modifications to the GCT model or inference logic are necessary.

Frequently Asked Questions

Do I need to modify code to switch between KITTI and Oxford Spires?

No. Because both datasets implement the same BaseDataset interface and are registered in benchmark/benchmark/core/registry.py, you can switch between them by simply changing the --dataset flag from kitti to oxford_spires and updating the --raw_data_root path accordingly. The BSSLoader and LingbotMapMethod remain agnostic to the specific dataset source.

What image patch size does LingBot-Map expect?

All dataset adapters use a fixed 14-pixel patch size for alignment, which matches the internal patch stride of the Geometric Context Transformer (GCT) implemented in LingbotMapMethod. The BSSLoader automatically handles this alignment during frame preprocessing.

How does the benchmark handle different camera intrinsics?

Each dataset adapter loads intrinsics via its load_frame_data() or load_global_data() method. For KITTI, intrinsics are parsed from the calibration files; for Oxford Spires, they are read from intrinsics.txt; and for 7-Scenes, fixed intrinsics are provided per scene. The adapters scale these parameters automatically when resizing is requested, ensuring the GCT receives consistent geometric context.

Can I evaluate on a specific sequence or scene only?

Yes. When using the Python API, instantiate the dataset adapter with specific arguments such as sequences=['00'] for KITTI or filter frame_ids after calling get_frame_list(scene) for Oxford Spires and 7-Scenes. The demo.py script processes all available scenes by default, but you can modify the loader loop to restrict evaluation to specific scene IDs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →