How to Integrate LingBot-Map with Custom RGBD Datasets: A Complete Implementation Guide
To integrate LingBot-Map with a custom RGBD dataset, subclass BaseDataset from benchmark/benchmark/dataset/base.py, implement the three required data access methods, ensure depth values are in meters with invalid pixels set to zero, and register your class via a YAML configuration file using its full Python import path.
The LingBot-Map framework separates data loading from model inference through a clean abstraction layer. By implementing the BaseDataset interface in the Robbyant/lingbot-map repository, you can stream arbitrary RGBD sequences through the neural mapping pipeline without touching the core model code in lingbot_map.models.
Understanding the BaseDataset Interface
LingBot-Map consumes RGBD data through a standardized dataset layer defined in benchmark/benchmark/dataset/base.py. This abstraction expects subclasses to implement three core methods that provide scene discovery, frame enumeration, and per-frame data loading. The built-in reference implementations in benchmark/datasets/neural_rgbd.py and benchmark/datasets/tum.py demonstrate how to handle custom folder layouts and file formats.
The framework wraps your dataset class with BSSLoader (located in benchmark/benchmark/core/loader.py), which handles frame sampling, caching, and conversion to the internal "Baselines Streaming Structure" (BSS) format. Because the dataset layer is stateless and pure-Python, you can plug in new sources without modifying the model inference logic.
Step-by-Step Integration Process
Step 1: Create a Dataset Subclass
Define a new Python class that inherits from BaseDataset and initializes the parent via super().__init__(raw_data_root, logger). Store any scene-level caches (such as pose files or intrinsic calibration) as instance variables to avoid redundant disk I/O.
- File location: Create your module at
benchmark/datasets/my_dataset.pyor as a standalone package. - Required import:
from benchmark.benchmark.dataset.base import BaseDataset. - Constructor signature:
def __init__(self, raw_data_root: str, logger=None):.
Step 2: Implement Required Data Access Methods
Your subclass must implement three methods that the BSSLoader calls during streaming:
get_scenes() -> list[str]: Return a list of scene identifiers (typically folder names) present in your dataset root.get_frame_list(scene: str) -> list[int]: Return a list of frame indices for the specified scene, usually derived by counting files in an image directory.load_frame_data(scene: str, frame_id: int) -> dict: Return a dictionary containing at minimum:'rgb': NumPy array(H, W, 3)of typeuint8.'depth': NumPy array(H, W)of typefloat32in meters, with invalid depth set to0.0.'intrinsics': NumPy array[fx, fy, cx, cy]of typefloat32.'pose': NumPy array(4, 4)of typefloat32in OpenCV convention, orNoneif ground truth is unavailable.
Step 3: Handle Depth Scaling and Intrinsics
The model expects metric depth. If your dataset stores depth in millimeters (as 16-bit PNG), convert it in load_frame_data:
depth = np.array(Image.open(depth_path), dtype=np.float32)
depth = depth / 1000.0 # Convert mm to meters
depth[depth > 10.0] = 0.0 # Clip far values
depth[depth < 0.001] = 0.0 # Clip near values
For intrinsics, always return a flat array [fx, fy, cx, cy]. The framework automatically handles image resizing, so you do not need to scale intrinsics manually when the model processes different resolutions.
Step 4: Register the Dataset via Configuration
Create a YAML configuration file to register your dataset class with the benchmark runner. The framework uses importlib to dynamically import your class:
# benchmark/config/my_custom_rgbd.yaml
dataset:
name: my_dataset.MyRgbdDataset # Full Python import path
raw_data_root: /absolute/path/to/data
benchmark:
evaluate_pointcloud: true
Run the benchmark with:
python benchmark/run.py \
--config benchmark/config/my_custom_rgbd.yaml \
--model_path /path/to/lingbot-map.pt \
--output_dir /tmp/results
Complete Custom Dataset Implementation
Below is a production-ready implementation for a dataset with the following structure:
/my_data_root/
scene_01/
rgb/00000.png, ...
depth/00000.png # 16-bit PNG (mm)
poses.txt # 4x4 matrices, one per line
focal.txt # Single float
# my_dataset.py
import numpy as np
from pathlib import Path
from PIL import Image
from benchmark.benchmark.dataset.base import BaseDataset
class MyRgbdDataset(BaseDataset):
"""Custom RGBD loader for LingBot-Map."""
def __init__(self, raw_data_root: str, logger=None):
super().__init__(raw_data_root, logger=logger)
self._pose_cache = {}
self._focal_cache = {}
def _load_poses(self, scene: str) -> np.ndarray:
"""Load 4x4 camera-to-world matrices."""
if scene in self._pose_cache:
return self._pose_cache[scene]
pose_file = self.raw_data_root / scene / "poses.txt"
if not pose_file.exists():
return np.empty((0, 4, 4), dtype=np.float32)
lines = [l.strip() for l in pose_file.read_text().splitlines() if l.strip()]
mats = []
for i in range(0, len(lines), 4):
mat = np.array(
[[float(x) for x in lines[i + r].split()] for r in range(4)],
dtype=np.float32,
)
# Convert OpenGL to OpenCV convention
mat[:, 1:3] *= -1.0
mats.append(mat)
self._pose_cache[scene] = np.stack(mats, axis=0)
return self._pose_cache[scene]
def _load_focal(self, scene: str) -> float:
"""Load focal length (assumes fx = fy)."""
if scene not in self._focal_cache:
focal_path = self.raw_data_root / scene / "focal.txt"
self._focal_cache[scene] = float(focal_path.read_text().strip())
return self._focal_cache[scene]
def get_scenes(self) -> list[str]:
"""Return scenes containing an 'rgb' subdirectory."""
return sorted(
d.name for d in self.raw_data_root.iterdir()
if d.is_dir() and (d / "rgb").exists()
)
def get_frame_list(self, scene: str) -> list[int]:
"""Return consecutive indices matching RGB file count."""
rgb_dir = self.raw_data_root / scene / "rgb"
return list(range(len(sorted(rgb_dir.glob("*.png")))))
def load_frame_data(self, scene: str, frame_id: int) -> dict:
"""Load RGB, depth, intrinsics, and pose for a specific frame."""
scene_dir = self.raw_data_root / scene
# Load RGB
rgb_path = scene_dir / "rgb" / f"{frame_id:05d}.png"
rgb = np.array(Image.open(rgb_path).convert("RGB"), dtype=np.uint8)
# Load and convert depth
depth_path = scene_dir / "depth" / f"{frame_id:05d}.png"
depth = np.array(Image.open(depth_path), dtype=np.float32)
depth = depth / 1000.0 # mm to meters
depth[depth > 10.0] = 0.0
# Load pose
poses = self._load_poses(scene)
pose = poses[frame_id] if poses.shape[0] > 0 else None
# Construct intrinsics
fx = fy = self._load_focal(scene)
cx, cy = 320.0, 240.0 # Adjust for your image resolution
intrinsics = np.array([fx, fy, cx, cy], dtype=np.float32)
return {
"rgb": rgb,
"depth": depth,
"pose": pose,
"intrinsics": intrinsics,
}
Running Inference on Custom Data
Once your dataset class is implemented, you can run LingBot-Map through multiple entry points:
Benchmark Evaluation: Use benchmark/run.py for quantitative evaluation. The runner instantiates your dataset class automatically based on the YAML configuration and computes metrics such as Chamfer distance and precision/recall if you implement the optional evaluate_pointcloud() method.
Interactive Demo: Run demo.py with the --dataset flag to visualize reconstruction in real-time:
python demo.py \
--model_path /path/to/lingbot-map.pt \
--dataset my_dataset.MyRgbdDataset \
--raw_data_root /absolute/path/to/data \
--scene scene_01 \
--mask_sky
Batch Rendering: For offline processing of long sequences, use demo_render/batch_demo.py, which also accepts the dataset class path and streams frames through the model without requiring a live display.
Summary
- Inherit from
BaseDatasetlocated atbenchmark/benchmark/dataset/base.pyto create a compatible data loader. - Implement three required methods:
get_scenes(),get_frame_list(), andload_frame_data()following the return type specifications. - Return metric depth (meters, not millimeters) and intrinsics as
[fx, fy, cx, cy]to ensure proper camera geometry handling. - Register via YAML by specifying the full Python import path in the
dataset.namefield; the framework resolves the class usingimportlib. - Use standard entry points (
benchmark/run.py,demo.py, orbatch_demo.py) without modification—theBSSLoaderautomatically wraps your dataset for streaming inference.
Frequently Asked Questions
What file format should I use for depth images?
LingBot-Map expects depth as a NumPy array of type float32 where values represent meters. If your source data uses 16-bit PNGs storing millimeters (common in RGBD datasets), divide by 1000.0 in your load_frame_data() implementation and set invalid pixels to 0.0, as demonstrated in benchmark/datasets/neural_rgbd.py.
Can I integrate datasets without ground truth poses?
Yes. If your dataset lacks camera poses, return None from load_frame_data() for the 'pose' key. The framework will populate the trajectory with NaN values and still run inference, though quantitative evaluation metrics that require pose alignment will be skipped.
How do I add custom evaluation metrics for my dataset?
Implement the optional evaluate_pointcloud() method in your BaseDataset subclass. This static method receives the predicted point cloud and ground truth data, returning a dictionary of metric names and float values. The benchmark runner in benchmark/run.py automatically calls this method when computing results if evaluate_pointcloud: true is set in your YAML configuration.
Where should I place my custom dataset Python file?
You can place your dataset module anywhere in the Python path. Common locations include benchmark/datasets/my_dataset.py (alongside neural_rgbd.py and tum.py) or as a separate pip-installable package. Ensure the module path in your YAML config matches the importable Python path (e.g., benchmark.datasets.my_dataset.MyRgbdDataset or my_package.my_dataset.MyRgbdDataset).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →