How to Implement Custom Camera Paths in the Offline Rendering Pipeline YAML Configuration
You can implement custom camera paths in lingbot-map's offline rendering pipeline by adding a camera_path list to your YAML configuration, implementing a load_camera_path() helper to parse poses into tensors, and modifying the viewer to apply these poses via client.camera.set_pose() instead of the learned camera head.
The lingbot-map repository provides an offline rendering system that generates scene videos by sequencing camera poses through a visualizer. To implement custom camera paths in the offline rendering pipeline YAML configuration, you extend the benchmark config with a trajectory definition and bypass the learned CameraHead prediction. This enables cinematic fly-throughs, evaluation trajectories, and debugging visualizations without retraining the model.
Understanding the Architecture
The offline rendering pipeline constructs videos frame-by-frame by feeding camera poses to the visualizer in lingbot_map/vis. By default, the system uses a learned CameraHead (lingbot_map/heads/camera_head.py) to predict viewpoints. To override this with custom trajectories, you provide a camera_path section in the YAML that specifies timestamped keyframes with positions and orientations. The configuration parser in benchmark/benchmark/core/config.py loads these entries as a dictionary, which a trajectory loader then converts into PyTorch tensors for the rendering loop.
Implementation Workflow
Step 1: Define Camera Keyframes in YAML
The benchmark expects a top-level camera_path list where each entry represents a camera keyframe. Every keyframe requires a timestamp, a position array [x, y, z], and an orientation quaternion [qw, qx, qy, qz].
Create or modify a method configuration file to include your trajectory:
# benchmark/configs/methods/lingbot_map_custom.yaml
model: lingbot_map
env: lingbot-map
_checkpoint: /path/to/lingbot-map.pt
_device: cuda
_mode: streaming
camera_path:
- timestamp: 0.0
position: [0.0, 0.0, 0.0]
orientation: [1.0, 0.0, 0.0, 0.0] # w, x, y, z
- timestamp: 0.5
position: [1.0, 0.2, 0.3]
orientation: [0.9239, 0.0, 0.3827, 0.0]
- timestamp: 1.0
position: [2.0, 0.5, 0.6]
orientation: [0.7071, 0.0, 0.7071, 0.0]
The parser in benchmark/benchmark/core/config.py loads this structure automatically without requiring code changes, returning the entire configuration as a Python dictionary.
Step 2: Implement the Trajectory Loader
You need a helper function to extract the camera_path list, validate the fields, and convert the data into tensors. Add a load_camera_path() function to benchmark/benchmark/io/trajectory.py:
import torch
from typing import List, Tuple
def load_camera_path(cfg: dict) -> List[Tuple[float, torch.Tensor]]:
"""
Convert the ``camera_path`` section into (timestamp, pose_tensor) pairs.
The pose tensor contains [tx, ty, tz, qw, qx, qy, qz].
"""
path_cfg = cfg.get("camera_path", [])
camera_path = []
for entry in path_cfg:
ts = float(entry["timestamp"])
pos = torch.tensor(entry["position"], dtype=torch.float32)
quat = torch.tensor(entry["orientation"], dtype=torch.float32)
pose = torch.cat([pos, quat]) # shape (7,)
camera_path.append((ts, pose))
return camera_path
This returns a list of tuples that the rendering loop can iterate over, with each pose represented as a 7-element tensor combining translation and quaternion rotation.
Step 3: Integrate Custom Poses into the Rendering Loop
Modify the visualizer in lingbot_map/vis/point_cloud_viewer.py to consume the loaded trajectory. Import the loader and apply each pose using client.camera.set_pose() before rendering:
from benchmark.benchmark.io.trajectory import load_camera_path
from lingbot_map.vis.utils import CameraState
def render_offline(cfg):
# Load custom trajectory
custom_path = load_camera_path(cfg)
for ts, pose in custom_path:
# Convert tensor to CameraState (position + quaternion)
cam_state = CameraState.from_pose_tensor(pose)
client.camera.set_pose(cam_state) # Apply custom pose
# Capture frame
render = client.camera.get_render(height=720, width=1280)
# Process frame...
The CameraState dataclass in lingbot_map/vis/utils.py handles the conversion from the 7-element tensor to the camera pose format expected by the renderer.
Step 4: Disable the Learned Camera Head
To ensure the pipeline uses your custom path exclusively, disable the CameraHead predictor by setting enable_camera_head: false in the same YAML configuration:
enable_camera_head: false
When this flag is disabled, benchmark/viewer.py will not invoke the camera prediction logic from lingbot_map/heads/camera_head.py, allowing the rendering loop to rely solely on the supplied camera_path poses.
Summary
- Add a
camera_pathsection to your YAML configuration with timestamped keyframes containingpositionandorientationarrays. - Implement
load_camera_path()inbenchmark/benchmark/io/trajectory.pyto parse YAML entries into(timestamp, pose_tensor)tuples. - Modify
lingbot_map/vis/point_cloud_viewer.pyto apply poses viaclient.camera.set_pose()using theCameraStateutility fromlingbot_map/vis/utils.py. - Set
enable_camera_head: falseto bypass the learned camera predictor and use only custom trajectories. - The configuration parser in
benchmark/benchmark/core/config.pyrequires no modification because it loads the entire YAML dictionary generically.
Frequently Asked Questions
What quaternion format does the orientation field expect?
The orientation field expects a quaternion in [qw, qx, qy, qz] format (w, x, y, z components). This corresponds to the scalar-first convention used throughout the lingbot_map geometry utilities. Ensure your quaternions are normalized to avoid rendering artifacts.
Do I need to modify the core configuration parser to add custom camera paths?
No. The generic configuration parser in benchmark/benchmark/core/config.py loads the entire YAML file into a Python dictionary without schema validation. You can add the camera_path key to any method configuration file, and it will be accessible via cfg.get("camera_path", []) without touching the parser code.
How do I verify that the custom path is being used instead of the learned camera head?
Set enable_camera_head: false in your YAML configuration and add print statements or logging inside load_camera_path() to confirm the trajectory is loading. You can also inspect the output video: if the camera follows your specified coordinates exactly rather than learned predictions, the custom path is active.
Can I combine custom camera paths with the learned camera head?
The architecture supports switching between modes via the enable_camera_head flag, but simultaneous use requires modifying benchmark/viewer.py to blend or select between sources. By default, the pipeline uses either the learned CameraHead predictions or the loaded camera_path, not both, to avoid pose conflicts during rendering.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →