How to Implement Custom Camera Paths in the Offline Rendering Pipeline YAML Configuration

You can implement custom camera paths in lingbot-map's offline rendering pipeline by adding a camera_path list to your YAML configuration, implementing a load_camera_path() helper to parse poses into tensors, and modifying the viewer to apply these poses via client.camera.set_pose() instead of the learned camera head.

The lingbot-map repository provides an offline rendering system that generates scene videos by sequencing camera poses through a visualizer. To implement custom camera paths in the offline rendering pipeline YAML configuration, you extend the benchmark config with a trajectory definition and bypass the learned CameraHead prediction. This enables cinematic fly-throughs, evaluation trajectories, and debugging visualizations without retraining the model.

Understanding the Architecture

The offline rendering pipeline constructs videos frame-by-frame by feeding camera poses to the visualizer in lingbot_map/vis. By default, the system uses a learned CameraHead (lingbot_map/heads/camera_head.py) to predict viewpoints. To override this with custom trajectories, you provide a camera_path section in the YAML that specifies timestamped keyframes with positions and orientations. The configuration parser in benchmark/benchmark/core/config.py loads these entries as a dictionary, which a trajectory loader then converts into PyTorch tensors for the rendering loop.

Implementation Workflow

Step 1: Define Camera Keyframes in YAML

The benchmark expects a top-level camera_path list where each entry represents a camera keyframe. Every keyframe requires a timestamp, a position array [x, y, z], and an orientation quaternion [qw, qx, qy, qz].

Create or modify a method configuration file to include your trajectory:


# benchmark/configs/methods/lingbot_map_custom.yaml

model: lingbot_map
env: lingbot-map
_checkpoint: /path/to/lingbot-map.pt
_device: cuda
_mode: streaming

camera_path:
  - timestamp: 0.0
    position: [0.0, 0.0, 0.0]
    orientation: [1.0, 0.0, 0.0, 0.0]   # w, x, y, z

  - timestamp: 0.5
    position: [1.0, 0.2, 0.3]
    orientation: [0.9239, 0.0, 0.3827, 0.0]
  - timestamp: 1.0
    position: [2.0, 0.5, 0.6]
    orientation: [0.7071, 0.0, 0.7071, 0.0]

The parser in benchmark/benchmark/core/config.py loads this structure automatically without requiring code changes, returning the entire configuration as a Python dictionary.

Step 2: Implement the Trajectory Loader

You need a helper function to extract the camera_path list, validate the fields, and convert the data into tensors. Add a load_camera_path() function to benchmark/benchmark/io/trajectory.py:

import torch
from typing import List, Tuple

def load_camera_path(cfg: dict) -> List[Tuple[float, torch.Tensor]]:
    """
    Convert the ``camera_path`` section into (timestamp, pose_tensor) pairs.
    The pose tensor contains [tx, ty, tz, qw, qx, qy, qz].
    """
    path_cfg = cfg.get("camera_path", [])
    camera_path = []
    for entry in path_cfg:
        ts = float(entry["timestamp"])
        pos = torch.tensor(entry["position"], dtype=torch.float32)
        quat = torch.tensor(entry["orientation"], dtype=torch.float32)
        pose = torch.cat([pos, quat])  # shape (7,)

        camera_path.append((ts, pose))
    return camera_path

This returns a list of tuples that the rendering loop can iterate over, with each pose represented as a 7-element tensor combining translation and quaternion rotation.

Step 3: Integrate Custom Poses into the Rendering Loop

Modify the visualizer in lingbot_map/vis/point_cloud_viewer.py to consume the loaded trajectory. Import the loader and apply each pose using client.camera.set_pose() before rendering:

from benchmark.benchmark.io.trajectory import load_camera_path
from lingbot_map.vis.utils import CameraState

def render_offline(cfg):
    # Load custom trajectory

    custom_path = load_camera_path(cfg)
    
    for ts, pose in custom_path:
        # Convert tensor to CameraState (position + quaternion)

        cam_state = CameraState.from_pose_tensor(pose)
        client.camera.set_pose(cam_state)  # Apply custom pose

        
        # Capture frame

        render = client.camera.get_render(height=720, width=1280)
        # Process frame...

The CameraState dataclass in lingbot_map/vis/utils.py handles the conversion from the 7-element tensor to the camera pose format expected by the renderer.

Step 4: Disable the Learned Camera Head

To ensure the pipeline uses your custom path exclusively, disable the CameraHead predictor by setting enable_camera_head: false in the same YAML configuration:

enable_camera_head: false

When this flag is disabled, benchmark/viewer.py will not invoke the camera prediction logic from lingbot_map/heads/camera_head.py, allowing the rendering loop to rely solely on the supplied camera_path poses.

Summary

Frequently Asked Questions

What quaternion format does the orientation field expect?

The orientation field expects a quaternion in [qw, qx, qy, qz] format (w, x, y, z components). This corresponds to the scalar-first convention used throughout the lingbot_map geometry utilities. Ensure your quaternions are normalized to avoid rendering artifacts.

Do I need to modify the core configuration parser to add custom camera paths?

No. The generic configuration parser in benchmark/benchmark/core/config.py loads the entire YAML file into a Python dictionary without schema validation. You can add the camera_path key to any method configuration file, and it will be accessible via cfg.get("camera_path", []) without touching the parser code.

How do I verify that the custom path is being used instead of the learned camera head?

Set enable_camera_head: false in your YAML configuration and add print statements or logging inside load_camera_path() to confirm the trajectory is loading. You can also inspect the output video: if the camera follows your specified coordinates exactly rather than learned predictions, the custom path is active.

Can I combine custom camera paths with the learned camera head?

The architecture supports switching between modes via the enable_camera_head flag, but simultaneous use requires modifying benchmark/viewer.py to blend or select between sources. By default, the pipeline uses either the learned CameraHead predictions or the loaded camera_path, not both, to avoid pose conflicts during rendering.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →