# How to Implement Custom Camera Paths in the Offline Rendering Pipeline YAML Configuration

> Learn to implement custom camera paths in the offline rendering pipeline YAML configuration. This guide details adding camera paths and modifying viewer settings for precise control.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-31

---

**You can implement custom camera paths in lingbot-map's offline rendering pipeline by adding a `camera_path` list to your YAML configuration, implementing a `load_camera_path()` helper to parse poses into tensors, and modifying the viewer to apply these poses via `client.camera.set_pose()` instead of the learned camera head.**

The lingbot-map repository provides an offline rendering system that generates scene videos by sequencing camera poses through a visualizer. To implement **custom camera paths in the offline rendering pipeline YAML configuration**, you extend the benchmark config with a trajectory definition and bypass the learned `CameraHead` prediction. This enables cinematic fly-throughs, evaluation trajectories, and debugging visualizations without retraining the model.

## Understanding the Architecture

The offline rendering pipeline constructs videos frame-by-frame by feeding camera poses to the visualizer in `lingbot_map/vis`. By default, the system uses a learned **CameraHead** ([`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py)) to predict viewpoints. To override this with custom trajectories, you provide a `camera_path` section in the YAML that specifies timestamped keyframes with positions and orientations. The configuration parser in [`benchmark/benchmark/core/config.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/config.py) loads these entries as a dictionary, which a trajectory loader then converts into PyTorch tensors for the rendering loop.

## Implementation Workflow

### Step 1: Define Camera Keyframes in YAML

The benchmark expects a top-level `camera_path` list where each entry represents a camera keyframe. Every keyframe requires a `timestamp`, a `position` array `[x, y, z]`, and an `orientation` quaternion `[qw, qx, qy, qz]`.

Create or modify a method configuration file to include your trajectory:

```yaml

# benchmark/configs/methods/lingbot_map_custom.yaml

model: lingbot_map
env: lingbot-map
_checkpoint: /path/to/lingbot-map.pt
_device: cuda
_mode: streaming

camera_path:
  - timestamp: 0.0
    position: [0.0, 0.0, 0.0]
    orientation: [1.0, 0.0, 0.0, 0.0]   # w, x, y, z

  - timestamp: 0.5
    position: [1.0, 0.2, 0.3]
    orientation: [0.9239, 0.0, 0.3827, 0.0]
  - timestamp: 1.0
    position: [2.0, 0.5, 0.6]
    orientation: [0.7071, 0.0, 0.7071, 0.0]

```

The parser in [`benchmark/benchmark/core/config.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/config.py) loads this structure automatically without requiring code changes, returning the entire configuration as a Python dictionary.

### Step 2: Implement the Trajectory Loader

You need a helper function to extract the `camera_path` list, validate the fields, and convert the data into tensors. Add a `load_camera_path()` function to [`benchmark/benchmark/io/trajectory.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/io/trajectory.py):

```python
import torch
from typing import List, Tuple

def load_camera_path(cfg: dict) -> List[Tuple[float, torch.Tensor]]:
    """
    Convert the ``camera_path`` section into (timestamp, pose_tensor) pairs.
    The pose tensor contains [tx, ty, tz, qw, qx, qy, qz].
    """
    path_cfg = cfg.get("camera_path", [])
    camera_path = []
    for entry in path_cfg:
        ts = float(entry["timestamp"])
        pos = torch.tensor(entry["position"], dtype=torch.float32)
        quat = torch.tensor(entry["orientation"], dtype=torch.float32)
        pose = torch.cat([pos, quat])  # shape (7,)

        camera_path.append((ts, pose))
    return camera_path

```

This returns a list of tuples that the rendering loop can iterate over, with each pose represented as a 7-element tensor combining translation and quaternion rotation.

### Step 3: Integrate Custom Poses into the Rendering Loop

Modify the visualizer in [`lingbot_map/vis/point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/point_cloud_viewer.py) to consume the loaded trajectory. Import the loader and apply each pose using `client.camera.set_pose()` before rendering:

```python
from benchmark.benchmark.io.trajectory import load_camera_path
from lingbot_map.vis.utils import CameraState

def render_offline(cfg):
    # Load custom trajectory

    custom_path = load_camera_path(cfg)
    
    for ts, pose in custom_path:
        # Convert tensor to CameraState (position + quaternion)

        cam_state = CameraState.from_pose_tensor(pose)
        client.camera.set_pose(cam_state)  # Apply custom pose

        
        # Capture frame

        render = client.camera.get_render(height=720, width=1280)
        # Process frame...

```

The `CameraState` dataclass in [`lingbot_map/vis/utils.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/utils.py) handles the conversion from the 7-element tensor to the camera pose format expected by the renderer.

### Step 4: Disable the Learned Camera Head

To ensure the pipeline uses your custom path exclusively, disable the `CameraHead` predictor by setting `enable_camera_head: false` in the same YAML configuration:

```yaml
enable_camera_head: false

```

When this flag is disabled, [`benchmark/viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/viewer.py) will not invoke the camera prediction logic from [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py), allowing the rendering loop to rely solely on the supplied `camera_path` poses.

## Summary

- **Add a `camera_path` section** to your YAML configuration with timestamped keyframes containing `position` and `orientation` arrays.
- **Implement `load_camera_path()`** in [`benchmark/benchmark/io/trajectory.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/io/trajectory.py) to parse YAML entries into `(timestamp, pose_tensor)` tuples.
- **Modify [`lingbot_map/vis/point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/point_cloud_viewer.py)** to apply poses via `client.camera.set_pose()` using the `CameraState` utility from [`lingbot_map/vis/utils.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/utils.py).
- **Set `enable_camera_head: false`** to bypass the learned camera predictor and use only custom trajectories.
- The configuration parser in [`benchmark/benchmark/core/config.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/config.py) requires no modification because it loads the entire YAML dictionary generically.

## Frequently Asked Questions

### What quaternion format does the `orientation` field expect?

The `orientation` field expects a quaternion in `[qw, qx, qy, qz]` format (w, x, y, z components). This corresponds to the scalar-first convention used throughout the `lingbot_map` geometry utilities. Ensure your quaternions are normalized to avoid rendering artifacts.

### Do I need to modify the core configuration parser to add custom camera paths?

No. The generic configuration parser in [`benchmark/benchmark/core/config.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/core/config.py) loads the entire YAML file into a Python dictionary without schema validation. You can add the `camera_path` key to any method configuration file, and it will be accessible via `cfg.get("camera_path", [])` without touching the parser code.

### How do I verify that the custom path is being used instead of the learned camera head?

Set `enable_camera_head: false` in your YAML configuration and add print statements or logging inside `load_camera_path()` to confirm the trajectory is loading. You can also inspect the output video: if the camera follows your specified coordinates exactly rather than learned predictions, the custom path is active.

### Can I combine custom camera paths with the learned camera head?

The architecture supports switching between modes via the `enable_camera_head` flag, but simultaneous use requires modifying [`benchmark/viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/viewer.py) to blend or select between sources. By default, the pipeline uses either the learned `CameraHead` predictions or the loaded `camera_path`, not both, to avoid pose conflicts during rendering.