How to Use the Offline Rendering Pipeline in LingBot-Map: Complete GLB Export Guide

The offline rendering pipeline in LingBot-Map converts raw GCT model predictions into self-contained GLB files using the predictions_to_glb function in lingbot_map/vis/glb_export.py, enabling fully local 3D visualization without network dependencies or external servers.

The LingBot-Map repository provides a fully-offline visualization system that transforms depth maps, point clouds, and camera poses into portable 3D scenes. This guide explains how to use the offline rendering pipeline in LingBot-Map to generate production-ready GLB files from your prediction dictionaries, covering confidence filtering, camera frustum generation, and scene alignment.

Understanding the Pipeline Architecture

The offline rendering pipeline processes raw predictions through a sequential 10-stage workflow implemented entirely in lingbot_map/vis/glb_export.py. The system ingests tensor dictionaries, applies optional masking and filtering, constructs point-cloud geometry, integrates camera frustums, aligns the scene to world coordinates, and serializes the result as a binary GLB file. Every stage operates locally without requiring network connectivity or external visualization servers.

Preparing Your Prediction Dictionary

The entry point for the offline rendering pipeline is a Python dictionary containing tensors produced by the GCT model (defined in lingbot_map/models/gct_stream.py).

Required Input Fields

Your prediction dictionary must contain these specific keys:

  • world_points or world_points_from_depth: Tensor of shape (S, H, W, 3) containing 3D world coordinates for each frame S
  • world_points_conf or depth_conf: Tensor of shape (S, H, W) with per-point confidence scores
  • images: Tensor of shape (S, 3, H, W) or (S, H, W, 3) containing RGB input frames
  • extrinsic: Tensor of shape (S, 3, 4) representing camera-to-world transformation matrices
  • intrinsic: Tensor of shape (S, 3, 3) containing camera intrinsics required for frustum generation

Configuring the Rendering Mode

The prediction_mode parameter in predictions_to_glb (lines 35-46 of glb_export.py) determines how 3D points are generated from your inputs.

Predicted Pointmap Mode

When set to "Predicted Pointmap", the pipeline uses pre-computed world_points directly from your input dictionary. This mode offers faster processing and is recommended when your model already outputs world-aligned point clouds.

Predicted Depthmap Mode

When set to "Predicted Depthmap", the pipeline re-projects depth maps on-the-fly using unproject_depth_map_to_point_map. Use this mode when working with raw depth predictions rather than pre-computed point clouds, allowing the pipeline to handle the unprojection math internally.

Filtering and Masking Strategies

The pipeline provides multiple filtering mechanisms to clean point clouds before export.

Confidence-Based Point Filtering

The pipeline applies percentile-based thresholding to remove low-confidence points. In lingbot_map/vis/glb_export.py, the confidence map is flattened and filtered according to your conf_thres parameter:

conf = pred_world_points_conf.reshape(-1)
conf_threshold = np.percentile(conf, conf_thres) if conf_thres > 0 else 0.0
conf_mask = (conf >= conf_threshold) & (conf > 1e-5)

Points below the threshold or with near-zero confidence are discarded from the final geometry.

Sky Segmentation Masking

When mask_sky=True and a target directory containing an images/ folder is provided, the pipeline runs a lightweight ONNX sky-segmentation model (skyseg.onnx). The _apply_sky_mask function (lines 115-144) downloads the model on first use, runs inference via run_skyseg, and multiplies the confidence map by the resulting binary mask. This effectively removes sky points that often cause artifacts in outdoor reconstructions.

Background Color Masking

Two mutually exclusive flags enable filtering of points matching pure black or white background colors. These boolean masks are applied on top of the confidence mask (lines 142-152), useful for cleaning indoor RGB-D captures with uniform studio backgrounds.

Constructing the 3D Scene

After filtering, surviving vertices and RGB colors are converted into a trimesh.PointCloud object (lines 170-174):

scene_3d = trimesh.Scene()
point_cloud_data = trimesh.PointCloud(vertices=vertices_3d, colors=colors_rgb)
scene_3d.add_geometry(point_cloud_data)

Camera Frustum Generation

Each camera matrix is converted to an OpenGL-compatible 4×4 homogeneous transform using get_opengl_conversion_matrix() (lines 336-342), which flips the Y and Z axes to match WebGL conventions. The integrate_camera_into_scene function (lines 447-508) builds cone meshes representing camera frustums, with configurable thickness and automatic scaling based on scene extent to maintain visible proportions regardless of world scale.

Scene Alignment

The apply_scene_alignment function (lines 191-208) re-orients the entire scene using the first camera's extrinsic matrix. This ensures the world-up axis points upward and the forward direction faces the camera, creating consistent orientation across different captures.

Trajectory Visualization

When exporting through the interactive viewer, the _build_trajectory_tube function (lines 996-1052 in lingbot_map/vis/point_cloud_viewer.py) creates cylindrical segments connecting sequential camera positions. These tubes are colored using Matplotlib colormaps to create smooth gradient trajectories that visualize camera movement through the scene.

Exporting GLB Files: Code Examples

Basic Offline Export

For headless export without UI interaction, use the core export function:

from lingbot_map.vis.glb_export import predictions_to_glb

# preds is your prediction dictionary from the GCT model

glb_scene = predictions_to_glb(
    predictions=preds,
    conf_thres=50.0,            # retain top 50% most confident points

    mask_sky=True,              # enable sky segmentation filtering

    target_dir="my_dataset/",   # directory containing images/ subfolder

    prediction_mode="Predicted Pointmap",
)

# Export to any file path

glb_scene.export("output_scene.glb")
print("GLB exported successfully")

Interactive Viewer Export

For visual validation before export, use the PointCloudViewer wrapper:

from lingbot_map.vis.point_cloud_viewer import PointCloudViewer

viewer = PointCloudViewer(
    pred_dict=preds,
    mask_sky=True,
    image_folder="my_dataset/",
    show_camera=True,
    glb_output_path="output_scene.glb",
)

viewer.animate()  # launches local Viser web UI for interactive exploration

Key Source Files and Functions

Understanding these modules helps with advanced customization:

Summary

  • The offline rendering pipeline in LingBot-Map converts prediction dictionaries to GLB files without requiring network connectivity or external visualization servers
  • Use predictions_to_glb in lingbot_map/vis/glb_export.py for programmatic, headless export operations
  • Configure rendering modes between pre-computed pointmaps and on-the-fly depth re-projection based on your model's output format
  • Apply confidence thresholding and optional sky segmentation to remove low-quality or distant sky points before export
  • Camera frustums and trajectory tubes provide essential spatial context for understanding camera coverage in the final 3D scene

Frequently Asked Questions

What input format does the offline rendering pipeline require?

The pipeline requires a Python dictionary containing world_points (or world_points_from_depth), confidence maps, RGB images, camera extrinsics of shape (S, 3, 4), and intrinsics of shape (S, 3, 3). These tensors are typically generated by the GCT model defined in lingbot_map/models/gct_stream.py and saved via the benchmark utilities.

How does the sky segmentation masking work?

When enabled via mask_sky=True, the pipeline downloads a lightweight ONNX model (skyseg.onnx) on first use through the _apply_sky_mask function. This model generates binary segmentation masks that remove sky pixels from the confidence calculation, preventing distant atmospheric points from creating artifacts in your point cloud.

Can I customize the camera frustums in the exported GLB?

Yes. The integrate_camera_into_scene function accepts parameters for frustum thickness and scene scaling. When using PointCloudViewer, you can adjust these properties interactively before export. The frustums automatically scale based on scene extent to maintain visible proportions across different capture scales.

Is the offline rendering pipeline compatible with custom prediction models?

Yes, provided your model outputs the required dictionary keys with compatible tensor shapes. The pipeline is model-agnostic regarding how predictions are generated, requiring only the standard keys: world_points, world_points_conf, images, extrinsic, and intrinsic. You can adapt custom models by formatting their outputs to match this specification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →