How to Use the Offline Rendering Pipeline for LingBot-Map: Complete GLB Export Guide

The offline rendering pipeline converts raw model predictions into viewable GLB files using the predictions_to_glb function in lingbot_map/vis/glb_export.py, requiring only a prediction dictionary containing world points, confidence maps, and camera parameters.

The LingBot-Map repository (Robbyant/lingbot-map) provides a fully self‑contained visualization system that runs without network dependencies. Located in the lingbot_map.vis package, this pipeline processes GCT model outputs (depth maps, point clouds, and camera poses) to generate standardized 3‑D scenes ready for any WebGL viewer.

Input Data Requirements

The offline rendering pipeline expects a Python dictionary (preds) containing specific tensor arrays produced by the GCT model. Your prediction dict must include these keys:

  • world_points or world_points_from_depth: Array of shape (S, H, W, 3) containing 3‑D coordinates in the world frame for S frames.
  • world_points_conf or depth_conf: Array of shape (S, H, W) providing per‑point confidence scores.
  • images: RGB frames of shape (S, 3, H, W) or (S, H, W, 3).
  • extrinsic: Camera‑to‑world matrices of shape (S, 3, 4).
  • intrinsic: Camera calibration matrices of shape (S, 3, 3) required for frustum generation.

These tensors are typically generated by the GCT model defined in lingbot_map/models/gct_stream.py and serialized via benchmark utilities.

Core Export Function: predictions_to_glb

The primary entry point resides in lingbot_map/vis/glb_export.py. The predictions_to_glb function orchestrates the entire conversion workflow, accepting the prediction dictionary and returning a trimesh.Scene object ready for export.

Key parameters include:

  • prediction_mode: Choose between "Predicted Pointmap" (uses pre‑computed world_points) or "Predicted Depthmap" (re‑projects depth maps on‑the‑fly using unproject_depth_map_to_point_map).
  • conf_thres: Percentile threshold (0‑100) for confidence filtering.
  • mask_sky: Boolean flag to enable ONNX‑based sky segmentation.
  • target_dir: Path to dataset directory containing an images/ folder (required for sky masking).

Step‑by‑Step Pipeline Architecture

1. Rendering Mode Selection

The pipeline branches based on the prediction_mode argument. When using "Predicted Depthmap", the system invokes depth unprojection functions to regenerate 3‑D coordinates from depth maps and camera intrinsics, while "Predicted Pointmap" directly consumes the cached world_points tensor.

2. Optional Sky Segmentation Masking

When mask_sky=True, the _apply_sky_mask function (lines 115‑144 in lingbot_map/vis/glb_export.py) downloads a lightweight ONNX model (skyseg.onnx) on first use. This model runs locally to generate binary masks that exclude sky pixels from the confidence calculations, preventing distant atmospheric points from cluttering the output.

3. Confidence‑Based Point Filtering

The pipeline flattens the confidence map and applies a percentile threshold to remove low‑confidence outliers:

conf = pred_world_points_conf.reshape(-1)
conf_threshold = np.percentile(conf, conf_thres) if conf_thres > 0 else 0.0
conf_mask = (conf >= conf_threshold) & (conf > 1e-5)

Points failing this mask are excluded from final geometry generation.

4. Background Masking

Two mutually exclusive flags enable removal of points matching black or white background colors—particularly useful for indoor RGB‑D captures with uniform backdrops. These boolean masks combine with the confidence mask (lines 142‑152) before geometry construction.

5. Point Cloud and Camera Geometry Construction

Surviving vertices and RGB colors feed into trimesh.PointCloud (lines 170‑174). Simultaneously, the integrate_camera_into_scene helper (lines 447‑508) converts camera extrinsics into OpenGL‑compatible 4×4 matrices using get_opengl_conversion_matrix(), then constructs scaled frustum cones for each viewpoint.

6. Scene Alignment

The apply_scene_alignment function (lines 191‑208) reorients the entire scene using the first camera’s extrinsic matrix, ensuring the world‑up axis points vertically and the forward direction faces the camera.

7. Trajectory Tube Generation

When enabled, _build_trajectory_tube (lines 996‑1052 in lingbot_map/vis/point_cloud_viewer.py) generates a continuous tube mesh connecting camera positions, colored with Matplotlib colormaps to visualize traversal paths through the scene.

Complete Offline Export Example

from lingbot_map.vis.glb_export import predictions_to_glb

# `preds` contains the tensors described in the Input Data Requirements section

glb_scene = predictions_to_glb(
    predictions=preds,
    conf_thres=50.0,            # Retain top 50% most confident points

    mask_sky=True,              # Activate sky segmentation

    target_dir="my_dataset/",   # Directory containing images/ folder

    prediction_mode="Predicted Pointmap",
)

# Export to binary GLB format

glb_scene.export("my_output_scene.glb")
print("Offline GLB exported successfully")

The resulting my_output_scene.glb opens directly in Blender, Sketchfab, or three.js without requiring server‑side components.

Interactive Alternative: PointCloudViewer

For exploratory visualization before final export, instantiate the PointCloudViewer class from lingbot_map/vis/point_cloud_viewer.py. This wrapper reuses the same offline pipeline while providing a local Viser web interface for screenshot capture and parameter tuning:

from lingbot_map.vis.point_cloud_viewer import PointCloudViewer

viewer = PointCloudViewer(
    pred_dict=preds,
    mask_sky=True,
    image_folder="my_dataset/",
    show_camera=True,
    glb_output_path="export.glb",
)

viewer.animate()  # Launches offline web UI

Key Source Files Reference

File Responsibility
lingbot_map/vis/glb_export.py Core offline export logic, confidence filtering, sky masking, camera integration, scene alignment
lingbot_map/vis/point_cloud_viewer.py Interactive UI with screenshot/video controls and trajectory visualization
lingbot_map/vis/sky_segmentation.py ONNX model loading and sky mask generation
lingbot_map/models/gct_stream.py GCT model definition producing prediction dictionaries

Summary

  • The offline rendering pipeline generates standalone GLB files from GCT model outputs without network dependencies.
  • Primary entry point is predictions_to_glb in lingbot_map/vis/glb_export.py, accepting prediction dictionaries with world_points, extrinsic, and intrinsic tensors.
  • Confidence filtering uses percentile thresholds to remove low‑quality points, while optional sky segmentation eliminates atmospheric artifacts via local ONNX inference.
  • Camera frustums are automatically generated with OpenGL‑compatible transforms and optional trajectory tubes.
  • Output files are standard binary GLB format compatible with any WebGL viewer.

Frequently Asked Questions

What input format does the offline rendering pipeline require?

The pipeline requires a Python dictionary containing world_points (or world_points_from_depth), confidence maps (world_points_conf), RGB images, extrinsic camera matrices, and intrinsic calibration matrices. These tensors typically come from the GCT model in lingbot_map/models/gct_stream.py.

Can the pipeline run completely without internet access?

Yes. Once the optional sky segmentation model (skyseg.onnx) is downloaded on first use, the pipeline operates entirely offline. All processing—from point cloud generation to GLB export—runs locally using trimesh and NumPy operations.

What is the difference between "Predicted Pointmap" and "Predicted Depthmap" modes?

"Predicted Pointmap" uses pre‑computed 3‑D coordinates directly from the world_points field, while "Predicted Depthmap" dynamically re‑projects depth values into 3‑D space using camera intrinsics via the unproject_depth_map_to_point_map function. Choose depthmap mode when you have raw depth predictions but no explicit point cloud.

How does confidence thresholding affect the output quality?

The conf_thres parameter (0‑100) sets a percentile cutoff for the confidence map. Points below this percentile are discarded, effectively filtering noise and uncertain geometry. A threshold of 50.0 keeps the top half of points by confidence, balancing density and accuracy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →