How to Use the Offline Rendering Pipeline in LingBot-Map: Complete GLB Export Guide
The offline rendering pipeline in LingBot-Map converts raw GCT model predictions into self-contained GLB files using the predictions_to_glb function in lingbot_map/vis/glb_export.py, enabling fully local 3D visualization without network dependencies or external servers.
The LingBot-Map repository provides a fully-offline visualization system that transforms depth maps, point clouds, and camera poses into portable 3D scenes. This guide explains how to use the offline rendering pipeline in LingBot-Map to generate production-ready GLB files from your prediction dictionaries, covering confidence filtering, camera frustum generation, and scene alignment.
Understanding the Pipeline Architecture
The offline rendering pipeline processes raw predictions through a sequential 10-stage workflow implemented entirely in lingbot_map/vis/glb_export.py. The system ingests tensor dictionaries, applies optional masking and filtering, constructs point-cloud geometry, integrates camera frustums, aligns the scene to world coordinates, and serializes the result as a binary GLB file. Every stage operates locally without requiring network connectivity or external visualization servers.
Preparing Your Prediction Dictionary
The entry point for the offline rendering pipeline is a Python dictionary containing tensors produced by the GCT model (defined in lingbot_map/models/gct_stream.py).
Required Input Fields
Your prediction dictionary must contain these specific keys:
world_pointsorworld_points_from_depth: Tensor of shape(S, H, W, 3)containing 3D world coordinates for each frameSworld_points_confordepth_conf: Tensor of shape(S, H, W)with per-point confidence scoresimages: Tensor of shape(S, 3, H, W)or(S, H, W, 3)containing RGB input framesextrinsic: Tensor of shape(S, 3, 4)representing camera-to-world transformation matricesintrinsic: Tensor of shape(S, 3, 3)containing camera intrinsics required for frustum generation
Configuring the Rendering Mode
The prediction_mode parameter in predictions_to_glb (lines 35-46 of glb_export.py) determines how 3D points are generated from your inputs.
Predicted Pointmap Mode
When set to "Predicted Pointmap", the pipeline uses pre-computed world_points directly from your input dictionary. This mode offers faster processing and is recommended when your model already outputs world-aligned point clouds.
Predicted Depthmap Mode
When set to "Predicted Depthmap", the pipeline re-projects depth maps on-the-fly using unproject_depth_map_to_point_map. Use this mode when working with raw depth predictions rather than pre-computed point clouds, allowing the pipeline to handle the unprojection math internally.
Filtering and Masking Strategies
The pipeline provides multiple filtering mechanisms to clean point clouds before export.
Confidence-Based Point Filtering
The pipeline applies percentile-based thresholding to remove low-confidence points. In lingbot_map/vis/glb_export.py, the confidence map is flattened and filtered according to your conf_thres parameter:
conf = pred_world_points_conf.reshape(-1)
conf_threshold = np.percentile(conf, conf_thres) if conf_thres > 0 else 0.0
conf_mask = (conf >= conf_threshold) & (conf > 1e-5)
Points below the threshold or with near-zero confidence are discarded from the final geometry.
Sky Segmentation Masking
When mask_sky=True and a target directory containing an images/ folder is provided, the pipeline runs a lightweight ONNX sky-segmentation model (skyseg.onnx). The _apply_sky_mask function (lines 115-144) downloads the model on first use, runs inference via run_skyseg, and multiplies the confidence map by the resulting binary mask. This effectively removes sky points that often cause artifacts in outdoor reconstructions.
Background Color Masking
Two mutually exclusive flags enable filtering of points matching pure black or white background colors. These boolean masks are applied on top of the confidence mask (lines 142-152), useful for cleaning indoor RGB-D captures with uniform studio backgrounds.
Constructing the 3D Scene
After filtering, surviving vertices and RGB colors are converted into a trimesh.PointCloud object (lines 170-174):
scene_3d = trimesh.Scene()
point_cloud_data = trimesh.PointCloud(vertices=vertices_3d, colors=colors_rgb)
scene_3d.add_geometry(point_cloud_data)
Camera Frustum Generation
Each camera matrix is converted to an OpenGL-compatible 4×4 homogeneous transform using get_opengl_conversion_matrix() (lines 336-342), which flips the Y and Z axes to match WebGL conventions. The integrate_camera_into_scene function (lines 447-508) builds cone meshes representing camera frustums, with configurable thickness and automatic scaling based on scene extent to maintain visible proportions regardless of world scale.
Scene Alignment
The apply_scene_alignment function (lines 191-208) re-orients the entire scene using the first camera's extrinsic matrix. This ensures the world-up axis points upward and the forward direction faces the camera, creating consistent orientation across different captures.
Trajectory Visualization
When exporting through the interactive viewer, the _build_trajectory_tube function (lines 996-1052 in lingbot_map/vis/point_cloud_viewer.py) creates cylindrical segments connecting sequential camera positions. These tubes are colored using Matplotlib colormaps to create smooth gradient trajectories that visualize camera movement through the scene.
Exporting GLB Files: Code Examples
Basic Offline Export
For headless export without UI interaction, use the core export function:
from lingbot_map.vis.glb_export import predictions_to_glb
# preds is your prediction dictionary from the GCT model
glb_scene = predictions_to_glb(
predictions=preds,
conf_thres=50.0, # retain top 50% most confident points
mask_sky=True, # enable sky segmentation filtering
target_dir="my_dataset/", # directory containing images/ subfolder
prediction_mode="Predicted Pointmap",
)
# Export to any file path
glb_scene.export("output_scene.glb")
print("GLB exported successfully")
Interactive Viewer Export
For visual validation before export, use the PointCloudViewer wrapper:
from lingbot_map.vis.point_cloud_viewer import PointCloudViewer
viewer = PointCloudViewer(
pred_dict=preds,
mask_sky=True,
image_folder="my_dataset/",
show_camera=True,
glb_output_path="output_scene.glb",
)
viewer.animate() # launches local Viser web UI for interactive exploration
Key Source Files and Functions
Understanding these modules helps with advanced customization:
lingbot_map/vis/glb_export.py: Containspredictions_to_glb, confidence filtering, sky masking, camera integration, and GLB serializationlingbot_map/vis/point_cloud_viewer.py: Provides thePointCloudViewerclass wrapping the offline pipeline with screenshot, video recording, and interactive export controlslingbot_map/vis/sky_segmentation.py: Handles ONNX model loading and sky mask generation used by both export and viewer moduleslingbot_map/vis/utils.py: Defines helper dataclasses includingCameraStatefor managing camera parameterslingbot_map/models/gct_stream.py: Defines the GCT model architecture that produces compatible prediction dictionarieslingbot_map/utils/load_fn.py: Provides utilities for loading NumPy and PyTorch tensors from disk storage
Summary
- The offline rendering pipeline in LingBot-Map converts prediction dictionaries to GLB files without requiring network connectivity or external visualization servers
- Use
predictions_to_glbinlingbot_map/vis/glb_export.pyfor programmatic, headless export operations - Configure rendering modes between pre-computed pointmaps and on-the-fly depth re-projection based on your model's output format
- Apply confidence thresholding and optional sky segmentation to remove low-quality or distant sky points before export
- Camera frustums and trajectory tubes provide essential spatial context for understanding camera coverage in the final 3D scene
Frequently Asked Questions
What input format does the offline rendering pipeline require?
The pipeline requires a Python dictionary containing world_points (or world_points_from_depth), confidence maps, RGB images, camera extrinsics of shape (S, 3, 4), and intrinsics of shape (S, 3, 3). These tensors are typically generated by the GCT model defined in lingbot_map/models/gct_stream.py and saved via the benchmark utilities.
How does the sky segmentation masking work?
When enabled via mask_sky=True, the pipeline downloads a lightweight ONNX model (skyseg.onnx) on first use through the _apply_sky_mask function. This model generates binary segmentation masks that remove sky pixels from the confidence calculation, preventing distant atmospheric points from creating artifacts in your point cloud.
Can I customize the camera frustums in the exported GLB?
Yes. The integrate_camera_into_scene function accepts parameters for frustum thickness and scene scaling. When using PointCloudViewer, you can adjust these properties interactively before export. The frustums automatically scale based on scene extent to maintain visible proportions across different capture scales.
Is the offline rendering pipeline compatible with custom prediction models?
Yes, provided your model outputs the required dictionary keys with compatible tensor shapes. The pipeline is model-agnostic regarding how predictions are generated, requiring only the standard keys: world_points, world_points_conf, images, extrinsic, and intrinsic. You can adapt custom models by formatting their outputs to match this specification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →