How the Offline Rendering Pipeline Works with batch_demo.py in lingbot-map
The offline rendering pipeline processes entire image sequences through the GCT streaming model, aggregates per-frame world points and camera poses via the stream aggregator, and renders the final reconstruction using Open3D-based viewers or GLB exporters from the lingbot_map.vis module.
The lingbot-map repository implements a decoupled architecture where 3D reconstruction and visualization operate independently. When executing batch_demo.py (the batch inference entry point), the system loads a complete folder of images, runs the GCT model in offline mode without KV-cache streaming constraints, and passes the accumulated geometry to specialized rendering utilities. This guide examines the source code path from raw image tensors to exported point clouds.
Stage 1: Loading and Preprocessing the Image Batch
The pipeline begins in demo.py where the load_images() function reads a directory of frames or extracts them from a video file. This utility, defined in lingbot_map/utils/load_fn.py, returns a batched tensor of shape [S, 3, H, W] representing S frames with RGB channels.
For batch processing, load_and_preprocess_images() applies normalization and resizing before returning the tensor stack. Unlike the interactive mode which loads frames on-demand, this stage materializes the entire sequence in memory to enable parallel or sequential processing without I/O bottlenecks.
Stage 2: Model Initialization
After loading the data, the script invokes load_model() to instantiate either GCTStream from lingbot_map/models/gct_stream.py (for standard streaming) or the windowed variant from gct_stream_window.py. The model receives hyperparameters directly from CLI arguments, including the checkpoint path and device configuration.
The initialized model runs in evaluation mode, ready to process the full batch without the memory-efficient KV-cache updates used in real-time streaming scenarios.
Stage 3: Offline Inference and Aggregation
The core reconstruction loop iterates over the pre-loaded frames, calling the model's forward() method for each image. According to the implementation in gct_stream.py, this method returns a dictionary containing:
world_points: 3D coordinates in world space with shape [1, N, 3]camera_pose: SE(3) transformation matrices with shape [1, 4, 4]
Rather than displaying results immediately, the pipeline appends these tensors to accumulators. The aggregator/stream.py module manages this collection, concatenating per-frame outputs into global point cloud and pose tensors. This aggregation happens entirely on the GPU when CUDA is available, minimizing host-device transfer overhead.
Stage 4: Rendering and Export
Once inference completes, the aggregated data flows into the lingbot_map.vis package for offline rendering. Three primary utilities handle different output formats:
point_cloud_viewer.py constructs an Open3D point cloud from the accumulated world points. It utilizes helper functions from vis/utils.py such as points_to_o3d() to convert PyTorch tensors to Open3D format and apply_pose() to transform geometry. The PointCloudViewer class launches an interactive window supporting camera navigation and point colorization.
glb_export.py provides the GLBExporter class for web-compatible exports. This utility packages the point cloud and camera trajectory into a GLB (glTF binary) file suitable for three.js viewers or Blender import.
viser_wrapper.py offers a thin abstraction around the viser library, enabling real-time streaming of the reconstructed scene if remote visualization is required during batch processing.
Complete Pipeline Implementation
The following code illustrates the full offline rendering workflow implemented in batch_demo.py:
import torch
from lingbot_map.utils.load_fn import load_images, load_and_preprocess_images
from lingbot_map.models.gct_stream import GCTStream
from lingbot_map.aggregator.stream import StreamAggregator
from lingbot_map.vis.point_cloud_viewer import PointCloudViewer
from lingbot_map.vis.glb_export import GLBExporter
# Stage 1: Load batch
images, paths, folder = load_images(image_folder="/path/to/frames")
images = load_and_preprocess_images(images) # [S, 3, H, W]
# Stage 2: Initialize model
model = GCTStream(checkpoint_path="model.ckpt").cuda().eval()
# Stage 3: Offline inference with aggregation
aggregator = StreamAggregator()
with torch.no_grad():
for i in range(images.shape[0]):
output = model(images[i:i+1]) # Process single frame
aggregator.add_frame(
world_points=output["world_points"],
camera_pose=output["camera_pose"]
)
# Retrieve accumulated data
world_pts, cam_poses = aggregator.get_accumulated()
# Stage 4: Render and export
viewer = PointCloudViewer(points=world_pts, poses=cam_poses)
viewer.show() # Interactive Open3D window
# Export to GLB for web viewing
exporter = GLBExporter(
points=world_pts,
poses=cam_poses,
output_path="reconstruction.glb"
)
exporter.save()
Summary
- Batch loading occurs via
load_images()andload_and_preprocess_images()inlingbot_map/utils/load_fn.py, producing a [S, 3, H, W] tensor stack. - Model inference uses
GCTStreamfromlingbot_map/models/gct_stream.py, returning per-frameworld_pointsandcamera_posedictionaries. - Aggregation is handled by
StreamAggregatorinlingbot_map/aggregator/stream.py, which concatenates frame data into global point clouds. - Rendering utilizes
lingbot_map.visutilities:point_cloud_viewer.pyfor Open3D visualization,glb_export.pyfor web formats, andviser_wrapper.pyfor streaming servers.
Frequently Asked Questions
What distinguishes batch_demo.py from the interactive demo.py?
The interactive demo.py processes frames sequentially with KV-cache optimizations for real-time performance, updating the visualization after each frame. In contrast, batch_demo.py disables streaming optimizations, loads the entire dataset upfront, and defers all rendering until the complete reconstruction is aggregated, enabling higher throughput for offline processing.
How does the aggregator manage memory for large image sequences?
The StreamAggregator class in lingbot_map/aggregator/stream.py maintains running tensors for world points and camera poses, concatenating new frames via torch.cat(). For very long sequences, you can modify the aggregator to implement windowed accumulation or downsample points before storage to prevent GPU memory exhaustion.
Can the pipeline export to formats other than GLB?
Currently, the repository provides first-class support for GLB export via glb_export.py and Open3D native formats through point_cloud_viewer.py. To support additional formats like PLY or LAS, extend the vis/utils.py module with conversion functions and call them after aggregation, leveraging the existing points_to_o3d() utility as a starting point.
Where is camera trajectory data stored during offline processing?
Camera poses are extracted from the model output dictionary under the key camera_pose and accumulated alongside world points in the StreamAggregator. Both tensors persist in GPU memory during inference and are only transferred to the host when rendering begins, ensuring efficient batch processing of high-resolution sequences.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →