# How the Offline Rendering Pipeline Works with batch_demo.py in lingbot-map

> Explore the offline rendering pipeline in lingbot-map using batch_demo.py. Learn how it processes image sequences, aggregates world points and camera poses, and renders final reconstructions with Open3D or GLB export.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-25

---

**The offline rendering pipeline processes entire image sequences through the GCT streaming model, aggregates per-frame world points and camera poses via the stream aggregator, and renders the final reconstruction using Open3D-based viewers or GLB exporters from the `lingbot_map.vis` module.**

The lingbot-map repository implements a decoupled architecture where 3D reconstruction and visualization operate independently. When executing [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py) (the batch inference entry point), the system loads a complete folder of images, runs the GCT model in offline mode without KV-cache streaming constraints, and passes the accumulated geometry to specialized rendering utilities. This guide examines the source code path from raw image tensors to exported point clouds.

## Stage 1: Loading and Preprocessing the Image Batch

The pipeline begins in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) where the `load_images()` function reads a directory of frames or extracts them from a video file. This utility, defined in [`lingbot_map/utils/load_fn.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/load_fn.py), returns a batched tensor of shape **[S, 3, H, W]** representing S frames with RGB channels.

For batch processing, `load_and_preprocess_images()` applies normalization and resizing before returning the tensor stack. Unlike the interactive mode which loads frames on-demand, this stage materializes the entire sequence in memory to enable parallel or sequential processing without I/O bottlenecks.

## Stage 2: Model Initialization

After loading the data, the script invokes `load_model()` to instantiate either `GCTStream` from [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) (for standard streaming) or the windowed variant from [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py). The model receives hyperparameters directly from CLI arguments, including the checkpoint path and device configuration.

The initialized model runs in evaluation mode, ready to process the full batch without the memory-efficient KV-cache updates used in real-time streaming scenarios.

## Stage 3: Offline Inference and Aggregation

The core reconstruction loop iterates over the pre-loaded frames, calling the model's `forward()` method for each image. According to the implementation in [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py), this method returns a dictionary containing:

- **`world_points`**: 3D coordinates in world space with shape **[1, N, 3]**
- **`camera_pose`**: SE(3) transformation matrices with shape **[1, 4, 4]**

Rather than displaying results immediately, the pipeline appends these tensors to accumulators. The [`aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/aggregator/stream.py) module manages this collection, concatenating per-frame outputs into global point cloud and pose tensors. This aggregation happens entirely on the GPU when CUDA is available, minimizing host-device transfer overhead.

## Stage 4: Rendering and Export

Once inference completes, the aggregated data flows into the `lingbot_map.vis` package for offline rendering. Three primary utilities handle different output formats:

**[`point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/point_cloud_viewer.py)** constructs an Open3D point cloud from the accumulated world points. It utilizes helper functions from [`vis/utils.py`](https://github.com/Robbyant/lingbot-map/blob/main/vis/utils.py) such as `points_to_o3d()` to convert PyTorch tensors to Open3D format and `apply_pose()` to transform geometry. The `PointCloudViewer` class launches an interactive window supporting camera navigation and point colorization.

**[`glb_export.py`](https://github.com/Robbyant/lingbot-map/blob/main/glb_export.py)** provides the `GLBExporter` class for web-compatible exports. This utility packages the point cloud and camera trajectory into a GLB (glTF binary) file suitable for three.js viewers or Blender import.

**[`viser_wrapper.py`](https://github.com/Robbyant/lingbot-map/blob/main/viser_wrapper.py)** offers a thin abstraction around the *viser* library, enabling real-time streaming of the reconstructed scene if remote visualization is required during batch processing.

## Complete Pipeline Implementation

The following code illustrates the full offline rendering workflow implemented in [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py):

```python
import torch
from lingbot_map.utils.load_fn import load_images, load_and_preprocess_images
from lingbot_map.models.gct_stream import GCTStream
from lingbot_map.aggregator.stream import StreamAggregator
from lingbot_map.vis.point_cloud_viewer import PointCloudViewer
from lingbot_map.vis.glb_export import GLBExporter

# Stage 1: Load batch

images, paths, folder = load_images(image_folder="/path/to/frames")
images = load_and_preprocess_images(images)  # [S, 3, H, W]

# Stage 2: Initialize model

model = GCTStream(checkpoint_path="model.ckpt").cuda().eval()

# Stage 3: Offline inference with aggregation

aggregator = StreamAggregator()
with torch.no_grad():
    for i in range(images.shape[0]):
        output = model(images[i:i+1])  # Process single frame

        aggregator.add_frame(
            world_points=output["world_points"],
            camera_pose=output["camera_pose"]
        )

# Retrieve accumulated data

world_pts, cam_poses = aggregator.get_accumulated()

# Stage 4: Render and export

viewer = PointCloudViewer(points=world_pts, poses=cam_poses)
viewer.show()  # Interactive Open3D window

# Export to GLB for web viewing

exporter = GLBExporter(
    points=world_pts,
    poses=cam_poses,
    output_path="reconstruction.glb"
)
exporter.save()

```

## Summary

- **Batch loading** occurs via `load_images()` and `load_and_preprocess_images()` in [`lingbot_map/utils/load_fn.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/load_fn.py), producing a **[S, 3, H, W]** tensor stack.
- **Model inference** uses `GCTStream` from [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), returning per-frame `world_points` and `camera_pose` dictionaries.
- **Aggregation** is handled by `StreamAggregator` in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), which concatenates frame data into global point clouds.
- **Rendering** utilizes `lingbot_map.vis` utilities: [`point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/point_cloud_viewer.py) for Open3D visualization, [`glb_export.py`](https://github.com/Robbyant/lingbot-map/blob/main/glb_export.py) for web formats, and [`viser_wrapper.py`](https://github.com/Robbyant/lingbot-map/blob/main/viser_wrapper.py) for streaming servers.

## Frequently Asked Questions

### What distinguishes batch_demo.py from the interactive demo.py?

The interactive [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) processes frames sequentially with KV-cache optimizations for real-time performance, updating the visualization after each frame. In contrast, [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py) disables streaming optimizations, loads the entire dataset upfront, and defers all rendering until the complete reconstruction is aggregated, enabling higher throughput for offline processing.

### How does the aggregator manage memory for large image sequences?

The `StreamAggregator` class in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) maintains running tensors for world points and camera poses, concatenating new frames via `torch.cat()`. For very long sequences, you can modify the aggregator to implement windowed accumulation or downsample points before storage to prevent GPU memory exhaustion.

### Can the pipeline export to formats other than GLB?

Currently, the repository provides first-class support for GLB export via [`glb_export.py`](https://github.com/Robbyant/lingbot-map/blob/main/glb_export.py) and Open3D native formats through [`point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/point_cloud_viewer.py). To support additional formats like PLY or LAS, extend the [`vis/utils.py`](https://github.com/Robbyant/lingbot-map/blob/main/vis/utils.py) module with conversion functions and call them after aggregation, leveraging the existing `points_to_o3d()` utility as a starting point.

### Where is camera trajectory data stored during offline processing?

Camera poses are extracted from the model output dictionary under the key `camera_pose` and accumulated alongside world points in the `StreamAggregator`. Both tensors persist in GPU memory during inference and are only transferred to the host when rendering begins, ensuring efficient batch processing of high-resolution sequences.