# How to Perform Offline Batch Rendering with LingBot-Map: A Complete Guide

> Learn to perform offline batch rendering with LingBot-Map. This guide details its four-stage pipeline for efficient video and image processing, overcoming memory limits.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-24

---

**LingBot-Map processes long video sequences or large image folders through a four-stage offline pipeline—data ingestion, model inference, prediction export, and headless rendering—without the memory constraints of the interactive viewer.**

The Robbyant/lingbot-map repository provides a dedicated offline batch rendering pipeline designed for production-scale workloads. Unlike the interactive [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) viewer, this headless system in [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py) handles arbitrarily long inputs through efficient caching, parallel I/O, and CUDA-accelerated voxelization.

## Understanding the Offline Batch Rendering Pipeline

The offline pipeline orchestrates four distinct stages to transform raw media into annotated 3D fly-through videos. According to the source code in [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py), the workflow handles scene discovery, frame extraction, neural inference, and final rendering through a unified command-line interface. Each stage is optimized for headless server deployment, enabling batch processing of multiple scenes without GPU display dependencies.

## Stage 1: Data Ingestion and Preprocessing

The pipeline accepts two input modalities: video files or sorted image folders. All frames undergo canonical preprocessing before reaching the model.

### Video Frame Extraction

When processing video inputs via `--video_path`, the system invokes `extract_frames_from_video` using OpenCV. Frames are extracted on-the-fly and optionally cached as PNG sequences to `--save_frames_dir` for reuse across multiple rendering passes. This prevents redundant decoding when iterating on visualization parameters.

### Image Folder Processing

For image sequences, the `list_image_paths` utility discovers frames within `--input_folder` and applies temporal filtering through `--first_k`, `--last_k`, or stride parameters. This allows processing of specific sub-sequences from large photogrammetry datasets.

### Canonical Preprocessing

All loaded frames pass through `load_and_preprocess_images` in [`lingbot_map/utils/load_fn.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/load_fn.py). This function performs:
- Resizing to model input dimensions
- Canonical cropping to remove lens distortion artifacts
- Padding to multiples of the patch size for transformer compatibility

## Stage 2: Model Inference Modes

After preprocessing, the system loads the LingBot-Map checkpoint via `load_model` and executes inference based on the `--mode` argument.

### Streaming Inference with KV-Cache

**Streaming mode** (`--mode streaming`) maintains a persistent KV-cache across frames using `model.inference_streaming`. This approach minimizes redundant computation for continuous video sequences. Control the cache refresh rate with `--keyframe_interval`, which determines how frequently the cache resets to prevent drift in long sequences.

### Windowed Inference for Long Sequences

**Windowed mode** (`--mode windowed`) splits sequences into overlapping chunks processed by `model.inference_windowed`. The overlap is controlled via `--overlap_keyframes` or `--overlap_size`, ensuring temporal consistency at segment boundaries. Use this mode for hour-long recordings that exceed GPU memory limits for full-sequence attention.

## Stage 3: Prediction Export and Caching

Following inference, per-frame predictions are serialized as compressed `.npz` files via `save_predictions_npz`. This format enables fast parallel I/O during the rendering stage through `load_predictions_from_npz`. Enable this behavior with `--save_predictions` to decouple expensive inference from iterative visualization tuning.

## Stage 4: Headless RGBD Rendering

The final stage invokes the **rgbd-render** pipeline through `render_with_pipeline`, operating entirely without display servers.

### Voxelization and Camera Path Generation

A `PipelineConfig` object is instantiated from an optional YAML preset (`--config`) and overridden by CLI flags. The pipeline calls `voxelize_frame` from the CUDA extension in `demo_render/render_cuda_ext` to build dense 3D representations. Camera trajectories are generated via `rgbd_render.camera.build_camera_path`, creating smooth fly-throughs based on the predicted camera poses.

### Overlay Composition and Video Encoding

The `OfflinePipeline` class composites optional visualizations including trajectory trails, camera frustums, and sky-mask visualizations. When `--mask_sky` is enabled, the system uses the ONNX segmentation model in [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py) to exclude sky regions from the point cloud. Final output is encoded as MP4 video or exported as GLB assets when `--save_glb` is specified.

## Practical Usage Examples

The following commands demonstrate typical workflows for the offline batch rendering pipeline.

Process multiple scene folders using streaming inference with sky masking:

```bash
python demo_render/batch_demo.py \
    --input_folder /data/scenes \
    --output_folder /data/outputs \
    --model_path /path/to/lingbot-map-long.pt \
    --mode streaming \
    --keyframe_interval 2 \
    --mask_sky \
    --save_predictions

```

Extract frames from a single video using windowed inference:

```bash
python demo_render/batch_demo.py \
    --video_path /data/video/indoor_travel.MP4 \
    --output_folder /data/outputs \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8 \
    --keyframe_interval 13 \
    --mask_sky \
    --save_predictions \
    --save_glb

```

Render from cached predictions without re-running inference:

```bash
python demo_render/batch_demo.py \
    --load_predictions /data/outputs/indoor_travel.npz \
    --output_folder /data/outputs \
    --config demo_render/config/indoor.yaml \
    --no_render

```

## Key Source Files and Architecture

Understanding the following components is essential for customizing the offline batch rendering pipeline:

- **[`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py)** – Entry point handling argument parsing, scene discovery, and stage orchestration.
- **[`lingbot_map/utils/load_fn.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/load_fn.py)** – Contains `load_and_preprocess_images` for canonical cropping and tensor preparation.
- **[`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py)** – ONNX-based sky segmentation used when `--mask_sky` is enabled.
- **[`rgbd_render/pipeline/builder.py`](https://github.com/Robbyant/lingbot-map/blob/main/rgbd_render/pipeline/builder.py)** and **[`rgbd_render/pipeline/offline.py`](https://github.com/Robbyant/lingbot-map/blob/main/rgbd_render/pipeline/offline.py)** – Construct scene representations and execute headless rendering.
- **`demo_render/render_cuda_ext/`** – CUDA extensions including `voxelize_frame.cu` for accelerated voxelization and frustum culling.
- **`demo_render/config/*.yaml`** – Preset configurations (e.g., [`indoor.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/indoor.yaml), [`outdoor_drive.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/outdoor_drive.yaml)) defining default camera paths and rendering parameters.

## Summary

- **LingBot-Map** provides a production-ready offline pipeline in [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py) for processing long sequences without interactive display requirements.
- The pipeline supports both **streaming inference** (with KV-cache management) and **windowed inference** (for memory-constrained long videos).
- Predictions are cached as **`.npz` files** to enable decoupled rendering and avoid redundant model inference.
- **CUDA-accelerated voxelization** and headless RGBD rendering generate MP4 fly-throughs and GLB assets from processed sequences.
- Sky masking, temporal filtering, and YAML-based configuration presets provide fine-grained control over the output visualization.

## Frequently Asked Questions

### What is the difference between streaming and windowed inference in LingBot-Map?

Streaming inference maintains a persistent KV-cache across frames using `model.inference_streaming`, making it efficient for continuous video where temporal consistency is maintained through cached attention states. Windowed inference uses `model.inference_windowed` to process sequences in overlapping chunks, which is necessary for very long videos that exceed GPU memory capacity but requires explicit overlap management via `--overlap_keyframes`.

### How do I handle sky masking in offline batch rendering?

Enable the `--mask_sky` flag to activate the ONNX segmentation model defined in [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py). The system segments sky regions from each frame and excludes them from the point cloud generation, preventing distant sky pixels from creating artifacts in the 3D reconstruction. Mask files are cached to disk to avoid recomputing segmentation during rendering iterations.

### Can I render videos from previously saved predictions without re-running inference?

Yes, use the `--load_predictions` argument to point to existing `.npz` files created by `save_predictions_npz`. When loading cached predictions, the pipeline skips the inference stage entirely and proceeds directly to the `render_with_pipeline` step. Combine with `--no_render` to perform format conversion (e.g., to GLB) without generating MP4 videos.

### What hardware requirements exist for the CUDA voxelization extensions?

The voxelization stage requires an NVIDIA GPU with CUDA support to run `voxelize_frame` from `demo_render/render_cuda_ext`. The extensions perform frustum culling and dense voxel grid generation on the GPU. While CPU fallback is not available for these specific operations, the inference stages can run on any PyTorch-supported device, though GPU acceleration is strongly recommended for real-time processing.