# How to Set Up the Offline Rendering Pipeline Using batch_demo.py in LingBot-Map

> Effortlessly set up LingBot-Map's offline rendering pipeline with batch_demo.py. Process large datasets to generate MP4 and GLB assets, overcoming memory limits for seamless video creation.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-30

---

**LingBot-Map's [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py) provides a headless offline pipeline that processes long video sequences or large image folders through four stages—data ingestion, model inference, prediction export, and headless rendering—to generate MP4 fly-throughs and GLB assets without the memory constraints of the interactive viewer.**

LingBot-Map is an open-source mapping system that includes a dedicated batch processing tool for datasets too large for the interactive [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) viewer. The [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py) script located in `demo_render/` orchestrates the complete workflow from raw input to final visualization. This guide walks through the architecture, configuration, and execution of the offline rendering pipeline using the actual source implementation from the Robbyant/lingbot-map repository.

## Understanding the Four-Stage Pipeline Architecture

The offline pipeline in [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py) is architected as a linear sequence of four distinct stages. Each stage is designed to handle memory-intensive operations through disk caching and incremental processing.

### Stage 1: Data Ingestion and Preprocessing

The pipeline begins by ingesting frames from either a video file or a directory of images. When processing video input via `--video_path`, the script calls `extract_frames_from_video` to decode frames using OpenCV. These frames can be optionally cached as PNG files to disk using `--save_frames_dir` to avoid re-extraction during iterative development.

For image folder inputs via `--input_folder`, the script invokes `list_image_paths` to build a filtered file list. Filtering supports range selection (`--first_k`, `--last_k`), stride sampling, and explicit frame indices. All frames are then passed through `load_and_preprocess_images` from [`lingbot_map/utils/load_fn.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/load_fn.py), which performs canonical cropping to a standardized resolution and pads dimensions to multiples of the model's patch size.

### Stage 2: Model Inference (Streaming vs. Windowed)

Once preprocessed, frames are fed into the LingBot-Map model loaded via `load_model`. The script supports two inference modes controlled by the `--mode` argument:

- **Streaming inference** (`model.inference_streaming`): Maintains a persistent KV-cache across consecutive frames for temporal consistency. The cache can be pruned using `--keyframe_interval` to limit memory growth during long sequences.
- **Windowed inference** (`model.inference_windowed`): Processes the sequence in overlapping chunks defined by `--window_size`. Overlap is calculated from `--overlap_keyframes` or `--overlap_size` to ensure smooth transitions between windows.

### Stage 3: Prediction Export to NPZ

After inference completes, per-frame predictions are serialized as compressed `.npz` files using `save_predictions_npz`. This format enables fast parallel I/O during the subsequent rendering phase via `load_predictions_from_npz`. The NPZ layout stores depth maps, camera poses, and optional semantic masks in a structure optimized for random access.

### Stage 4: Headless RGBD Rendering

The final stage executes the **rgbd-render** pipeline through `render_with_pipeline`. A `PipelineConfig` object is first instantiated from an optional YAML preset (`--config`) and then overridden by explicit CLI flags. The rendering subprocess performs:

1. **Voxelization**: CUDA-accelerated voxel grid construction via `voxelize_frame` from `demo_render/render_cuda_ext`
2. **Camera Path Generation**: Virtual trajectory creation using `rgbd_render.camera.build_camera_path`
3. **Overlay Composition**: Optional trajectory trails, frustum visualization, and sky-mask overlays
4. **Video Encoding**: Final MP4 output via the `OfflinePipeline` class

## Configuration and Key Parameters

The pipeline behavior is controlled through a hierarchical configuration system. Base parameters are defined in YAML presets located in `demo_render/config/` (such as [`indoor.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/indoor.yaml) or [`outdoor_drive.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/outdoor_drive.yaml)), then augmented by command-line arguments.

Critical parameters include:

- `--model_path`: Path to the LingBot-Map checkpoint file (e.g., `lingbot-map-long.pt`)
- `--mode`: Inference strategy (`streaming` or `windowed`)
- `--keyframe_interval`: Frame interval for KV-cache trimming in streaming mode
- `--mask_sky`: Enables ONNX-based sky segmentation using [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py)
- `--save_predictions`: Persists inference outputs to NPZ format for later reuse
- `--save_glb`: Exports a textured mesh in GLB format alongside the video
- `--no_render`: Skips video generation (useful when only GLB export is required)

## Practical Usage Examples

The following commands demonstrate typical workflow patterns. Replace placeholder paths with your local directory structure.

### Batch Processing Image Folders

Process multiple scene folders where each sub-directory contains an image sequence:

```bash
python demo_render/batch_demo.py \
    --input_folder /data/scenes \
    --output_folder /data/outputs \
    --model_path /path/to/lingbot-map-long.pt \
    --mode streaming \
    --keyframe_interval 2 \
    --mask_sky \
    --save_predictions

```

### Processing Single Video Files

Extract frames on-the-fly from a video using windowed inference with custom overlap:

```bash
python demo_render/batch_demo.py \
    --video_path /data/video/indoor_travel.MP4 \
    --output_folder /data/outputs \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8 \
    --keyframe_interval 13 \
    --mask_sky \
    --save_predictions \
    --save_glb

```

### Rendering from Saved Predictions

Skip inference entirely and render from previously cached NPZ files:

```bash
python demo_render/batch_demo.py \
    --load_predictions /data/outputs/indoor_travel.npz \
    --output_folder /data/outputs \
    --config demo_render/config/indoor.yaml \
    --no_render

```

## Key Source Files and Implementation Details

Understanding the underlying source structure helps with debugging and customization:

| File | Purpose |
|------|---------|
| [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py) | Main orchestration script handling argument parsing, scene discovery, and pipeline coordination |
| [`lingbot_map/utils/load_fn.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/load_fn.py) | Contains `load_and_preprocess_images` for canonical cropping and patch-size alignment |
| [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py) | ONNX runtime wrapper for sky-mask generation when `--mask_sky` is enabled |
| `demo_render/render_cuda_ext/` | CUDA extensions for `voxelize_frame` and frustum culling operations |
| [`rgbd_render/pipeline/offline.py`](https://github.com/Robbyant/lingbot-map/blob/main/rgbd_render/pipeline/offline.py) | Implements `OfflinePipeline` for headless video encoding |

## Summary

- The [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py) script provides a **four-stage pipeline** (ingestion, inference, export, render) for processing large datasets offline.
- **Streaming mode** uses KV-caching for temporal consistency, while **windowed mode** processes sequences in overlapping chunks to manage memory.
- Predictions are cached as **NPZ files** to decouple inference from rendering, enabling iterative visualization adjustments without re-running the model.
- **YAML configuration presets** combined with CLI overrides provide flexible control over camera paths, voxelization parameters, and output formats.
- The pipeline supports both **video input** (with OpenCV frame extraction) and **image folder** input (with filtering and stride control).

## Frequently Asked Questions

### What is the difference between streaming and windowed inference in batch_demo.py?

**Streaming inference** maintains a key-value cache across consecutive frames using `model.inference_streaming`, making it ideal for long continuous sequences where temporal consistency is critical. **Windowed inference** splits the sequence into discrete overlapping windows processed by `model.inference_windowed`, which limits memory usage to a fixed budget regardless of sequence length. Use streaming for high-fidelity reconstructions of continuous camera motion; use windowed mode when processing hours-long footage on limited GPU memory.

### How does the sky masking feature work in the offline pipeline?

When `--mask_sky` is specified, the pipeline invokes the ONNX runtime through [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py) to generate per-frame segmentation masks. These masks identify sky pixels which are then excluded from voxelization and mesh reconstruction to improve geometric accuracy. The masks are cached to disk to avoid redundant computation if the pipeline is re-run.

### Can I run the rendering stage without re-running inference?

Yes. By specifying `--load_predictions` pointing to a directory of previously saved `.npz` files, the pipeline skips the model inference stage entirely. This is useful for adjusting rendering parameters—such as camera paths, overlays, or output resolution—without the computational cost of re-processing the source video. Combine this with `--no_render` to export only GLB meshes from existing predictions.

### Where is the voxelization implemented in the source code?

The CUDA-accelerated voxelization kernel is implemented in `demo_render/render_cuda_ext/voxelize_frame.cu` and exposed to Python through the accompanying [`__init__.py`](https://github.com/Robbyant/lingbot-map/blob/main/__init__.py). This extension is called by `render_with_pipeline` within the `rgbd_render` pipeline to convert per-frame depth maps into a unified voxel grid for mesh extraction and visualization.