How to Set Up the Offline Rendering Pipeline Using batch_demo.py in LingBot-Map

LingBot-Map's batch_demo.py provides a headless offline pipeline that processes long video sequences or large image folders through four stages—data ingestion, model inference, prediction export, and headless rendering—to generate MP4 fly-throughs and GLB assets without the memory constraints of the interactive viewer.

LingBot-Map is an open-source mapping system that includes a dedicated batch processing tool for datasets too large for the interactive demo.py viewer. The batch_demo.py script located in demo_render/ orchestrates the complete workflow from raw input to final visualization. This guide walks through the architecture, configuration, and execution of the offline rendering pipeline using the actual source implementation from the Robbyant/lingbot-map repository.

Understanding the Four-Stage Pipeline Architecture

The offline pipeline in demo_render/batch_demo.py is architected as a linear sequence of four distinct stages. Each stage is designed to handle memory-intensive operations through disk caching and incremental processing.

Stage 1: Data Ingestion and Preprocessing

The pipeline begins by ingesting frames from either a video file or a directory of images. When processing video input via --video_path, the script calls extract_frames_from_video to decode frames using OpenCV. These frames can be optionally cached as PNG files to disk using --save_frames_dir to avoid re-extraction during iterative development.

For image folder inputs via --input_folder, the script invokes list_image_paths to build a filtered file list. Filtering supports range selection (--first_k, --last_k), stride sampling, and explicit frame indices. All frames are then passed through load_and_preprocess_images from lingbot_map/utils/load_fn.py, which performs canonical cropping to a standardized resolution and pads dimensions to multiples of the model's patch size.

Stage 2: Model Inference (Streaming vs. Windowed)

Once preprocessed, frames are fed into the LingBot-Map model loaded via load_model. The script supports two inference modes controlled by the --mode argument:

  • Streaming inference (model.inference_streaming): Maintains a persistent KV-cache across consecutive frames for temporal consistency. The cache can be pruned using --keyframe_interval to limit memory growth during long sequences.
  • Windowed inference (model.inference_windowed): Processes the sequence in overlapping chunks defined by --window_size. Overlap is calculated from --overlap_keyframes or --overlap_size to ensure smooth transitions between windows.

Stage 3: Prediction Export to NPZ

After inference completes, per-frame predictions are serialized as compressed .npz files using save_predictions_npz. This format enables fast parallel I/O during the subsequent rendering phase via load_predictions_from_npz. The NPZ layout stores depth maps, camera poses, and optional semantic masks in a structure optimized for random access.

Stage 4: Headless RGBD Rendering

The final stage executes the rgbd-render pipeline through render_with_pipeline. A PipelineConfig object is first instantiated from an optional YAML preset (--config) and then overridden by explicit CLI flags. The rendering subprocess performs:

  1. Voxelization: CUDA-accelerated voxel grid construction via voxelize_frame from demo_render/render_cuda_ext
  2. Camera Path Generation: Virtual trajectory creation using rgbd_render.camera.build_camera_path
  3. Overlay Composition: Optional trajectory trails, frustum visualization, and sky-mask overlays
  4. Video Encoding: Final MP4 output via the OfflinePipeline class

Configuration and Key Parameters

The pipeline behavior is controlled through a hierarchical configuration system. Base parameters are defined in YAML presets located in demo_render/config/ (such as indoor.yaml or outdoor_drive.yaml), then augmented by command-line arguments.

Critical parameters include:

  • --model_path: Path to the LingBot-Map checkpoint file (e.g., lingbot-map-long.pt)
  • --mode: Inference strategy (streaming or windowed)
  • --keyframe_interval: Frame interval for KV-cache trimming in streaming mode
  • --mask_sky: Enables ONNX-based sky segmentation using lingbot_map/vis/sky_segmentation.py
  • --save_predictions: Persists inference outputs to NPZ format for later reuse
  • --save_glb: Exports a textured mesh in GLB format alongside the video
  • --no_render: Skips video generation (useful when only GLB export is required)

Practical Usage Examples

The following commands demonstrate typical workflow patterns. Replace placeholder paths with your local directory structure.

Batch Processing Image Folders

Process multiple scene folders where each sub-directory contains an image sequence:

python demo_render/batch_demo.py \
    --input_folder /data/scenes \
    --output_folder /data/outputs \
    --model_path /path/to/lingbot-map-long.pt \
    --mode streaming \
    --keyframe_interval 2 \
    --mask_sky \
    --save_predictions

Processing Single Video Files

Extract frames on-the-fly from a video using windowed inference with custom overlap:

python demo_render/batch_demo.py \
    --video_path /data/video/indoor_travel.MP4 \
    --output_folder /data/outputs \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8 \
    --keyframe_interval 13 \
    --mask_sky \
    --save_predictions \
    --save_glb

Rendering from Saved Predictions

Skip inference entirely and render from previously cached NPZ files:

python demo_render/batch_demo.py \
    --load_predictions /data/outputs/indoor_travel.npz \
    --output_folder /data/outputs \
    --config demo_render/config/indoor.yaml \
    --no_render

Key Source Files and Implementation Details

Understanding the underlying source structure helps with debugging and customization:

File Purpose
demo_render/batch_demo.py Main orchestration script handling argument parsing, scene discovery, and pipeline coordination
lingbot_map/utils/load_fn.py Contains load_and_preprocess_images for canonical cropping and patch-size alignment
lingbot_map/vis/sky_segmentation.py ONNX runtime wrapper for sky-mask generation when --mask_sky is enabled
demo_render/render_cuda_ext/ CUDA extensions for voxelize_frame and frustum culling operations
rgbd_render/pipeline/offline.py Implements OfflinePipeline for headless video encoding

Summary

  • The batch_demo.py script provides a four-stage pipeline (ingestion, inference, export, render) for processing large datasets offline.
  • Streaming mode uses KV-caching for temporal consistency, while windowed mode processes sequences in overlapping chunks to manage memory.
  • Predictions are cached as NPZ files to decouple inference from rendering, enabling iterative visualization adjustments without re-running the model.
  • YAML configuration presets combined with CLI overrides provide flexible control over camera paths, voxelization parameters, and output formats.
  • The pipeline supports both video input (with OpenCV frame extraction) and image folder input (with filtering and stride control).

Frequently Asked Questions

What is the difference between streaming and windowed inference in batch_demo.py?

Streaming inference maintains a key-value cache across consecutive frames using model.inference_streaming, making it ideal for long continuous sequences where temporal consistency is critical. Windowed inference splits the sequence into discrete overlapping windows processed by model.inference_windowed, which limits memory usage to a fixed budget regardless of sequence length. Use streaming for high-fidelity reconstructions of continuous camera motion; use windowed mode when processing hours-long footage on limited GPU memory.

How does the sky masking feature work in the offline pipeline?

When --mask_sky is specified, the pipeline invokes the ONNX runtime through lingbot_map/vis/sky_segmentation.py to generate per-frame segmentation masks. These masks identify sky pixels which are then excluded from voxelization and mesh reconstruction to improve geometric accuracy. The masks are cached to disk to avoid redundant computation if the pipeline is re-run.

Can I run the rendering stage without re-running inference?

Yes. By specifying --load_predictions pointing to a directory of previously saved .npz files, the pipeline skips the model inference stage entirely. This is useful for adjusting rendering parameters—such as camera paths, overlays, or output resolution—without the computational cost of re-processing the source video. Combine this with --no_render to export only GLB meshes from existing predictions.

Where is the voxelization implemented in the source code?

The CUDA-accelerated voxelization kernel is implemented in demo_render/render_cuda_ext/voxelize_frame.cu and exposed to Python through the accompanying __init__.py. This extension is called by render_with_pipeline within the rgbd_render pipeline to convert per-frame depth maps into a unified voxel grid for mesh extraction and visualization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →