How to Perform Offline Batch Rendering with LingBot-Map: A Complete Guide

LingBot-Map processes long video sequences or large image folders through a four-stage offline pipeline—data ingestion, model inference, prediction export, and headless rendering—without the memory constraints of the interactive viewer.

The Robbyant/lingbot-map repository provides a dedicated offline batch rendering pipeline designed for production-scale workloads. Unlike the interactive demo.py viewer, this headless system in demo_render/batch_demo.py handles arbitrarily long inputs through efficient caching, parallel I/O, and CUDA-accelerated voxelization.

Understanding the Offline Batch Rendering Pipeline

The offline pipeline orchestrates four distinct stages to transform raw media into annotated 3D fly-through videos. According to the source code in demo_render/batch_demo.py, the workflow handles scene discovery, frame extraction, neural inference, and final rendering through a unified command-line interface. Each stage is optimized for headless server deployment, enabling batch processing of multiple scenes without GPU display dependencies.

Stage 1: Data Ingestion and Preprocessing

The pipeline accepts two input modalities: video files or sorted image folders. All frames undergo canonical preprocessing before reaching the model.

Video Frame Extraction

When processing video inputs via --video_path, the system invokes extract_frames_from_video using OpenCV. Frames are extracted on-the-fly and optionally cached as PNG sequences to --save_frames_dir for reuse across multiple rendering passes. This prevents redundant decoding when iterating on visualization parameters.

Image Folder Processing

For image sequences, the list_image_paths utility discovers frames within --input_folder and applies temporal filtering through --first_k, --last_k, or stride parameters. This allows processing of specific sub-sequences from large photogrammetry datasets.

Canonical Preprocessing

All loaded frames pass through load_and_preprocess_images in lingbot_map/utils/load_fn.py. This function performs:

  • Resizing to model input dimensions
  • Canonical cropping to remove lens distortion artifacts
  • Padding to multiples of the patch size for transformer compatibility

Stage 2: Model Inference Modes

After preprocessing, the system loads the LingBot-Map checkpoint via load_model and executes inference based on the --mode argument.

Streaming Inference with KV-Cache

Streaming mode (--mode streaming) maintains a persistent KV-cache across frames using model.inference_streaming. This approach minimizes redundant computation for continuous video sequences. Control the cache refresh rate with --keyframe_interval, which determines how frequently the cache resets to prevent drift in long sequences.

Windowed Inference for Long Sequences

Windowed mode (--mode windowed) splits sequences into overlapping chunks processed by model.inference_windowed. The overlap is controlled via --overlap_keyframes or --overlap_size, ensuring temporal consistency at segment boundaries. Use this mode for hour-long recordings that exceed GPU memory limits for full-sequence attention.

Stage 3: Prediction Export and Caching

Following inference, per-frame predictions are serialized as compressed .npz files via save_predictions_npz. This format enables fast parallel I/O during the rendering stage through load_predictions_from_npz. Enable this behavior with --save_predictions to decouple expensive inference from iterative visualization tuning.

Stage 4: Headless RGBD Rendering

The final stage invokes the rgbd-render pipeline through render_with_pipeline, operating entirely without display servers.

Voxelization and Camera Path Generation

A PipelineConfig object is instantiated from an optional YAML preset (--config) and overridden by CLI flags. The pipeline calls voxelize_frame from the CUDA extension in demo_render/render_cuda_ext to build dense 3D representations. Camera trajectories are generated via rgbd_render.camera.build_camera_path, creating smooth fly-throughs based on the predicted camera poses.

Overlay Composition and Video Encoding

The OfflinePipeline class composites optional visualizations including trajectory trails, camera frustums, and sky-mask visualizations. When --mask_sky is enabled, the system uses the ONNX segmentation model in lingbot_map/vis/sky_segmentation.py to exclude sky regions from the point cloud. Final output is encoded as MP4 video or exported as GLB assets when --save_glb is specified.

Practical Usage Examples

The following commands demonstrate typical workflows for the offline batch rendering pipeline.

Process multiple scene folders using streaming inference with sky masking:

python demo_render/batch_demo.py \
    --input_folder /data/scenes \
    --output_folder /data/outputs \
    --model_path /path/to/lingbot-map-long.pt \
    --mode streaming \
    --keyframe_interval 2 \
    --mask_sky \
    --save_predictions

Extract frames from a single video using windowed inference:

python demo_render/batch_demo.py \
    --video_path /data/video/indoor_travel.MP4 \
    --output_folder /data/outputs \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8 \
    --keyframe_interval 13 \
    --mask_sky \
    --save_predictions \
    --save_glb

Render from cached predictions without re-running inference:

python demo_render/batch_demo.py \
    --load_predictions /data/outputs/indoor_travel.npz \
    --output_folder /data/outputs \
    --config demo_render/config/indoor.yaml \
    --no_render

Key Source Files and Architecture

Understanding the following components is essential for customizing the offline batch rendering pipeline:

Summary

  • LingBot-Map provides a production-ready offline pipeline in demo_render/batch_demo.py for processing long sequences without interactive display requirements.
  • The pipeline supports both streaming inference (with KV-cache management) and windowed inference (for memory-constrained long videos).
  • Predictions are cached as .npz files to enable decoupled rendering and avoid redundant model inference.
  • CUDA-accelerated voxelization and headless RGBD rendering generate MP4 fly-throughs and GLB assets from processed sequences.
  • Sky masking, temporal filtering, and YAML-based configuration presets provide fine-grained control over the output visualization.

Frequently Asked Questions

What is the difference between streaming and windowed inference in LingBot-Map?

Streaming inference maintains a persistent KV-cache across frames using model.inference_streaming, making it efficient for continuous video where temporal consistency is maintained through cached attention states. Windowed inference uses model.inference_windowed to process sequences in overlapping chunks, which is necessary for very long videos that exceed GPU memory capacity but requires explicit overlap management via --overlap_keyframes.

How do I handle sky masking in offline batch rendering?

Enable the --mask_sky flag to activate the ONNX segmentation model defined in lingbot_map/vis/sky_segmentation.py. The system segments sky regions from each frame and excludes them from the point cloud generation, preventing distant sky pixels from creating artifacts in the 3D reconstruction. Mask files are cached to disk to avoid recomputing segmentation during rendering iterations.

Can I render videos from previously saved predictions without re-running inference?

Yes, use the --load_predictions argument to point to existing .npz files created by save_predictions_npz. When loading cached predictions, the pipeline skips the inference stage entirely and proceeds directly to the render_with_pipeline step. Combine with --no_render to perform format conversion (e.g., to GLB) without generating MP4 videos.

What hardware requirements exist for the CUDA voxelization extensions?

The voxelization stage requires an NVIDIA GPU with CUDA support to run voxelize_frame from demo_render/render_cuda_ext. The extensions perform frustum culling and dense voxel grid generation on the GPU. While CPU fallback is not available for these specific operations, the inference stages can run on any PyTorch-supported device, though GPU acceleration is strongly recommended for real-time processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →