How to Perform Offline Batch Rendering with LingBot-Map: A Complete Guide
LingBot-Map processes long video sequences or large image folders through a four-stage offline pipeline—data ingestion, model inference, prediction export, and headless rendering—without the memory constraints of the interactive viewer.
The Robbyant/lingbot-map repository provides a dedicated offline batch rendering pipeline designed for production-scale workloads. Unlike the interactive demo.py viewer, this headless system in demo_render/batch_demo.py handles arbitrarily long inputs through efficient caching, parallel I/O, and CUDA-accelerated voxelization.
Understanding the Offline Batch Rendering Pipeline
The offline pipeline orchestrates four distinct stages to transform raw media into annotated 3D fly-through videos. According to the source code in demo_render/batch_demo.py, the workflow handles scene discovery, frame extraction, neural inference, and final rendering through a unified command-line interface. Each stage is optimized for headless server deployment, enabling batch processing of multiple scenes without GPU display dependencies.
Stage 1: Data Ingestion and Preprocessing
The pipeline accepts two input modalities: video files or sorted image folders. All frames undergo canonical preprocessing before reaching the model.
Video Frame Extraction
When processing video inputs via --video_path, the system invokes extract_frames_from_video using OpenCV. Frames are extracted on-the-fly and optionally cached as PNG sequences to --save_frames_dir for reuse across multiple rendering passes. This prevents redundant decoding when iterating on visualization parameters.
Image Folder Processing
For image sequences, the list_image_paths utility discovers frames within --input_folder and applies temporal filtering through --first_k, --last_k, or stride parameters. This allows processing of specific sub-sequences from large photogrammetry datasets.
Canonical Preprocessing
All loaded frames pass through load_and_preprocess_images in lingbot_map/utils/load_fn.py. This function performs:
- Resizing to model input dimensions
- Canonical cropping to remove lens distortion artifacts
- Padding to multiples of the patch size for transformer compatibility
Stage 2: Model Inference Modes
After preprocessing, the system loads the LingBot-Map checkpoint via load_model and executes inference based on the --mode argument.
Streaming Inference with KV-Cache
Streaming mode (--mode streaming) maintains a persistent KV-cache across frames using model.inference_streaming. This approach minimizes redundant computation for continuous video sequences. Control the cache refresh rate with --keyframe_interval, which determines how frequently the cache resets to prevent drift in long sequences.
Windowed Inference for Long Sequences
Windowed mode (--mode windowed) splits sequences into overlapping chunks processed by model.inference_windowed. The overlap is controlled via --overlap_keyframes or --overlap_size, ensuring temporal consistency at segment boundaries. Use this mode for hour-long recordings that exceed GPU memory limits for full-sequence attention.
Stage 3: Prediction Export and Caching
Following inference, per-frame predictions are serialized as compressed .npz files via save_predictions_npz. This format enables fast parallel I/O during the rendering stage through load_predictions_from_npz. Enable this behavior with --save_predictions to decouple expensive inference from iterative visualization tuning.
Stage 4: Headless RGBD Rendering
The final stage invokes the rgbd-render pipeline through render_with_pipeline, operating entirely without display servers.
Voxelization and Camera Path Generation
A PipelineConfig object is instantiated from an optional YAML preset (--config) and overridden by CLI flags. The pipeline calls voxelize_frame from the CUDA extension in demo_render/render_cuda_ext to build dense 3D representations. Camera trajectories are generated via rgbd_render.camera.build_camera_path, creating smooth fly-throughs based on the predicted camera poses.
Overlay Composition and Video Encoding
The OfflinePipeline class composites optional visualizations including trajectory trails, camera frustums, and sky-mask visualizations. When --mask_sky is enabled, the system uses the ONNX segmentation model in lingbot_map/vis/sky_segmentation.py to exclude sky regions from the point cloud. Final output is encoded as MP4 video or exported as GLB assets when --save_glb is specified.
Practical Usage Examples
The following commands demonstrate typical workflows for the offline batch rendering pipeline.
Process multiple scene folders using streaming inference with sky masking:
python demo_render/batch_demo.py \
--input_folder /data/scenes \
--output_folder /data/outputs \
--model_path /path/to/lingbot-map-long.pt \
--mode streaming \
--keyframe_interval 2 \
--mask_sky \
--save_predictions
Extract frames from a single video using windowed inference:
python demo_render/batch_demo.py \
--video_path /data/video/indoor_travel.MP4 \
--output_folder /data/outputs \
--model_path /path/to/lingbot-map.pt \
--config demo_render/config/indoor.yaml \
--mode windowed \
--window_size 128 \
--overlap_keyframes 8 \
--keyframe_interval 13 \
--mask_sky \
--save_predictions \
--save_glb
Render from cached predictions without re-running inference:
python demo_render/batch_demo.py \
--load_predictions /data/outputs/indoor_travel.npz \
--output_folder /data/outputs \
--config demo_render/config/indoor.yaml \
--no_render
Key Source Files and Architecture
Understanding the following components is essential for customizing the offline batch rendering pipeline:
demo_render/batch_demo.py– Entry point handling argument parsing, scene discovery, and stage orchestration.lingbot_map/utils/load_fn.py– Containsload_and_preprocess_imagesfor canonical cropping and tensor preparation.lingbot_map/vis/sky_segmentation.py– ONNX-based sky segmentation used when--mask_skyis enabled.rgbd_render/pipeline/builder.pyandrgbd_render/pipeline/offline.py– Construct scene representations and execute headless rendering.demo_render/render_cuda_ext/– CUDA extensions includingvoxelize_frame.cufor accelerated voxelization and frustum culling.demo_render/config/*.yaml– Preset configurations (e.g.,indoor.yaml,outdoor_drive.yaml) defining default camera paths and rendering parameters.
Summary
- LingBot-Map provides a production-ready offline pipeline in
demo_render/batch_demo.pyfor processing long sequences without interactive display requirements. - The pipeline supports both streaming inference (with KV-cache management) and windowed inference (for memory-constrained long videos).
- Predictions are cached as
.npzfiles to enable decoupled rendering and avoid redundant model inference. - CUDA-accelerated voxelization and headless RGBD rendering generate MP4 fly-throughs and GLB assets from processed sequences.
- Sky masking, temporal filtering, and YAML-based configuration presets provide fine-grained control over the output visualization.
Frequently Asked Questions
What is the difference between streaming and windowed inference in LingBot-Map?
Streaming inference maintains a persistent KV-cache across frames using model.inference_streaming, making it efficient for continuous video where temporal consistency is maintained through cached attention states. Windowed inference uses model.inference_windowed to process sequences in overlapping chunks, which is necessary for very long videos that exceed GPU memory capacity but requires explicit overlap management via --overlap_keyframes.
How do I handle sky masking in offline batch rendering?
Enable the --mask_sky flag to activate the ONNX segmentation model defined in lingbot_map/vis/sky_segmentation.py. The system segments sky regions from each frame and excludes them from the point cloud generation, preventing distant sky pixels from creating artifacts in the 3D reconstruction. Mask files are cached to disk to avoid recomputing segmentation during rendering iterations.
Can I render videos from previously saved predictions without re-running inference?
Yes, use the --load_predictions argument to point to existing .npz files created by save_predictions_npz. When loading cached predictions, the pipeline skips the inference stage entirely and proceeds directly to the render_with_pipeline step. Combine with --no_render to perform format conversion (e.g., to GLB) without generating MP4 videos.
What hardware requirements exist for the CUDA voxelization extensions?
The voxelization stage requires an NVIDIA GPU with CUDA support to run voxelize_frame from demo_render/render_cuda_ext. The extensions perform frustum culling and dense voxel grid generation on the GPU. While CPU fallback is not available for these specific operations, the inference stages can run on any PyTorch-supported device, though GPU acceleration is strongly recommended for real-time processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →