How Sky Segmentation Works with ONNX Runtime in LingBot-Map
LingBot-Map performs sky segmentation by running a lightweight ONNX model on each input image and converting the raw network output into a binary mask that suppresses sky regions in confidence maps.
LingBot-Map is an open-source SLAM benchmark pipeline that evaluates visual mapping algorithms using automated preprocessing techniques. The repository implements sky segmentation with ONNX runtime to identify and exclude sky pixels from feature matching, preventing false correspondences in outdoor scenarios. This CPU-optimized approach processes images through a pre-trained segmentation model without requiring GPU acceleration, making it suitable for large-scale benchmark datasets.
The Sky Segmentation Pipeline
The implementation in lingbot_map/vis/sky_segmentation.py follows a modular pipeline that acquires the model, executes inference, and integrates results into downstream confidence calculations.
Model Acquisition and Caching
If the model file skyseg.onnx is not present locally, the system automatically downloads it from Hugging Face. The download_skyseg_model() function (lines 30‑58) handles the retrieval and storage of the binary weights. To avoid redundant processing across runs, generated masks are cached in a directory specified by <image_dir>_sky_masks or a user-defined sky_mask_dir. The _prepare_sky_mask_cache function (lines 35‑44) validates cache integrity using a version stamp (_SKYSEG_CACHE_VERSION), ensuring stale masks are regenerated when the implementation changes.
ONNX Runtime Session Initialization
The pipeline creates a single onnxruntime.InferenceSession instance per execution context within load_or_create_sky_masks (line 52). This session persists across the batch processing of multiple images, minimizing overhead from repeated model loading. The session operates on CPU by default, leveraging ONNX Runtime’s optimized execution providers for fast forward passes.
Input Preprocessing
Before inference, images are resized to the model’s expected input dimensions defined by _SKYSEG_INPUT_SIZE = (320, 320). The preprocessing logic inside run_skyseg (lines 54‑60) applies ImageNet mean and standard deviation normalization to match the training distribution of the segmentation network. This standardization ensures consistent output logits regardless of input image resolution or color space.
Inference and Post-Processing
The pre-processed tensor is fed to the session via session.run() (lines 62‑65), producing raw segmentation logits. These outputs undergo min‑max normalization to the range [0, 255] and are cast to uint8 (lines 66‑72). The _result_map_to_non_sky_conf function (lines 92‑95) inverts the mask (1 − mask) to generate a non‑sky confidence map ranging [0, 1], where values near 0 indicate sky regions and values near 1 indicate non‑sky regions.
Integration with Confidence Scores
The apply_sky_segmentation function (lines 73‑126) loads cached masks or generates them on demand, then applies a soft threshold defined by _SKYSEG_SOFT_THRESHOLD = 0.1 to create a binary mask. This mask is multiplied element‑wise into the per‑frame confidence tensor, effectively zeroing out or attenuating confidence values in sky regions. This suppression prevents SLAM frontends from attempting to track features on clouds or distant atmospheric haze.
Optional Visualization
For debugging purposes, the pipeline can generate side‑by-side visualization panels showing the original image, raw segmentation mask, and overlay composites. The _save_sky_mask_visualization function (lines 83‑108) writes these diagnostic images to sky_mask_visualization_dir when specified.
Practical Code Examples
Running Sky Segmentation on a Single Image
import cv2
import onnxruntime
import os
from lingbot_map.vis.sky_segmentation import segment_sky, download_skyseg_model
# Ensure the model is present locally
model_path = "skyseg.onnx"
if not os.path.exists(model_path):
download_skyseg_model(model_path)
# Create an ONNX Runtime session
sess = onnxruntime.InferenceSession(model_path)
# Run segmentation
mask = segment_sky("path/to/image.jpg", sess, output_path="mask.png")
# `mask` is a float32 array in [0, 1] where 0 = sky, 1 = non-sky
Integrating Sky Masks into Confidence Volumes
import numpy as np
from lingbot_map.vis.sky_segmentation import apply_sky_segmentation
# Load confidence tensor with shape (num_frames, H, W)
conf = np.load("confidence.npy")
# Apply sky masks to suppress sky regions
conf_no_sky = apply_sky_segmentation(
conf,
image_folder="path/to/images",
skyseg_model_path="skyseg.onnx",
sky_mask_dir="cache_sky_masks",
sky_mask_visualization_dir="visualizations"
)
np.save("confidence_no_sky.npy", conf_no_sky)
Batch Processing from NumPy Arrays
import numpy as np
from lingbot_map.vis.sky_segmentation import load_or_create_sky_masks
# Load image stack with shape (S, H, W, 3) where S is number of frames
imgs = np.load("frames.npy")
# Generate or load cached masks
sky_masks = load_or_create_sky_masks(
images=imgs,
skyseg_model_path="skyseg.onnx",
target_shape=(imgs.shape[1], imgs.shape[2])
)
# Result is a (S, H, W) float mask array
Summary
- LingBot-Map implements sky segmentation using a lightweight ONNX model executed via
onnxruntime.InferenceSessioninlingbot_map/vis/sky_segmentation.py. - The pipeline resizes inputs to 320×320 pixels, applies ImageNet normalization, and runs inference on CPU for hardware flexibility.
- Raw outputs are inverted to create non-sky confidence maps, which are cached with version control to prevent stale data.
- The
apply_sky_segmentationfunction applies a 0.1 soft threshold to generate binary masks that suppress sky regions in SLAM confidence tensors. - Optional visualization utilities generate side-by-side diagnostic panels for debugging segmentation quality.
Frequently Asked Questions
What model format does LingBot-Map use for sky segmentation?
LingBot-Map uses an ONNX model stored as skyseg.onnx. The implementation relies on the onnxruntime Python package to load and execute the model, enabling cross-platform deployment without framework-specific dependencies like PyTorch or TensorFlow.
How does the caching mechanism prevent stale segmentation masks?
The system uses _SKYSEG_CACHE_VERSION as a version stamp within _prepare_sky_mask_cache. If the implementation version changes or the cached mask format becomes incompatible, the version mismatch triggers automatic regeneration of masks, ensuring consistency with the current codebase.
Can the sky segmentation pipeline run on machines without a GPU?
Yes. The ONNX Runtime session is configured for CPU inference by default, making the pipeline suitable for headless servers and large-scale batch processing without GPU acceleration. The lightweight model architecture (320×320 input) ensures fast inference even on standard CPU hardware.
How is the sky mask applied to confidence scores during SLAM evaluation?
The apply_sky_segmentation function loads the generated masks and thresholds them using _SKYSEG_SOFT_THRESHOLD = 0.1 to create a binary mask. This mask is multiplied into the per-frame confidence tensor, effectively setting confidence values to zero (or near-zero) in sky regions to prevent feature tracking on unstable atmospheric content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →