# How Sky Segmentation Works with ONNX Runtime in LingBot-Map

> Discover how LingBot-Map achieves sky segmentation using ONNX Runtime. This article explains the process of running an ONNX model and creating binary masks to effectively suppress sky regions.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-25

---

**LingBot-Map performs sky segmentation by running a lightweight ONNX model on each input image and converting the raw network output into a binary mask that suppresses sky regions in confidence maps.**

LingBot-Map is an open-source SLAM benchmark pipeline that evaluates visual mapping algorithms using automated preprocessing techniques. The repository implements **sky segmentation with ONNX runtime** to identify and exclude sky pixels from feature matching, preventing false correspondences in outdoor scenarios. This CPU-optimized approach processes images through a pre-trained segmentation model without requiring GPU acceleration, making it suitable for large-scale benchmark datasets.

## The Sky Segmentation Pipeline

The implementation in [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py) follows a modular pipeline that acquires the model, executes inference, and integrates results into downstream confidence calculations.

### Model Acquisition and Caching

If the model file `skyseg.onnx` is not present locally, the system automatically downloads it from Hugging Face. The `download_skyseg_model()` function (lines 30‑58) handles the retrieval and storage of the binary weights. To avoid redundant processing across runs, generated masks are cached in a directory specified by `<image_dir>_sky_masks` or a user-defined `sky_mask_dir`. The `_prepare_sky_mask_cache` function (lines 35‑44) validates cache integrity using a version stamp (`_SKYSEG_CACHE_VERSION`), ensuring stale masks are regenerated when the implementation changes.

### ONNX Runtime Session Initialization

The pipeline creates a single `onnxruntime.InferenceSession` instance per execution context within `load_or_create_sky_masks` (line 52). This session persists across the batch processing of multiple images, minimizing overhead from repeated model loading. The session operates on CPU by default, leveraging ONNX Runtime’s optimized execution providers for fast forward passes.

### Input Preprocessing

Before inference, images are resized to the model’s expected input dimensions defined by `_SKYSEG_INPUT_SIZE = (320, 320)`. The preprocessing logic inside `run_skyseg` (lines 54‑60) applies ImageNet mean and standard deviation normalization to match the training distribution of the segmentation network. This standardization ensures consistent output logits regardless of input image resolution or color space.

### Inference and Post-Processing

The pre-processed tensor is fed to the session via `session.run()` (lines 62‑65), producing raw segmentation logits. These outputs undergo min‑max normalization to the range `[0, 255]` and are cast to `uint8` (lines 66‑72). The `_result_map_to_non_sky_conf` function (lines 92‑95) inverts the mask (`1 − mask`) to generate a non‑sky confidence map ranging `[0, 1]`, where values near 0 indicate sky regions and values near 1 indicate non‑sky regions.

### Integration with Confidence Scores

The `apply_sky_segmentation` function (lines 73‑126) loads cached masks or generates them on demand, then applies a soft threshold defined by `_SKYSEG_SOFT_THRESHOLD = 0.1` to create a binary mask. This mask is multiplied element‑wise into the per‑frame confidence tensor, effectively zeroing out or attenuating confidence values in sky regions. This suppression prevents SLAM frontends from attempting to track features on clouds or distant atmospheric haze.

### Optional Visualization

For debugging purposes, the pipeline can generate side‑by-side visualization panels showing the original image, raw segmentation mask, and overlay composites. The `_save_sky_mask_visualization` function (lines 83‑108) writes these diagnostic images to `sky_mask_visualization_dir` when specified.

## Practical Code Examples

### Running Sky Segmentation on a Single Image

```python
import cv2
import onnxruntime
import os
from lingbot_map.vis.sky_segmentation import segment_sky, download_skyseg_model

# Ensure the model is present locally

model_path = "skyseg.onnx"
if not os.path.exists(model_path):
    download_skyseg_model(model_path)

# Create an ONNX Runtime session

sess = onnxruntime.InferenceSession(model_path)

# Run segmentation

mask = segment_sky("path/to/image.jpg", sess, output_path="mask.png")

# `mask` is a float32 array in [0, 1] where 0 = sky, 1 = non-sky

```

### Integrating Sky Masks into Confidence Volumes

```python
import numpy as np
from lingbot_map.vis.sky_segmentation import apply_sky_segmentation

# Load confidence tensor with shape (num_frames, H, W)

conf = np.load("confidence.npy")

# Apply sky masks to suppress sky regions

conf_no_sky = apply_sky_segmentation(
    conf,
    image_folder="path/to/images",
    skyseg_model_path="skyseg.onnx",
    sky_mask_dir="cache_sky_masks",
    sky_mask_visualization_dir="visualizations"
)

np.save("confidence_no_sky.npy", conf_no_sky)

```

### Batch Processing from NumPy Arrays

```python
import numpy as np
from lingbot_map.vis.sky_segmentation import load_or_create_sky_masks

# Load image stack with shape (S, H, W, 3) where S is number of frames

imgs = np.load("frames.npy")

# Generate or load cached masks

sky_masks = load_or_create_sky_masks(
    images=imgs,
    skyseg_model_path="skyseg.onnx",
    target_shape=(imgs.shape[1], imgs.shape[2])
)

# Result is a (S, H, W) float mask array

```

## Summary

- **LingBot-Map** implements sky segmentation using a lightweight ONNX model executed via `onnxruntime.InferenceSession` in [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py).
- The pipeline resizes inputs to **320×320 pixels**, applies ImageNet normalization, and runs inference on CPU for hardware flexibility.
- Raw outputs are inverted to create **non-sky confidence maps**, which are cached with version control to prevent stale data.
- The `apply_sky_segmentation` function applies a **0.1 soft threshold** to generate binary masks that suppress sky regions in SLAM confidence tensors.
- Optional visualization utilities generate side-by-side diagnostic panels for debugging segmentation quality.

## Frequently Asked Questions

### What model format does LingBot-Map use for sky segmentation?

LingBot-Map uses an **ONNX model** stored as `skyseg.onnx`. The implementation relies on the `onnxruntime` Python package to load and execute the model, enabling cross-platform deployment without framework-specific dependencies like PyTorch or TensorFlow.

### How does the caching mechanism prevent stale segmentation masks?

The system uses `_SKYSEG_CACHE_VERSION` as a version stamp within `_prepare_sky_mask_cache`. If the implementation version changes or the cached mask format becomes incompatible, the version mismatch triggers automatic regeneration of masks, ensuring consistency with the current codebase.

### Can the sky segmentation pipeline run on machines without a GPU?

Yes. The **ONNX Runtime** session is configured for CPU inference by default, making the pipeline suitable for headless servers and large-scale batch processing without GPU acceleration. The lightweight model architecture (320×320 input) ensures fast inference even on standard CPU hardware.

### How is the sky mask applied to confidence scores during SLAM evaluation?

The `apply_sky_segmentation` function loads the generated masks and thresholds them using `_SKYSEG_SOFT_THRESHOLD = 0.1` to create a binary mask. This mask is multiplied into the per-frame confidence tensor, effectively setting confidence values to zero (or near-zero) in sky regions to prevent feature tracking on unstable atmospheric content.