# How to Export and Reuse Cached Sky Masks for Batch Processing in Ling-Bot-Map

> Export and reuse cached sky masks in Ling-Bot-Map for faster batch processing. Save time by bypassing ONNX inference and directly applying saved masks.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-29

---

**You export cached sky masks by running `load_or_create_sky_masks` with a `sky_mask_dir` parameter, which writes PNG files to disk that can be reused in subsequent batch runs by passing the same directory path to `apply_sky_segmentation` or `ViserWrapper`, bypassing ONNX inference entirely.**

The Ling-Bot-Map pipeline generates per-frame **sky masks** to filter sky points before 3D reconstruction, but recomputing these via ONNX inference for every batch wastes significant compute. According to the Robbyant/lingbot-map source code, the [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py) module implements a robust **cache-first strategy** that persists masks as portable PNG files with automatic version validation.

## Understanding the Sky Mask Cache Architecture

The caching system operates on a simple principle: generate once, validate always, reuse forever. When you supply a `sky_mask_dir` argument, the pipeline invokes `_prepare_sky_mask_cache` to create the directory structure and stamp it with a hidden version file named `.skyseg_cache_version` (lines 28-29 and 35-44 in [`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py)). This version stamp ensures that cached masks are only reused when they match the current model and preprocessing pipeline version.

The actual mask data is stored as standard **uint8 PNG images** produced by `_mask_to_uint8`, making the cache portable and inspectable with standard tools. Before invoking the **ONNX segmentation model**, `load_or_create_sky_masks` performs a cache lookup via `cv2.imread` (lines 80-88); if a valid mask file exists and matches the expected dimensions, the function returns the cached data immediately, skipping neural network inference entirely.

## Step-by-Step: Exporting and Reusing Sky Masks

### Step 1 – Initialize the Cache Directory

The internal function `_prepare_sky_mask_cache` automatically handles directory creation and version stamping when you first specify a `sky_mask_dir`. It writes the `_SKYSEG_CACHE_VERSION` constant to a hidden file inside the directory, establishing a contract that prevents the pipeline from accidentally loading incompatible masks after code updates.

### Step 2 – Generate and Export Masks

Call `load_or_create_sky_masks` with your image folder and a target cache directory. The function converts boolean sky masks to uint8 PNG format via `_mask_to_uint8` and writes them to `sky_mask_dir`. This constitutes your export operation—the resulting folder contains everything needed for future batch runs.

### Step 3 – Validate Cache Version

Each time the cache is accessed, the pipeline checks for the `.skyseg_cache_version` file. If a version mismatch occurs (indicating a model or preprocessing change), the cache is invalidated and fresh masks are generated. This automatic validation protects processing integrity without manual intervention.

### Step 4 – Reuse in Batch Processing

For subsequent runs, pass the same `sky_mask_dir` to `apply_sky_segmentation` or the `ViserWrapper` class. The function checks for existing PNG files before running the ONNX model (lines 80-88), loading cached masks directly when available. This reduces batch processing time from minutes to seconds while maintaining identical results.

## Practical Implementation Examples

The following examples demonstrate the complete workflow from initial mask generation to reuse in high-level visualization wrappers.

Generate masks once and export the cache:

```python
from lingbot_map.vis.sky_segmentation import load_or_create_sky_masks

# Create cache directory with version stamp

sky_mask_dir = "./data/images_sky_masks"
load_or_create_sky_masks(
    image_folder="./data/images",
    sky_mask_dir=sky_mask_dir,          # Cache written here

    sky_mask_visualization_dir=None,   # Optional: set path for debug visuals

)
print(f"Cache exported to {sky_mask_dir}")

```

Reuse cached masks in batch processing:

```python
import numpy as np
from lingbot_map.vis.sky_segmentation import apply_sky_segmentation

# Confidence tensor from previous reconstruction step

conf = np.random.rand(100, 480, 640).astype(np.float32)

# Reuse existing masks; ONNX model is skipped if cache is valid

conf_no_sky = apply_sky_segmentation(
    conf,
    image_folder="./data/images",          # Required for image ordering

    sky_mask_dir="./data/images_sky_masks", # Reuse cached PNGs

)

```

Integrate with `ViserWrapper` for interactive visualization:

```python
from lingbot_map.vis.viser_wrapper import ViserWrapper

viewer = ViserWrapper(
    conf,                                   # Confidence volume

    image_folder="./data/images",
    mask_sky=True,                         # Enable masking

    sky_mask_dir="./data/images_sky_masks", # Loads cached masks automatically

)
viewer.run()

```

## Integration Points for Large-Scale Workflows

For benchmark suites and large-scale experiments, the caching logic is also available in [`benchmark/benchmark/utils/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/utils/sky_segmentation.py), which mirrors the core implementation. You can inject cached masks at the entry point level in [`benchmark/benchmark/method/base.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/benchmark/method/base.py), ensuring consistent sky filtering across entire evaluation campaigns without repeated model inference.

This architecture decouples the expensive segmentation step from the reconstruction pipeline, allowing you to pre-compute masks on a GPU-enabled machine and transfer the lightweight PNG cache to CPU-only batch processing nodes.

## Summary

- **Cache initialization** happens automatically via `_prepare_sky_mask_cache` when you specify `sky_mask_dir`, creating a version-stamped directory.
- **Mask export** produces standard PNG files via `_mask_to_uint8`, making the cache portable and archivable.
- **Version validation** via `.skyseg_cache_version` prevents stale mask reuse when the model or preprocessing changes.
- **Batch reuse** requires only passing the same directory to `apply_sky_segmentation` or `ViserWrapper`, bypassing ONNX inference entirely.
- **Integration** is supported in both the core visualization module and benchmark utilities for scalable workflows.

## Frequently Asked Questions

### What file format are cached sky masks stored in?

The pipeline stores cached masks as standard **PNG images** with uint8 encoding. The `_mask_to_uint8` function converts boolean segmentation masks to this format before writing to disk, allowing you to inspect masks with any image viewer or OpenCV-compatible tool.

### How does the cache prevent stale masks after model updates?

The `_prepare_sky_mask_cache` function writes a `_SKYSEG_CACHE_VERSION` stamp to a hidden `.skyseg_cache_version` file inside the cache directory (lines 28-29). If the version constant in the code does not match the stamp in the directory, the cache is invalidated and fresh masks are generated automatically.

### Can I share cached masks between different machines?

Yes. Because the cache consists of ordinary PNG files and a version metadata file, you can archive the `sky_mask_dir` directory, transfer it to another machine, and point any Ling-Bot-Map installation to that path. Ensure the destination environment uses the same code version to satisfy the cache validation check.

### Is the ONNX model executed if valid cached masks exist?

No. The `load_or_create_sky_masks` function checks for existing files using `cv2.imread` before any model inference (lines 80-88). If valid masks are found and pass dimension checks, the ONNX runtime is bypassed entirely, significantly reducing processing time for batch operations.