# Interoperability Between sdata and HDF5/NetCDF Formats: Usage and Conversion Guide

> Master interoperability between sdata and HDF5/NetCDF formats. Learn efficient usage and conversion techniques with this essential guide from lepy/sdata.

- Repository: [lepy/sdata](https://github.com/lepy/sdata)
- Tags: how-to-guide
- Published: 2026-03-06

---

**The sdata library provides native HDF5 serialization via `Data.to_hdf5()` and `Data.from_hdf5()` methods, while NetCDF interoperability requires converting DataFrames through xarray as a bridge format.**

The **sdata** library (lepy/sdata) offers a pandas-based interface for scientific data management, with native HDF5 support and documented pathways for NetCDF conversion. Understanding these interoperability mechanisms enables seamless data exchange between sdata's metadata-rich objects and standard scientific storage formats.

## Native HDF5 Support in sdata

### Core Implementation Details

The `Data` class in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py) implements the primary interoperability methods `to_hdf5()` and `from_hdf5()`. These methods wrap **pandas.HDFStore** (which relies on PyTables) for low-level binary storage, while the helper class in [`sdata/iolib/hdf.py`](https://github.com/lepy/sdata/blob/main/sdata/iolib/hdf.py) manages the HDF5 I/O operations. Metadata tables are serialized alongside the main data payload, creating self-describing files that preserve both values and descriptive attributes.

### Reading and Writing HDF5 Files

Creating and saving an sdata object to HDF5 requires only a few lines of code:

```python
import sdata
import pandas as pd

# Create an sdata.Data instance

df = pd.DataFrame({
    "time": pd.date_range("2023-01-01", periods=5, freq="D"),
    "temperature": [22.5, 23.0, 22.8, 23.2, 22.9],
})
data = sdata.Data(name="weather", table=df, comment="Demo weather data")

# Write to HDF5

h5_path = "weather.h5"
data.to_hdf5(h5_path)  # stores data + metadata in a Pandas HDFStore

```

Loading the file restores the complete object, including the metadata DataFrame:

```python
loaded = sdata.Data.from_hdf5("weather.h5")
print(loaded.table)       # pandas DataFrame

print(loaded.metadata.df) # metadata table

```

## Converting sdata to NetCDF Format

### Current NetCDF Support Status

According to the project README (line 26), NetCDF is explicitly listed as a **future** supported format. The codebase currently contains no dedicated NetCDF reader or writer, meaning all conversion operations must be handled externally using xarray as an interoperability bridge.

### Exporting sdata Objects to NetCDF

Convert sdata objects to NetCDF by first extracting the pandas DataFrame, then using xarray to write the file. Map sdata metadata to NetCDF global attributes to preserve descriptive information:

```python
import xarray as xr

# Assume `data` is an sdata.Data instance

df = data.table  # pandas DataFrame

ds = xr.Dataset.from_dataframe(df)

# Attach sdata metadata as NetCDF attributes

for _, row in data.metadata.df.iterrows():
    ds.attrs[row["key"]] = row["value"]

nc_path = "weather.nc"
ds.to_netcdf(nc_path, engine="netcdf4")  # writes NetCDF-4 file

```

### Importing NetCDF Data into sdata

To import NetCDF files, load them with xarray, convert to a pandas DataFrame, and reconstruct the sdata object:

```python
import xarray as xr
import sdata

nc_path = "weather.nc"
ds = xr.open_dataset(nc_path, engine="netcdf4")

# Convert back to pandas DataFrame

df = ds.to_dataframe().reset_index()

# Re-create sdata.Data object

data_from_nc = sdata.Data(name="weather", table=df, comment="Loaded from NetCDF")

# Restore metadata from NetCDF attributes

for key, val in ds.attrs.items():
    data_from_nc.metadata.add(key, val)

```

## Complete Round-Trip Workflow

The following script demonstrates a full interoperability cycle: sdata → HDF5 → sdata → NetCDF → sdata:

```python
import sdata, pandas as pd, xarray as xr

# 1. Create sdata object

df = pd.DataFrame({"x": [0, 1, 2], "y": [10, 20, 30]})
sd = sdata.Data(name="sample", table=df)
sd.metadata.add("author", "Alice", unit="", description="creator")

# 2. Export to HDF5

h5_file = "sample.h5"
sd.to_hdf5(h5_file)

# 3. Load from HDF5

sd_h5 = sdata.Data.from_hdf5(h5_file)

# 4. Convert to NetCDF

ds = xr.Dataset.from_dataframe(sd_h5.table)
for _, row in sd_h5.metadata.df.iterrows():
    ds.attrs[row["key"]] = row["value"]
nc_file = "sample.nc"
ds.to_netcdf(nc_file)

# 5. Load NetCDF and rebuild sdata

ds_loaded = xr.open_dataset(nc_file)
df_loaded = ds_loaded.to_dataframe().reset_index()
sd_nc = sdata.Data(name="sample", table=df_loaded)
for key, val in ds_loaded.attrs.items():
    sd_nc.metadata.add(key, val)

print(sd_nc.table)
print(sd_nc.metadata.df)

```

## Key Source Files and Implementation

- **[`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py)** — Contains the `Data` class with `to_hdf5()` and `from_hdf5()` methods
- **[`sdata/iolib/hdf.py`](https://github.com/lepy/sdata/blob/main/sdata/iolib/hdf.py)** — Wraps pandas HDFStore (PyTables) for low-level read/write operations
- **[`sdata/iolib/vault.py`](https://github.com/lepy/sdata/blob/main/sdata/iolib/vault.py)** — Manages optional vault index persistence (not preserved in NetCDF conversion)
- **`docs/source/usage/usage.rst`** — Documents usage patterns and design notes
- **[`README.md`](https://github.com/lepy/sdata/blob/main/README.md)** (line 26) — Lists NetCDF as a future supported format

## Summary

- **Native HDF5 support**: The `Data` class provides first-class HDF5 serialization using pandas HDFStore, preserving both data tables and metadata in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py).
- **NetCDF bridge required**: NetCDF is not natively implemented; conversion requires xarray's `from_dataframe()` and `to_netcdf()` methods.
- **Metadata mapping**: HDF5 preserves the full metadata structure automatically, while NetCDF conversion requires manual mapping of attributes to `ds.attrs`.
- **Vault limitations**: The internal vault index from [`sdata/iolib/vault.py`](https://github.com/lepy/sdata/blob/main/sdata/iolib/vault.py) is lost during NetCDF conversion; only the primary data table and explicit metadata survive the round-trip.

## Frequently Asked Questions

### Does sdata support NetCDF files natively?

No, sdata does not currently implement native NetCDF read or write methods. According to the repository's README.md (line 26), NetCDF is designated as a future supported format. Until native support is added, users must convert between sdata and NetCDF using xarray as an intermediate format.

### How does sdata store metadata when writing to HDF5?

The `to_hdf5()` method in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py) serializes the DataFrame alongside a metadata attribute table within the pandas HDFStore structure. This creates a self-describing HDF5 file where both the numerical data and descriptive metadata are preserved and automatically restored when using `from_hdf5()`.

### Is the sdata vault index preserved in NetCDF conversions?

No, the internal vault index managed by [`sdata/iolib/vault.py`](https://github.com/lepy/sdata/blob/main/sdata/iolib/vault.py) is not preserved when converting to NetCDF. Only the primary data table and explicit metadata attributes are transferred. Users requiring vault functionality should remain in the native HDF5 ecosystem or implement custom indexing for NetCDF workflows.

### What dependencies are required for HDF5 and NetCDF interoperability?

For HDF5 support, sdata requires **pandas** with **PyTables** (the `tables` package) installed. For NetCDF conversion, you need **xarray** and a NetCDF backend such as **netcdf4** or **h5netcdf**. These are not hard dependencies of sdata and must be installed separately when performing format conversions.