Interoperability Between sdata and HDF5/NetCDF Formats: Usage and Conversion Guide
The sdata library provides native HDF5 serialization via Data.to_hdf5() and Data.from_hdf5() methods, while NetCDF interoperability requires converting DataFrames through xarray as a bridge format.
The sdata library (lepy/sdata) offers a pandas-based interface for scientific data management, with native HDF5 support and documented pathways for NetCDF conversion. Understanding these interoperability mechanisms enables seamless data exchange between sdata's metadata-rich objects and standard scientific storage formats.
Native HDF5 Support in sdata
Core Implementation Details
The Data class in sdata/data.py implements the primary interoperability methods to_hdf5() and from_hdf5(). These methods wrap pandas.HDFStore (which relies on PyTables) for low-level binary storage, while the helper class in sdata/iolib/hdf.py manages the HDF5 I/O operations. Metadata tables are serialized alongside the main data payload, creating self-describing files that preserve both values and descriptive attributes.
Reading and Writing HDF5 Files
Creating and saving an sdata object to HDF5 requires only a few lines of code:
import sdata
import pandas as pd
# Create an sdata.Data instance
df = pd.DataFrame({
"time": pd.date_range("2023-01-01", periods=5, freq="D"),
"temperature": [22.5, 23.0, 22.8, 23.2, 22.9],
})
data = sdata.Data(name="weather", table=df, comment="Demo weather data")
# Write to HDF5
h5_path = "weather.h5"
data.to_hdf5(h5_path) # stores data + metadata in a Pandas HDFStore
Loading the file restores the complete object, including the metadata DataFrame:
loaded = sdata.Data.from_hdf5("weather.h5")
print(loaded.table) # pandas DataFrame
print(loaded.metadata.df) # metadata table
Converting sdata to NetCDF Format
Current NetCDF Support Status
According to the project README (line 26), NetCDF is explicitly listed as a future supported format. The codebase currently contains no dedicated NetCDF reader or writer, meaning all conversion operations must be handled externally using xarray as an interoperability bridge.
Exporting sdata Objects to NetCDF
Convert sdata objects to NetCDF by first extracting the pandas DataFrame, then using xarray to write the file. Map sdata metadata to NetCDF global attributes to preserve descriptive information:
import xarray as xr
# Assume `data` is an sdata.Data instance
df = data.table # pandas DataFrame
ds = xr.Dataset.from_dataframe(df)
# Attach sdata metadata as NetCDF attributes
for _, row in data.metadata.df.iterrows():
ds.attrs[row["key"]] = row["value"]
nc_path = "weather.nc"
ds.to_netcdf(nc_path, engine="netcdf4") # writes NetCDF-4 file
Importing NetCDF Data into sdata
To import NetCDF files, load them with xarray, convert to a pandas DataFrame, and reconstruct the sdata object:
import xarray as xr
import sdata
nc_path = "weather.nc"
ds = xr.open_dataset(nc_path, engine="netcdf4")
# Convert back to pandas DataFrame
df = ds.to_dataframe().reset_index()
# Re-create sdata.Data object
data_from_nc = sdata.Data(name="weather", table=df, comment="Loaded from NetCDF")
# Restore metadata from NetCDF attributes
for key, val in ds.attrs.items():
data_from_nc.metadata.add(key, val)
Complete Round-Trip Workflow
The following script demonstrates a full interoperability cycle: sdata → HDF5 → sdata → NetCDF → sdata:
import sdata, pandas as pd, xarray as xr
# 1. Create sdata object
df = pd.DataFrame({"x": [0, 1, 2], "y": [10, 20, 30]})
sd = sdata.Data(name="sample", table=df)
sd.metadata.add("author", "Alice", unit="", description="creator")
# 2. Export to HDF5
h5_file = "sample.h5"
sd.to_hdf5(h5_file)
# 3. Load from HDF5
sd_h5 = sdata.Data.from_hdf5(h5_file)
# 4. Convert to NetCDF
ds = xr.Dataset.from_dataframe(sd_h5.table)
for _, row in sd_h5.metadata.df.iterrows():
ds.attrs[row["key"]] = row["value"]
nc_file = "sample.nc"
ds.to_netcdf(nc_file)
# 5. Load NetCDF and rebuild sdata
ds_loaded = xr.open_dataset(nc_file)
df_loaded = ds_loaded.to_dataframe().reset_index()
sd_nc = sdata.Data(name="sample", table=df_loaded)
for key, val in ds_loaded.attrs.items():
sd_nc.metadata.add(key, val)
print(sd_nc.table)
print(sd_nc.metadata.df)
Key Source Files and Implementation
sdata/data.py— Contains theDataclass withto_hdf5()andfrom_hdf5()methodssdata/iolib/hdf.py— Wraps pandas HDFStore (PyTables) for low-level read/write operationssdata/iolib/vault.py— Manages optional vault index persistence (not preserved in NetCDF conversion)docs/source/usage/usage.rst— Documents usage patterns and design notesREADME.md(line 26) — Lists NetCDF as a future supported format
Summary
- Native HDF5 support: The
Dataclass provides first-class HDF5 serialization using pandas HDFStore, preserving both data tables and metadata insdata/data.py. - NetCDF bridge required: NetCDF is not natively implemented; conversion requires xarray's
from_dataframe()andto_netcdf()methods. - Metadata mapping: HDF5 preserves the full metadata structure automatically, while NetCDF conversion requires manual mapping of attributes to
ds.attrs. - Vault limitations: The internal vault index from
sdata/iolib/vault.pyis lost during NetCDF conversion; only the primary data table and explicit metadata survive the round-trip.
Frequently Asked Questions
Does sdata support NetCDF files natively?
No, sdata does not currently implement native NetCDF read or write methods. According to the repository's README.md (line 26), NetCDF is designated as a future supported format. Until native support is added, users must convert between sdata and NetCDF using xarray as an intermediate format.
How does sdata store metadata when writing to HDF5?
The to_hdf5() method in sdata/data.py serializes the DataFrame alongside a metadata attribute table within the pandas HDFStore structure. This creates a self-describing HDF5 file where both the numerical data and descriptive metadata are preserved and automatically restored when using from_hdf5().
Is the sdata vault index preserved in NetCDF conversions?
No, the internal vault index managed by sdata/iolib/vault.py is not preserved when converting to NetCDF. Only the primary data table and explicit metadata attributes are transferred. Users requiring vault functionality should remain in the native HDF5 ecosystem or implement custom indexing for NetCDF workflows.
What dependencies are required for HDF5 and NetCDF interoperability?
For HDF5 support, sdata requires pandas with PyTables (the tables package) installed. For NetCDF conversion, you need xarray and a NetCDF backend such as netcdf4 or h5netcdf. These are not hard dependencies of sdata and must be installed separately when performing format conversions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →