Integrating sdata with pandas DataFrames: Best Practices for Scientific Data Manipulation
Integrate sdata with pandas DataFrames by wrapping your data in the sdata.sclass.dataframe.DataFrame class to attach persistent column-level metadata, unique identifiers, and scientific provenance that survives serialization across Parquet, JSON, and Excel formats.
The lepy/sdata repository extends pandas with a lightweight metadata layer designed specifically for reproducible scientific computing. Integrating sdata with pandas DataFrames allows researchers to embed units, descriptions, and persistent object identities directly into their data structures without breaking compatibility with the pandas ecosystem.
Understanding the sdata Architecture
The core integration lives in sdata/sclass/dataframe.py, where the DataFrame class inherits from sdata.base.Base. Unlike standard pandas objects, every sdata DataFrame carries a SUID (short name stored in _sdata_sname) and a UUID (_sdata_uuid), ensuring reproducible references across experiments and file formats.
Column-level metadata is handled through the Metadata class defined in sdata/metadata.py. This specialized dictionary stores Attribute objects containing labels, units, and descriptions for each column, accessed via the column_metadata property. While pandas offers DataFrame.attrs, sdata enforces structured validation and separates metadata from the raw data array, preventing pollution of the numerical contents.
Creating DataFrames with Scientific Metadata
Begin by loading your raw data into a standard pandas DataFrame, then wrap it with sdata to inject scientific context. The constructor validates that every column in your data has a corresponding entry in the metadata dictionary.
import pandas as pd
from sdata.sclass.dataframe import DataFrame
# Load raw experimental data
raw_df = pd.read_csv("experiment.csv")
# Define column-level metadata
col_meta = {
"force": {"label": "Force", "unit": "N"},
"displacement": {"label": "Displacement", "unit": "mm"},
}
# Wrap with sdata
sdf = DataFrame(
df=raw_df,
column_metadata=col_meta,
name="TensileTest",
description="Pull-test of alloy X at room temperature"
)
After creation, manipulate the underlying data through sdf.df using standard pandas operations. The metadata remains attached and accessible via sdf.column_metadata or the convenience property sdf.cmdf, which returns the metadata as a DataFrame for inspection.
Serializing and Persisting Data
sdata implements format-aware serialization that preserves metadata where pandas alone would lose it. The to_parquet() method in sdata/sclass/dataframe.py writes binary Parquet files while embedding the full metadata hierarchy into df.attrs["_sdata"].
# Export to Parquet with embedded metadata
parquet_path = sdf.to_parquet(path="/data/parquet")
# Export to JSON with base64-encoded parquet bytes for portability
json_str = sdf.to_json()
# Export to Excel with metadata in a hidden sheet
sdf.to_xlsx("/data/tensile_test.xlsx")
Restore objects using the corresponding classmethods, which automatically reconstruct the UUID, SUID, and column metadata:
# Full round-trip preservation
loaded = DataFrame.from_parquet(parquet_path)
assert loaded.uuid == sdf.uuid
assert loaded.column_metadata.to_dict() == sdf.column_metadata.to_dict()
The metadata field _sdata_version is set automatically during serialization, enabling version checking when loading legacy archives.
Managing Multiple Experiments with DataFrameGroup
For batch experiments or time-series collections, use DataFrameGroup from sdata/sclass/dataframegroup.py. This container aggregates multiple DataFrames, serializing each entry as base64-encoded parquet while preserving per-frame column metadata.
from sdata.sclass.dataframegroup import DataFrameGroup
import glob
group = DataFrameGroup(name="BatchTensileTests")
for i, fpath in enumerate(sorted(glob.glob("batch/*.csv"))):
df = pd.read_csv(fpath)
meta = {
"stress": {"label": "Engineering Stress", "unit": "MPa"},
"strain": {"label": "Engineering Strain", "unit": "-"},
}
key = f"specimen_{i}"
group.add_dataframe(key=key, df=df, column_metadata=meta)
# Serialize entire collection
group_dict = group.to_dict()
This pattern supports workflows where each measurement requires its own metadata context, such as varying test conditions across specimens.
Downstream Integration and Provenance
Pass sdata objects to downstream libraries that expect plain pandas DataFrames using the to_dataframe() method. This returns a copy of the internal DataFrame with a special !sdata attribute containing the full metadata and description dictionary.
# Convert for scikit-learn or visualization libraries
plain_df = sdf.to_dataframe()
# Access provenance in downstream code
sdata_meta = plain_df.attrs["!sdata"]
print(sdata_meta["!sdata_name"]["value"]) # "TensileTest"
print(sdata_meta["!sdata_description"]["value"]) # "Pull-test of alloy X"
This mechanism ensures that scientific context travels with the data even when crossing API boundaries into tools that lack native sdata support.
Best Practices for Scientific Workflows
Unit Conversions: Perform calculations directly on sdf.df columns, but keep the original unit in column_metadata. This makes conversions explicit while preserving the source unit for audit trails.
Column Renaming: After modifying column names in the underlying DataFrame, call Data.set_columnnames_from_metadata() (available in sdata/data.py) to synchronize the !sdata_column_* attributes.
Large Datasets: Prefer to_parquet() for memory-efficient streaming. Unlike JSON serialization, this method does not require loading the entire object into memory as a Python dictionary.
Concatenation: When combining multiple sdata DataFrames manually, merge their column_metadata dictionaries explicitly to avoid key collisions. The library does not automatically reconcile metadata during concatenation operations.
Metadata Validation: The constructor enforces that every DataFrame column has a metadata entry. Use this validation to catch schema drift early in data pipelines.
Summary
- Object Identity: Every sdata DataFrame carries a UUID and SUID stored in
_sdata_uuidand_sdata_sname, enabling reproducible references across file formats. - Column Metadata: Structured metadata lives in a dedicated
Metadataobject accessible viacolumn_metadata, supporting units, labels, and descriptions without polluting data arrays. - Serialization: Use
to_parquet()andfrom_parquet()for binary storage with full metadata round-tripping, orto_dict()for JSON-compatible portability. - Batch Processing: The
DataFrameGroupclass manages collections of related DataFrames with individual metadata contexts, ideal for experimental batches. - Interoperability: The
to_dataframe()method provides pandas-compatible objects with metadata attached to the!sdataattribute, ensuring provenance persists through downstream analysis.
Frequently Asked Questions
How does sdata store metadata differently from pandas DataFrame.attrs?
While pandas provides DataFrame.attrs as a catch-all dictionary, sdata enforces a structured schema through the Metadata class in sdata/metadata.py. Column-level attributes are validated during construction, and sdata-specific fields like _sdata_uuid and _sdata_version are managed automatically. This prevents accidental metadata loss and ensures consistent serialization across Parquet, JSON, and Excel formats.
Can I convert an sdata DataFrame back to a plain pandas DataFrame?
Yes. Call to_dataframe() on any sdata DataFrame to receive a standard pandas DataFrame copy. The returned object includes the full metadata dictionary under df.attrs["!sdata"], allowing downstream libraries to access provenance information without requiring the sdata package as a dependency.
What file formats support full metadata round-tripping?
Parquet provides the most robust support, storing metadata in df.attrs["_sdata"] alongside the binary data. JSON serialization embeds base64-encoded parquet bytes within the dictionary structure. Excel exports store metadata in hidden worksheets. CSV export is not recommended, as it cannot preserve the structured metadata hierarchy.
How do I handle unit conversions without losing original metadata?
Modify the numerical data in sdf.df directly (e.g., sdf.df["force"] = sdf.df["force"] * 1000 to convert kN to N), but leave the column_metadata entry unchanged or update it to reflect the new unit explicitly. sdata treats the metadata as documentation of the current state while maintaining an audit trail through the name, description, and versioning fields.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →