# Integrating sdata with pandas DataFrames: Best Practices for Scientific Data Manipulation

> Learn best practices for integrating sdata with pandas DataFrames. Attach metadata, IDs, and provenance to your scientific data for seamless serialization and analysis.

- Repository: [lepy/sdata](https://github.com/lepy/sdata)
- Tags: best-practices
- Published: 2026-03-05

---

**Integrate sdata with pandas DataFrames by wrapping your data in the `sdata.sclass.dataframe.DataFrame` class to attach persistent column-level metadata, unique identifiers, and scientific provenance that survives serialization across Parquet, JSON, and Excel formats.**

The `lepy/sdata` repository extends pandas with a lightweight metadata layer designed specifically for reproducible scientific computing. Integrating sdata with pandas DataFrames allows researchers to embed units, descriptions, and persistent object identities directly into their data structures without breaking compatibility with the pandas ecosystem.

## Understanding the sdata Architecture

The core integration lives in [`sdata/sclass/dataframe.py`](https://github.com/lepy/sdata/blob/main/sdata/sclass/dataframe.py), where the **`DataFrame`** class inherits from `sdata.base.Base`. Unlike standard pandas objects, every sdata DataFrame carries a **SUID** (short name stored in `_sdata_sname`) and a **UUID** (`_sdata_uuid`), ensuring reproducible references across experiments and file formats.

**Column-level metadata** is handled through the `Metadata` class defined in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py). This specialized dictionary stores `Attribute` objects containing labels, units, and descriptions for each column, accessed via the `column_metadata` property. While pandas offers `DataFrame.attrs`, sdata enforces structured validation and separates metadata from the raw data array, preventing pollution of the numerical contents.

## Creating DataFrames with Scientific Metadata

Begin by loading your raw data into a standard pandas DataFrame, then wrap it with sdata to inject scientific context. The constructor validates that every column in your data has a corresponding entry in the metadata dictionary.

```python
import pandas as pd
from sdata.sclass.dataframe import DataFrame

# Load raw experimental data

raw_df = pd.read_csv("experiment.csv")

# Define column-level metadata

col_meta = {
    "force": {"label": "Force", "unit": "N"},
    "displacement": {"label": "Displacement", "unit": "mm"},
}

# Wrap with sdata

sdf = DataFrame(
    df=raw_df,
    column_metadata=col_meta,
    name="TensileTest",
    description="Pull-test of alloy X at room temperature"
)

```

After creation, manipulate the underlying data through `sdf.df` using standard pandas operations. The metadata remains attached and accessible via `sdf.column_metadata` or the convenience property `sdf.cmdf`, which returns the metadata as a DataFrame for inspection.

## Serializing and Persisting Data

sdata implements format-aware serialization that preserves metadata where pandas alone would lose it. The `to_parquet()` method in [`sdata/sclass/dataframe.py`](https://github.com/lepy/sdata/blob/main/sdata/sclass/dataframe.py) writes binary Parquet files while embedding the full metadata hierarchy into `df.attrs["_sdata"]`.

```python

# Export to Parquet with embedded metadata

parquet_path = sdf.to_parquet(path="/data/parquet")

# Export to JSON with base64-encoded parquet bytes for portability

json_str = sdf.to_json()

# Export to Excel with metadata in a hidden sheet

sdf.to_xlsx("/data/tensile_test.xlsx")

```

Restore objects using the corresponding classmethods, which automatically reconstruct the UUID, SUID, and column metadata:

```python

# Full round-trip preservation

loaded = DataFrame.from_parquet(parquet_path)
assert loaded.uuid == sdf.uuid
assert loaded.column_metadata.to_dict() == sdf.column_metadata.to_dict()

```

The metadata field `_sdata_version` is set automatically during serialization, enabling version checking when loading legacy archives.

## Managing Multiple Experiments with DataFrameGroup

For batch experiments or time-series collections, use **`DataFrameGroup`** from [`sdata/sclass/dataframegroup.py`](https://github.com/lepy/sdata/blob/main/sdata/sclass/dataframegroup.py). This container aggregates multiple DataFrames, serializing each entry as base64-encoded parquet while preserving per-frame column metadata.

```python
from sdata.sclass.dataframegroup import DataFrameGroup
import glob

group = DataFrameGroup(name="BatchTensileTests")

for i, fpath in enumerate(sorted(glob.glob("batch/*.csv"))):
    df = pd.read_csv(fpath)
    meta = {
        "stress": {"label": "Engineering Stress", "unit": "MPa"},
        "strain": {"label": "Engineering Strain", "unit": "-"},
    }
    key = f"specimen_{i}"
    group.add_dataframe(key=key, df=df, column_metadata=meta)

# Serialize entire collection

group_dict = group.to_dict()

```

This pattern supports workflows where each measurement requires its own metadata context, such as varying test conditions across specimens.

## Downstream Integration and Provenance

Pass sdata objects to downstream libraries that expect plain pandas DataFrames using the **`to_dataframe()`** method. This returns a copy of the internal DataFrame with a special `!sdata` attribute containing the full metadata and description dictionary.

```python

# Convert for scikit-learn or visualization libraries

plain_df = sdf.to_dataframe()

# Access provenance in downstream code

sdata_meta = plain_df.attrs["!sdata"]
print(sdata_meta["!sdata_name"]["value"])  # "TensileTest"

print(sdata_meta["!sdata_description"]["value"])  # "Pull-test of alloy X"

```

This mechanism ensures that scientific context travels with the data even when crossing API boundaries into tools that lack native sdata support.

## Best Practices for Scientific Workflows

**Unit Conversions**: Perform calculations directly on `sdf.df` columns, but keep the original unit in `column_metadata`. This makes conversions explicit while preserving the source unit for audit trails.

**Column Renaming**: After modifying column names in the underlying DataFrame, call `Data.set_columnnames_from_metadata()` (available in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py)) to synchronize the `!sdata_column_*` attributes.

**Large Datasets**: Prefer `to_parquet()` for memory-efficient streaming. Unlike JSON serialization, this method does not require loading the entire object into memory as a Python dictionary.

**Concatenation**: When combining multiple sdata DataFrames manually, merge their `column_metadata` dictionaries explicitly to avoid key collisions. The library does not automatically reconcile metadata during concatenation operations.

**Metadata Validation**: The constructor enforces that every DataFrame column has a metadata entry. Use this validation to catch schema drift early in data pipelines.

## Summary

- **Object Identity**: Every sdata DataFrame carries a UUID and SUID stored in `_sdata_uuid` and `_sdata_sname`, enabling reproducible references across file formats.
- **Column Metadata**: Structured metadata lives in a dedicated `Metadata` object accessible via `column_metadata`, supporting units, labels, and descriptions without polluting data arrays.
- **Serialization**: Use `to_parquet()` and `from_parquet()` for binary storage with full metadata round-tripping, or `to_dict()` for JSON-compatible portability.
- **Batch Processing**: The `DataFrameGroup` class manages collections of related DataFrames with individual metadata contexts, ideal for experimental batches.
- **Interoperability**: The `to_dataframe()` method provides pandas-compatible objects with metadata attached to the `!sdata` attribute, ensuring provenance persists through downstream analysis.

## Frequently Asked Questions

### How does sdata store metadata differently from pandas DataFrame.attrs?

While pandas provides `DataFrame.attrs` as a catch-all dictionary, sdata enforces a structured schema through the `Metadata` class in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py). Column-level attributes are validated during construction, and sdata-specific fields like `_sdata_uuid` and `_sdata_version` are managed automatically. This prevents accidental metadata loss and ensures consistent serialization across Parquet, JSON, and Excel formats.

### Can I convert an sdata DataFrame back to a plain pandas DataFrame?

Yes. Call **`to_dataframe()`** on any sdata DataFrame to receive a standard pandas DataFrame copy. The returned object includes the full metadata dictionary under `df.attrs["!sdata"]`, allowing downstream libraries to access provenance information without requiring the sdata package as a dependency.

### What file formats support full metadata round-tripping?

**Parquet** provides the most robust support, storing metadata in `df.attrs["_sdata"]` alongside the binary data. **JSON** serialization embeds base64-encoded parquet bytes within the dictionary structure. **Excel** exports store metadata in hidden worksheets. CSV export is not recommended, as it cannot preserve the structured metadata hierarchy.

### How do I handle unit conversions without losing original metadata?

Modify the numerical data in `sdf.df` directly (e.g., `sdf.df["force"] = sdf.df["force"] * 1000` to convert kN to N), but leave the `column_metadata` entry unchanged or update it to reflect the new unit explicitly. sdata treats the metadata as documentation of the current state while maintaining an audit trail through the `name`, `description`, and versioning fields.