# Understanding sdata's Metadata Schema: Design Rationale and Attribute Types

> Explore sdata's metadata schema design rationale and attribute types. Learn how it ensures type safety, deterministic ordering, and bidirectional serialization across formats. Lightweight and dependency-free.

- Repository: [lepy/sdata](https://github.com/lepy/sdata)
- Tags: deep-dive
- Published: 2026-03-06

---

**The `sdata` library implements a self-describing metadata schema built around the `Attribute` and `Metadata` classes in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py), designed to ensure type safety, deterministic ordering, and bidirectional serialization across CSV, JSON, and DataFrame formats while maintaining a lightweight, dependency-free core.**

The `lepy/sdata` repository provides a Python framework for storing experimental and simulation data with rich, portable metadata. At its heart lies a carefully designed metadata schema that balances semantic richness with practical interoperability. Understanding the design rationale behind sdata's metadata schema reveals how the library achieves both human readability and machine-friendly structured storage without heavy external dependencies.

## Core Architecture: Attribute and Metadata Classes

The schema rests on two cooperating concepts defined in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py): the atomic `Attribute` class and the container `Metadata` class.

### The Attribute Class as a Typed Data Primitive

The `Attribute` class, defined at line 65 of [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py), represents a single metadata entry (e.g., *force_x* = 1.2 kN). It encapsulates **value**, **unit**, **type**, **description**, **label**, **required flag**, and optional **ontology** information.

A static `DTYPES` map (lines 68–74) normalizes string identifiers to concrete Python types (`float`, `int`, `str`, `datetime`, `bool`, `list`). This guarantees a **consistent internal representation** regardless of how the user supplies data. The constructor automatically infers the dtype via `guess_dtype` when the user omits explicit type declarations, allowing a flexible API where callers need not specify types explicitly.

### The Metadata Container with Deterministic Ordering

The `Metadata` class acts as an ordered container of `Attribute` objects. Internally, it uses a `SortedDict` (initialized in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py) lines 13–15 and provided by [`sdata/base.py`](https://github.com/lepy/sdata/blob/main/sdata/base.py)) to ensure attribute order is deterministic. This determinism is critical when serializing to CSV or generating content-based hashes.

The class defines a fixed set of **attribute keys** via the `ATTRIBUTEKEYS` constant (lines 6–8): `name`, `value`, `unit`, `dtype`, `description`, `label`, `required`, and `ontology`. Every metadata entry must expose these fields, creating a predictable tabular schema.

## Type Safety and Automatic Inference

The schema emphasizes **explicitness** and **safety** through controlled type coercion. The `guess_dtype_from_value` and `guess_value_dtype` methods provide automatic type inference when loading loosely typed sources like CSV files. Type conversion is guarded: empty strings become `None`, while numeric empties become `NaN`.

The `DTYPES` mapping ensures that once a type is assigned, it remains consistent throughout the object's lifecycle, preventing silent data corruption during serialization round-trips.

## Unit Extraction and Semantic Richness

Beyond primitive storage, the schema supports semantic richness through the `ontology` field and automatic unit detection. The `set_unit_from_name` method (lines 19–55 in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py)) parses common naming patterns—such as `"Force (N)"`, `"Length [mm]"`, or `"Thickness <µm>"`—to auto-populate the `unit` field without manual user input.

The `required` flag enables schema validation via the `is_complete` method, which verifies that all attributes marked as required contain non-empty values.

## Serialization and Interoperability

The design prioritizes **portability** through bidirectional conversion helpers:

- `to_dict` / `from_dict`
- `to_dataframe` / `from_dataframe`
- `to_json` / `from_json`
- `to_csv` / `from_csv`

These methods support round-trip serialization while preserving types, units, and ordering. A SHA-3 hash (accessed via the `sha3_256` property) provides a **content-based identifier** for each metadata block, enabling reproducible data pipelines and cheap integrity checks across storage backends like SQLite, HDF5, or MinIO.

## Practical Implementation Examples

Create a metadata container and populate it with typed attributes:

```python
from sdata.metadata import Metadata

# Initialize container

md = Metadata(name="experiment_01")

# Add attributes with automatic type inference

md.set_attr("force_x", 1.23, unit="N", description="Axial force")
md.set_attr("temperature", "23.5", unit="°C")  # Stored as str unless cast

md.set_attr("valid", True, dtype="bool")

# Extract units from attribute names automatically

md.set_unit_from_name("stress (MPa)")

```

Convert to a pandas DataFrame for analysis:

```python
df = md.to_dataframe()
print(df)

```

```

               value  unit  dtype description label required ontology
name                                                                
force_x         1.23     N  float  Axial force    ""    False       ""
stress         None   MPa   None                      ""    False       ""
temperature    23.5    °C    str                      ""    False       ""
valid          True        bool                      ""    False       ""

```

Export to portable JSON format:

```python
json_str = md.to_json()
print(json_str)

```

```json
{
  "force_x": {"name":"force_x","value":1.23,"unit":"N","dtype":"float","description":"Axial force","label":"","required":false,"ontology":""},
  "temperature": {"name":"temperature","value":"23.5","unit":"°C","dtype":"str","description":"","label":"","required":false,"ontology":""},
  "valid": {"name":"valid","value":true,"unit":"","dtype":"bool","description":"","label":"","required":false,"ontology":""}
}

```

Compute a SHA-3 checksum for integrity verification:

```python
checksum = md.sha3_256
print("Metadata hash:", checksum)

```

```text
Metadata hash: 3f9c2e7a8c5b1d4e6f7a9c3e2b1a6d5f8c2e9b7a4d6e1c3f0a2b4c6d8e9f0a1

```

Verify schema completeness:

```python

# Check if all required attributes are present

is_valid = md.is_complete()

```

## Summary

- **Dual-class architecture**: The `Attribute` class handles individual typed entries while `Metadata` manages ordered collections using a `SortedDict` from [`sdata/base.py`](https://github.com/lepy/sdata/blob/main/sdata/base.py).
- **Strict type normalization**: The `DTYPES` map in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py) (lines 68–74) ensures consistent internal representation across Python primitives.
- **Automatic inference**: Methods like `guess_dtype` and `set_unit_from_name` reduce boilerplate while maintaining schema integrity.
- **Deterministic ordering**: The use of `SortedDict` guarantees consistent CSV serialization and reproducible hashes.
- **Bidirectional serialization**: Native support for dict, DataFrame, JSON, and CSV formats with type preservation.
- **Cryptographic integrity**: The `sha3_256` property provides content-based addressing for reproducible pipelines.

## Frequently Asked Questions

### How does sdata handle type conversion for metadata attributes?

The `Attribute` class uses a static `DTYPES` map to normalize string identifiers (like `"float"` or `"int"`) to Python types. When loading data, `guess_dtype_from_value` automatically infers types from raw values, while guarded conversion rules ensure empty strings become `None` and invalid numerics become `NaN`. This prevents silent data corruption during CSV or JSON ingestion.

### What is the purpose of the SortedDict in the Metadata class?

The `Metadata` container uses `SortedDict` (defined in [`sdata/base.py`](https://github.com/lepy/sdata/blob/main/sdata/base.py)) to maintain deterministic attribute ordering. This ensures that CSV exports and SHA-3 hashes remain consistent across different Python sessions and platforms, which is essential for reproducible scientific workflows and content-based data verification.

### How does sdata ensure metadata integrity across different storage formats?

The schema provides bidirectional conversion methods (`to_json`/`from_json`, `to_csv`/`from_csv`, etc.) that preserve the fixed `ATTRIBUTEKEYS` structure. Additionally, the `sha3_256` property computes a SHA-3 hash of the metadata's JSON representation, creating a portable fingerprint that can verify integrity regardless of whether the data resides in memory, SQLite, HDF5, or external object storage.

### Can sdata extract units automatically from attribute names?

Yes. The `set_unit_from_name` method in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py) (lines 19–55) parses common scientific naming conventions—such as parentheses `"Force (N)"`, brackets `"Length [mm]"`, or angle brackets `"Thickness <µm>"`—to automatically populate the `unit` field. This reduces manual data entry errors while maintaining consistent unit metadata across experimental datasets.