Understanding sdata's Metadata Schema: Design Rationale and Attribute Types

The sdata library implements a self-describing metadata schema built around the Attribute and Metadata classes in sdata/metadata.py, designed to ensure type safety, deterministic ordering, and bidirectional serialization across CSV, JSON, and DataFrame formats while maintaining a lightweight, dependency-free core.

The lepy/sdata repository provides a Python framework for storing experimental and simulation data with rich, portable metadata. At its heart lies a carefully designed metadata schema that balances semantic richness with practical interoperability. Understanding the design rationale behind sdata's metadata schema reveals how the library achieves both human readability and machine-friendly structured storage without heavy external dependencies.

Core Architecture: Attribute and Metadata Classes

The schema rests on two cooperating concepts defined in sdata/metadata.py: the atomic Attribute class and the container Metadata class.

The Attribute Class as a Typed Data Primitive

The Attribute class, defined at line 65 of sdata/metadata.py, represents a single metadata entry (e.g., force_x = 1.2 kN). It encapsulates value, unit, type, description, label, required flag, and optional ontology information.

A static DTYPES map (lines 68–74) normalizes string identifiers to concrete Python types (float, int, str, datetime, bool, list). This guarantees a consistent internal representation regardless of how the user supplies data. The constructor automatically infers the dtype via guess_dtype when the user omits explicit type declarations, allowing a flexible API where callers need not specify types explicitly.

The Metadata Container with Deterministic Ordering

The Metadata class acts as an ordered container of Attribute objects. Internally, it uses a SortedDict (initialized in sdata/metadata.py lines 13–15 and provided by sdata/base.py) to ensure attribute order is deterministic. This determinism is critical when serializing to CSV or generating content-based hashes.

The class defines a fixed set of attribute keys via the ATTRIBUTEKEYS constant (lines 6–8): name, value, unit, dtype, description, label, required, and ontology. Every metadata entry must expose these fields, creating a predictable tabular schema.

Type Safety and Automatic Inference

The schema emphasizes explicitness and safety through controlled type coercion. The guess_dtype_from_value and guess_value_dtype methods provide automatic type inference when loading loosely typed sources like CSV files. Type conversion is guarded: empty strings become None, while numeric empties become NaN.

The DTYPES mapping ensures that once a type is assigned, it remains consistent throughout the object's lifecycle, preventing silent data corruption during serialization round-trips.

Unit Extraction and Semantic Richness

Beyond primitive storage, the schema supports semantic richness through the ontology field and automatic unit detection. The set_unit_from_name method (lines 19–55 in sdata/metadata.py) parses common naming patterns—such as "Force (N)", "Length [mm]", or "Thickness <µm>"—to auto-populate the unit field without manual user input.

The required flag enables schema validation via the is_complete method, which verifies that all attributes marked as required contain non-empty values.

Serialization and Interoperability

The design prioritizes portability through bidirectional conversion helpers:

  • to_dict / from_dict
  • to_dataframe / from_dataframe
  • to_json / from_json
  • to_csv / from_csv

These methods support round-trip serialization while preserving types, units, and ordering. A SHA-3 hash (accessed via the sha3_256 property) provides a content-based identifier for each metadata block, enabling reproducible data pipelines and cheap integrity checks across storage backends like SQLite, HDF5, or MinIO.

Practical Implementation Examples

Create a metadata container and populate it with typed attributes:

from sdata.metadata import Metadata

# Initialize container

md = Metadata(name="experiment_01")

# Add attributes with automatic type inference

md.set_attr("force_x", 1.23, unit="N", description="Axial force")
md.set_attr("temperature", "23.5", unit="°C")  # Stored as str unless cast

md.set_attr("valid", True, dtype="bool")

# Extract units from attribute names automatically

md.set_unit_from_name("stress (MPa)")

Convert to a pandas DataFrame for analysis:

df = md.to_dataframe()
print(df)

               value  unit  dtype description label required ontology
name                                                                
force_x         1.23     N  float  Axial force    ""    False       ""
stress         None   MPa   None                      ""    False       ""
temperature    23.5    °C    str                      ""    False       ""
valid          True        bool                      ""    False       ""

Export to portable JSON format:

json_str = md.to_json()
print(json_str)
{
  "force_x": {"name":"force_x","value":1.23,"unit":"N","dtype":"float","description":"Axial force","label":"","required":false,"ontology":""},
  "temperature": {"name":"temperature","value":"23.5","unit":"°C","dtype":"str","description":"","label":"","required":false,"ontology":""},
  "valid": {"name":"valid","value":true,"unit":"","dtype":"bool","description":"","label":"","required":false,"ontology":""}
}

Compute a SHA-3 checksum for integrity verification:

checksum = md.sha3_256
print("Metadata hash:", checksum)
Metadata hash: 3f9c2e7a8c5b1d4e6f7a9c3e2b1a6d5f8c2e9b7a4d6e1c3f0a2b4c6d8e9f0a1

Verify schema completeness:


# Check if all required attributes are present

is_valid = md.is_complete()

Summary

  • Dual-class architecture: The Attribute class handles individual typed entries while Metadata manages ordered collections using a SortedDict from sdata/base.py.
  • Strict type normalization: The DTYPES map in sdata/metadata.py (lines 68–74) ensures consistent internal representation across Python primitives.
  • Automatic inference: Methods like guess_dtype and set_unit_from_name reduce boilerplate while maintaining schema integrity.
  • Deterministic ordering: The use of SortedDict guarantees consistent CSV serialization and reproducible hashes.
  • Bidirectional serialization: Native support for dict, DataFrame, JSON, and CSV formats with type preservation.
  • Cryptographic integrity: The sha3_256 property provides content-based addressing for reproducible pipelines.

Frequently Asked Questions

How does sdata handle type conversion for metadata attributes?

The Attribute class uses a static DTYPES map to normalize string identifiers (like "float" or "int") to Python types. When loading data, guess_dtype_from_value automatically infers types from raw values, while guarded conversion rules ensure empty strings become None and invalid numerics become NaN. This prevents silent data corruption during CSV or JSON ingestion.

What is the purpose of the SortedDict in the Metadata class?

The Metadata container uses SortedDict (defined in sdata/base.py) to maintain deterministic attribute ordering. This ensures that CSV exports and SHA-3 hashes remain consistent across different Python sessions and platforms, which is essential for reproducible scientific workflows and content-based data verification.

How does sdata ensure metadata integrity across different storage formats?

The schema provides bidirectional conversion methods (to_json/from_json, to_csv/from_csv, etc.) that preserve the fixed ATTRIBUTEKEYS structure. Additionally, the sha3_256 property computes a SHA-3 hash of the metadata's JSON representation, creating a portable fingerprint that can verify integrity regardless of whether the data resides in memory, SQLite, HDF5, or external object storage.

Can sdata extract units automatically from attribute names?

Yes. The set_unit_from_name method in sdata/metadata.py (lines 19–55) parses common scientific naming conventions—such as parentheses "Force (N)", brackets "Length [mm]", or angle brackets "Thickness <µm>"—to automatically populate the unit field. This reduces manual data entry errors while maintaining consistent unit metadata across experimental datasets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →