Understanding sdata's Metadata Schema: Design Rationale and Attribute Types
The sdata library implements a self-describing metadata schema built around the Attribute and Metadata classes in sdata/metadata.py, designed to ensure type safety, deterministic ordering, and bidirectional serialization across CSV, JSON, and DataFrame formats while maintaining a lightweight, dependency-free core.
The lepy/sdata repository provides a Python framework for storing experimental and simulation data with rich, portable metadata. At its heart lies a carefully designed metadata schema that balances semantic richness with practical interoperability. Understanding the design rationale behind sdata's metadata schema reveals how the library achieves both human readability and machine-friendly structured storage without heavy external dependencies.
Core Architecture: Attribute and Metadata Classes
The schema rests on two cooperating concepts defined in sdata/metadata.py: the atomic Attribute class and the container Metadata class.
The Attribute Class as a Typed Data Primitive
The Attribute class, defined at line 65 of sdata/metadata.py, represents a single metadata entry (e.g., force_x = 1.2 kN). It encapsulates value, unit, type, description, label, required flag, and optional ontology information.
A static DTYPES map (lines 68–74) normalizes string identifiers to concrete Python types (float, int, str, datetime, bool, list). This guarantees a consistent internal representation regardless of how the user supplies data. The constructor automatically infers the dtype via guess_dtype when the user omits explicit type declarations, allowing a flexible API where callers need not specify types explicitly.
The Metadata Container with Deterministic Ordering
The Metadata class acts as an ordered container of Attribute objects. Internally, it uses a SortedDict (initialized in sdata/metadata.py lines 13–15 and provided by sdata/base.py) to ensure attribute order is deterministic. This determinism is critical when serializing to CSV or generating content-based hashes.
The class defines a fixed set of attribute keys via the ATTRIBUTEKEYS constant (lines 6–8): name, value, unit, dtype, description, label, required, and ontology. Every metadata entry must expose these fields, creating a predictable tabular schema.
Type Safety and Automatic Inference
The schema emphasizes explicitness and safety through controlled type coercion. The guess_dtype_from_value and guess_value_dtype methods provide automatic type inference when loading loosely typed sources like CSV files. Type conversion is guarded: empty strings become None, while numeric empties become NaN.
The DTYPES mapping ensures that once a type is assigned, it remains consistent throughout the object's lifecycle, preventing silent data corruption during serialization round-trips.
Unit Extraction and Semantic Richness
Beyond primitive storage, the schema supports semantic richness through the ontology field and automatic unit detection. The set_unit_from_name method (lines 19–55 in sdata/metadata.py) parses common naming patterns—such as "Force (N)", "Length [mm]", or "Thickness <µm>"—to auto-populate the unit field without manual user input.
The required flag enables schema validation via the is_complete method, which verifies that all attributes marked as required contain non-empty values.
Serialization and Interoperability
The design prioritizes portability through bidirectional conversion helpers:
to_dict/from_dictto_dataframe/from_dataframeto_json/from_jsonto_csv/from_csv
These methods support round-trip serialization while preserving types, units, and ordering. A SHA-3 hash (accessed via the sha3_256 property) provides a content-based identifier for each metadata block, enabling reproducible data pipelines and cheap integrity checks across storage backends like SQLite, HDF5, or MinIO.
Practical Implementation Examples
Create a metadata container and populate it with typed attributes:
from sdata.metadata import Metadata
# Initialize container
md = Metadata(name="experiment_01")
# Add attributes with automatic type inference
md.set_attr("force_x", 1.23, unit="N", description="Axial force")
md.set_attr("temperature", "23.5", unit="°C") # Stored as str unless cast
md.set_attr("valid", True, dtype="bool")
# Extract units from attribute names automatically
md.set_unit_from_name("stress (MPa)")
Convert to a pandas DataFrame for analysis:
df = md.to_dataframe()
print(df)
value unit dtype description label required ontology
name
force_x 1.23 N float Axial force "" False ""
stress None MPa None "" False ""
temperature 23.5 °C str "" False ""
valid True bool "" False ""
Export to portable JSON format:
json_str = md.to_json()
print(json_str)
{
"force_x": {"name":"force_x","value":1.23,"unit":"N","dtype":"float","description":"Axial force","label":"","required":false,"ontology":""},
"temperature": {"name":"temperature","value":"23.5","unit":"°C","dtype":"str","description":"","label":"","required":false,"ontology":""},
"valid": {"name":"valid","value":true,"unit":"","dtype":"bool","description":"","label":"","required":false,"ontology":""}
}
Compute a SHA-3 checksum for integrity verification:
checksum = md.sha3_256
print("Metadata hash:", checksum)
Metadata hash: 3f9c2e7a8c5b1d4e6f7a9c3e2b1a6d5f8c2e9b7a4d6e1c3f0a2b4c6d8e9f0a1
Verify schema completeness:
# Check if all required attributes are present
is_valid = md.is_complete()
Summary
- Dual-class architecture: The
Attributeclass handles individual typed entries whileMetadatamanages ordered collections using aSortedDictfromsdata/base.py. - Strict type normalization: The
DTYPESmap insdata/metadata.py(lines 68–74) ensures consistent internal representation across Python primitives. - Automatic inference: Methods like
guess_dtypeandset_unit_from_namereduce boilerplate while maintaining schema integrity. - Deterministic ordering: The use of
SortedDictguarantees consistent CSV serialization and reproducible hashes. - Bidirectional serialization: Native support for dict, DataFrame, JSON, and CSV formats with type preservation.
- Cryptographic integrity: The
sha3_256property provides content-based addressing for reproducible pipelines.
Frequently Asked Questions
How does sdata handle type conversion for metadata attributes?
The Attribute class uses a static DTYPES map to normalize string identifiers (like "float" or "int") to Python types. When loading data, guess_dtype_from_value automatically infers types from raw values, while guarded conversion rules ensure empty strings become None and invalid numerics become NaN. This prevents silent data corruption during CSV or JSON ingestion.
What is the purpose of the SortedDict in the Metadata class?
The Metadata container uses SortedDict (defined in sdata/base.py) to maintain deterministic attribute ordering. This ensures that CSV exports and SHA-3 hashes remain consistent across different Python sessions and platforms, which is essential for reproducible scientific workflows and content-based data verification.
How does sdata ensure metadata integrity across different storage formats?
The schema provides bidirectional conversion methods (to_json/from_json, to_csv/from_csv, etc.) that preserve the fixed ATTRIBUTEKEYS structure. Additionally, the sha3_256 property computes a SHA-3 hash of the metadata's JSON representation, creating a portable fingerprint that can verify integrity regardless of whether the data resides in memory, SQLite, HDF5, or external object storage.
Can sdata extract units automatically from attribute names?
Yes. The set_unit_from_name method in sdata/metadata.py (lines 19–55) parses common scientific naming conventions—such as parentheses "Force (N)", brackets "Length [mm]", or angle brackets "Thickness <µm>"—to automatically populate the unit field. This reduces manual data entry errors while maintaining consistent unit metadata across experimental datasets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →