# Implementing Robust Data Validation and Required Attributes in sdata Projects

> Implement robust data validation and required attributes in sdata projects ensuring data integrity with type-safe casting and completeness checks through Attribute and Metadata classes.

- Repository: [lepy/sdata](https://github.com/lepy/sdata)
- Tags: how-to-guide
- Published: 2026-03-06

---

**The sdata library enforces data integrity through the `Attribute` and `Metadata` classes in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py), which provide type-safe casting, required-field flags, and completeness checks to ensure every dataset carries validated, self-describing metadata.**

The sdata repository provides a structured framework for storing experimental data alongside comprehensive metadata containers. By leveraging the built-in validation mechanisms in the `Attribute` and `Metadata` classes, developers can implement strict required-field policies and automatic type casting that prevent incomplete datasets from entering analysis pipelines.

## Core Architecture for Data Validation

The validation system centers on two primary classes defined in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py) that handle individual metadata entries and container-wide integrity checks.

### The Attribute Class

The **`Attribute`** class encapsulates a single metadata entry and enforces type safety through robust casting logic. Located at lines 65–104 of [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py), this class provides:

- **`__init__(name, value, **kwargs)`** – Initializes the attribute with dtype, unit, and required flag parameters.
- **`_set_value`** – Internal method that normalizes inputs: empty strings become `""`, missing numerics convert to `np.nan`, and booleans accept multiple string representations including `"true"` and `"1"`.
- **`required` property** – Normalizes truthy values to ensure consistent mandatory field tracking.
- **`to_dict` and `to_csv`** – Serialization methods for metadata interchange.

### The Metadata Container

The **`Metadata`** class acts as a dictionary-like container for `Attribute` objects, supplying bulk operations and completeness verification. Key methods include:

- **`set_attr(name, value, **kwargs)`** – Creates or updates an attribute with optional dtype and unit specifications.
- **`add_attribute(attr, **kwargs)`** – Inserts an existing `Attribute` instance with optional prefix handling.
- **`required_attributes`** – Dynamic property returning a view of all attributes marked as required.
- **`is_complete()`** – Validates that every required attribute contains a non-empty, non-None value.

### Helper Functions for Type Inference

The module includes utility functions that streamline data ingestion:

- **`extract_name_unit`** (lines 19–58) – Parses scientific unit strings from attribute names, enabling automatic unit extraction via `Metadata.set_unit_from_name`.
- **`guess_dtype_from_value`** (lines 79–106) – Heuristically determines appropriate dtypes (int, float, bool, str, list) when importing free-form dictionaries.

## The Validation Workflow in sdata

According to the source code in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py), sdata implements a five-stage validation pipeline that processes metadata from construction through completeness verification.

1. **Construction** – Instantiate a `Metadata` container and populate it via `set_attr` or `add_attribute` calls.
2. **Type Guessing** – When dtype is omitted, `Attribute.__init__` invokes `guess_dtype` (falling back to `guess_dtype_from_value`) to infer the correct Python type from raw values.
3. **Casting** – The `_set_value` method normalizes inputs, handling empty strings, NaN values, boolean strings, and comma-separated lists without raising exceptions.
4. **Required Flag Assignment** – The `required` property stores a boolean flag; `Metadata.required_attributes` builds a dynamic view of all mandatory fields.
5. **Completeness Verification** – `Metadata.is_complete()` iterates over `required_attributes` and returns `False` if any required field is `None` or empty, enabling programmatic enforcement of data integrity before persistence.

## Practical Implementation Examples

### Defining Mandatory Metadata Fields

Use `set_attr` with `required=True` to declare critical metadata that must be present before export:

```python
from sdata.metadata import Metadata

meta = Metadata(name="MyExperiment")

# Define required attributes with explicit types

meta.set_attr("sample_id", "AB-001", required=True, dtype="str")
meta.set_attr("temperature", 23.5, required=True, dtype="float", unit="°C")
meta.set_attr("start_time", "2023-07-01T10:00:00Z", dtype="timestamp")

```

The `required=True` flag marks `sample_id` and `temperature` as mandatory. When `meta.is_complete()` executes later, it verifies these fields contain non-empty values.

### Automatic Type Casting from Raw Input

When importing unstructured data, sdata automatically infers and converts types:

```python
raw = {
    "sample_id": "AB-001",
    "temperature": "23.5",          # String resembling float

    "valid": "true",                # Truthy string

    "tags": "steel, tensile"        # CSV-style list

}

meta.update_from_dict(raw)

```

The `guess_dtype_from_value` helper deduces that `"23.5"` should cast to float, `"true"` to boolean, and comma-separated strings to Python lists.

### Validating Completeness Before Export

Gate downstream processing with explicit completeness checks:

```python
if not meta.is_complete():
    missing = [attr.name for attr in meta.required_attributes if not attr.value]
    raise ValueError(f"Missing required metadata fields: {missing}")

json_str = meta.to_json()

```

### Round-Trip Serialization

Metadata with required flags persists through export and re-import:

```python

# Export to CSV with headers

csv_text = meta.to_csv_header(filepath="meta.csv")

# Later reconstruction preserves validation rules

meta2 = Metadata.from_csv("meta.csv")
assert meta2.is_complete() == meta.is_complete()

```

## Summary

- **Explicit required flags** (`required=True`) provide declarative enforcement of critical metadata fields in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py).
- **Robust casting logic** inside `Attribute._set_value` handles empty strings, NaN numerics, boolean representations, and list-style inputs without exceptions.
- **Programmatic validation** via `Metadata.is_complete()` enables simple gating of storage, analysis, or publication workflows.
- **Pandas interoperability** allows conversion to DataFrames for bulk inspection or merging with experimental data tables via `to_dataframe` and `from_dataframe`.
- **Serialization support** maintains required attribute states through CSV and JSON round-trips.

## Frequently Asked Questions

### How do I mark an attribute as required in sdata?

Pass `required=True` to `Metadata.set_attr()` or `Attribute.__init__()`. The `required` property normalizes truthy values, and `Metadata.required_attributes` provides a dynamic view of all mandatory fields for programmatic access.

### What happens if a required attribute is missing during validation?

`Metadata.is_complete()` returns `False` when any required attribute contains `None` or an empty string. Callers should check this method before exporting or processing data, typically raising a `ValueError` or prompting for missing values when the check fails.

### How does sdata handle type conversion for boolean and list values?

The `Attribute._set_value` method accepts multiple boolean representations (`"true"`, `"1"`, `"yes"`) and converts them to Python booleans. For lists, comma-separated strings are parsed into Python list objects. The `guess_dtype_from_value` helper in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py) automatically infers these types during dictionary import.

### Can I export metadata with required flags to CSV and JSON formats?

Yes. The `to_dict`, `to_csv`, and `to_json` methods serialize the `required` flag along with dtype and unit information. When re-importing via `from_csv` or `from_dict`, the `Metadata` container reconstructs the original validation rules, preserving required field constraints through the round-trip.