# Advanced sdata Serialization: Handling Complex Nested Structures to JSON, CSV, XLSX, and ZIP

> Master advanced sdata serialization. Export complex nested scientific data to JSON, CSV, XLSX, and ZIP formats deterministically, preserving metadata and relationships.

- Repository: [lepy/sdata](https://github.com/lepy/sdata)
- Tags: deep-dive
- Published: 2026-03-05

---

**Advanced sdata serialization enables deterministic export of hierarchical scientific datasets to JSON, CSV, XLSX, and ZIP formats while preserving metadata integrity and parent-child relationships through UUID-based identifiers.**

The `sdata` library provides a self-describing, hierarchical data model for scientific computing that couples tabular payloads with rich metadata blocks. This Python package implements advanced serialization capabilities in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py) that maintain complex nested structures across multiple export formats, making it ideal for experimental data management and reproducible research workflows.

## Core Architecture for Serialization

The serialization system rests on three interconnected classes that handle data containment, metadata management, and deterministic identification.

### The Data Container

The `Data` class in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py) serves as the primary export entry point. Each instance encapsulates a **pandas DataFrame** (the tabular payload), a `Metadata` object, a text description, and a hierarchical `group` dictionary containing child objects. The constructor automatically generates stable identifiers and timestamps, ensuring every object is ready for serialization upon instantiation.

```python
import sdata
import pandas as pd

df = pd.DataFrame({"Force [N]": [12.3, 15.6], "Displacement [mm]": [0.5, 1.0]})
data = sdata.Data(name="tensile_test", table=df)

```

### Metadata and Attributes

The `Metadata` class in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py) functions as an ordered dictionary of `Attribute` objects, storing both user-defined parameters and mandatory system attributes (version, UUID, timestamps). Each `Attribute` supports physical units, data types, and ontology references, enabling self-describing data files that carry their own schema.

```python
data.metadata.add("material", "Al-2024", dtype="str", unit="-")
data.metadata.add("temperature", 23.5, dtype="float", unit="°C")

```

### Deterministic Identifiers

The `SUUID` class in [`sdata/suuid.py`](https://github.com/lepy/sdata/blob/main/sdata/suuid.py) generates **Super UUIDs**—deterministic identifiers built from class name, object name, and parent UUID. This guarantees reproducible linking across data trees, ensuring that nested structures maintain stable references even after export and re-import.

## Exporting to Structured Formats

The `Data` class provides native methods for exporting to four primary formats, each preserving the metadata-table relationship differently.

### JSON Serialization

The `to_json()` method exports a dictionary containing `metadata`, `table`, and `description` keys to a JSON string or file. This format preserves the complete object state and supports round-trip deserialization via `Data.from_json()`.

```python
json_str = data.to_json()
reloaded = sdata.Data.from_json(json_str=json_str)

```

### CSV with Metadata Headers

The `to_csv()` method in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py) writes a header-prefixed metadata block (lines beginning with `#;`) followed by the CSV table content. This allows human-readable metadata inspection while maintaining pandas-compatible table parsing.

```python
data.to_csv("experiment.csv")

```

### Excel Workbooks

The `to_xlsx()` method creates multi-sheet Excel files with three distinct sheets: `metadata` (attribute key-value pairs), `table` (the DataFrame), and `description` (free-form text). This separation prevents metadata pollution of tabular data while keeping all information in a single file.

```python
data.to_xlsx("experiment.xlsx")

```

### ZIP Bundle Preservation

For complex nested structures, `sdata.tools.zip_folder()` aggregates folder hierarchies into portable archives. When combined with `Data.to_folder()`, this exports entire data trees—including subdirectories for child objects—as compressed ZIP bundles that preserve the full parent-child relationships.

```python
from sdata.tools import zip_folder

root.to_folder("export_dir", dtype="xlsx")
zip_folder("export_dir", "experiment_bundle.zip")

```

## Handling Hierarchical Data Trees

Advanced sdata serialization excels at maintaining complex nested structures through explicit parent-child relationships and recursive export capabilities.

### Parent-Child Relationships

The `Data.group` attribute stores child objects as a dictionary keyed by their SUUID strings. This enables construction of arbitrary-depth trees representing experimental campaigns, measurement series, or multi-modal datasets.

```python
root = sdata.Data(name="campaign")
subtest = sdata.Data(name="tensile_01", table=df)
root.add_data(subtest)  # Adds to root.group with deterministic SUUID

```

### Recursive Export

When calling `Data.to_folder()`, the method recursively iterates through `Data.group`, creating subdirectories for each child object. Each folder contains the object's metadata and table files, maintaining the hierarchy on disk. The `from_folder()` class method reconstructs this tree by traversing subdirectories and re-instantiating objects with their original UUID/SUUID relationships.

```python
root.to_folder("experiment_tree", dtype="csv")
restored = sdata.Data.from_folder("experiment_tree")

```

### Round-trip Integrity

The `sha3_256` property generates a reproducible checksum combining the JSON representation of metadata, the DataFrame content, and the description. This enables verification that serialization and deserialization preserved data integrity.

```python
assert data.sha3_256 == reloaded.sha3_256  # Verifies bit-perfect round-trip

```

## Practical Implementation Examples

### Exporting Nested Structures to Excel

```python
import sdata
import pandas as pd

# Create parent container

experiment = sdata.Data(name="fatigue_study")

# Add child measurements with metadata

for i in range(3):
    df = pd.DataFrame({"cycles": [1000, 2000], "stress": [200, 180]})
    run = sdata.Data(name=f"run_{i}", table=df)
    run.metadata.add("load_ratio", 0.1, unit="-")
    experiment.add_data(run)

# Export entire hierarchy to Excel (each object becomes a folder with .xlsx file)

experiment.to_folder("fatigue_data", dtype="xlsx")

```

### Creating Portable ZIP Archives

```python
from sdata.tools import zip_folder, unzip_folder

# Export and compress

experiment.to_folder("temp_export", dtype="json")
zip_folder("temp_export", "portable_experiment.zip")

# Later: extract and restore

unzip_folder("portable_experiment.zip", extract_to="restored")
full_tree = sdata.Data.from_folder("restored")

```

### SQLite Backend for Large Trees

For datasets too large for file-based export, `to_sqlite()` and `from_sqlite()` serialize the entire object tree to an SQLite database, preserving the `group` hierarchy in relational tables.

```python
experiment.to_sqlite("experiment.db")
database_restored = sdata.Data.from_sqlite("experiment.db")

```

## Summary

- **Advanced sdata serialization** handles complex nested structures through the `Data.group` hierarchy and deterministic SUUID linking.
- Export methods in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py) support **JSON** (`to_json`), **CSV** (`to_csv`), **XLSX** (`to_xlsx`), and **ZIP** (`to_folder` + `zip_folder`) formats.
- The `Metadata` and `Attribute` system in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py) preserves physical units and data types across all export formats.
- **SHA-3 checksums** verify round-trip integrity between serialization and deserialization.
- **Recursive folder export** maintains parent-child relationships, enabling reconstruction of complete experimental trees from ZIP archives or SQLite databases.

## Frequently Asked Questions

### How does sdata handle nested Data objects during export?

When calling `Data.to_folder()`, the method recursively traverses the `Data.group` dictionary and creates subdirectories for each child object. Each subdirectory contains the child's metadata and table files. The `from_folder()` class method reverses this process, reconstructing the object tree and restoring original UUID/SUUID relationships as implemented in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py).

### What is the difference between UUID and SUUID in sdata?

Standard UUIDs provide random unique identifiers, while **SUUIDs** (Super UUIDs) defined in [`sdata/suuid.py`](https://github.com/lepy/sdata/blob/main/sdata/suuid.py) are deterministic hashes built from the class name, object name, and parent UUID. SUUIDs guarantee reproducible identifiers across export/import cycles, enabling stable linking in complex nested data trees without central coordination.

### Can sdata preserve physical units during CSV export?

Yes. When using `to_csv()`, the metadata block (prefixed with `#;`) contains the `unit` attribute for each parameter. The `Metadata` class in [`sdata/metadata.py`](https://github.com/lepy/sdata/blob/main/sdata/metadata.py) serializes all Attribute properties—including units, data types, and descriptions—ensuring physical context survives the export process and is available after `Data.from_csv()` reconstruction.

### How do I verify data integrity after serialization?

Compare the `sha3_256` property values before and after serialization. This property calculates a SHA-3-256 hash over the JSON representation of metadata, the DataFrame content, and the description text. Matching checksums confirm that `to_json`/`from_json`, `to_csv`/`from_csv`, or other round-trip operations preserved data integrity exactly as implemented in the source code.