Advanced sdata Serialization: Handling Complex Nested Structures to JSON, CSV, XLSX, and ZIP
Advanced sdata serialization enables deterministic export of hierarchical scientific datasets to JSON, CSV, XLSX, and ZIP formats while preserving metadata integrity and parent-child relationships through UUID-based identifiers.
The sdata library provides a self-describing, hierarchical data model for scientific computing that couples tabular payloads with rich metadata blocks. This Python package implements advanced serialization capabilities in sdata/data.py that maintain complex nested structures across multiple export formats, making it ideal for experimental data management and reproducible research workflows.
Core Architecture for Serialization
The serialization system rests on three interconnected classes that handle data containment, metadata management, and deterministic identification.
The Data Container
The Data class in sdata/data.py serves as the primary export entry point. Each instance encapsulates a pandas DataFrame (the tabular payload), a Metadata object, a text description, and a hierarchical group dictionary containing child objects. The constructor automatically generates stable identifiers and timestamps, ensuring every object is ready for serialization upon instantiation.
import sdata
import pandas as pd
df = pd.DataFrame({"Force [N]": [12.3, 15.6], "Displacement [mm]": [0.5, 1.0]})
data = sdata.Data(name="tensile_test", table=df)
Metadata and Attributes
The Metadata class in sdata/metadata.py functions as an ordered dictionary of Attribute objects, storing both user-defined parameters and mandatory system attributes (version, UUID, timestamps). Each Attribute supports physical units, data types, and ontology references, enabling self-describing data files that carry their own schema.
data.metadata.add("material", "Al-2024", dtype="str", unit="-")
data.metadata.add("temperature", 23.5, dtype="float", unit="°C")
Deterministic Identifiers
The SUUID class in sdata/suuid.py generates Super UUIDs—deterministic identifiers built from class name, object name, and parent UUID. This guarantees reproducible linking across data trees, ensuring that nested structures maintain stable references even after export and re-import.
Exporting to Structured Formats
The Data class provides native methods for exporting to four primary formats, each preserving the metadata-table relationship differently.
JSON Serialization
The to_json() method exports a dictionary containing metadata, table, and description keys to a JSON string or file. This format preserves the complete object state and supports round-trip deserialization via Data.from_json().
json_str = data.to_json()
reloaded = sdata.Data.from_json(json_str=json_str)
CSV with Metadata Headers
The to_csv() method in sdata/data.py writes a header-prefixed metadata block (lines beginning with #;) followed by the CSV table content. This allows human-readable metadata inspection while maintaining pandas-compatible table parsing.
data.to_csv("experiment.csv")
Excel Workbooks
The to_xlsx() method creates multi-sheet Excel files with three distinct sheets: metadata (attribute key-value pairs), table (the DataFrame), and description (free-form text). This separation prevents metadata pollution of tabular data while keeping all information in a single file.
data.to_xlsx("experiment.xlsx")
ZIP Bundle Preservation
For complex nested structures, sdata.tools.zip_folder() aggregates folder hierarchies into portable archives. When combined with Data.to_folder(), this exports entire data trees—including subdirectories for child objects—as compressed ZIP bundles that preserve the full parent-child relationships.
from sdata.tools import zip_folder
root.to_folder("export_dir", dtype="xlsx")
zip_folder("export_dir", "experiment_bundle.zip")
Handling Hierarchical Data Trees
Advanced sdata serialization excels at maintaining complex nested structures through explicit parent-child relationships and recursive export capabilities.
Parent-Child Relationships
The Data.group attribute stores child objects as a dictionary keyed by their SUUID strings. This enables construction of arbitrary-depth trees representing experimental campaigns, measurement series, or multi-modal datasets.
root = sdata.Data(name="campaign")
subtest = sdata.Data(name="tensile_01", table=df)
root.add_data(subtest) # Adds to root.group with deterministic SUUID
Recursive Export
When calling Data.to_folder(), the method recursively iterates through Data.group, creating subdirectories for each child object. Each folder contains the object's metadata and table files, maintaining the hierarchy on disk. The from_folder() class method reconstructs this tree by traversing subdirectories and re-instantiating objects with their original UUID/SUUID relationships.
root.to_folder("experiment_tree", dtype="csv")
restored = sdata.Data.from_folder("experiment_tree")
Round-trip Integrity
The sha3_256 property generates a reproducible checksum combining the JSON representation of metadata, the DataFrame content, and the description. This enables verification that serialization and deserialization preserved data integrity.
assert data.sha3_256 == reloaded.sha3_256 # Verifies bit-perfect round-trip
Practical Implementation Examples
Exporting Nested Structures to Excel
import sdata
import pandas as pd
# Create parent container
experiment = sdata.Data(name="fatigue_study")
# Add child measurements with metadata
for i in range(3):
df = pd.DataFrame({"cycles": [1000, 2000], "stress": [200, 180]})
run = sdata.Data(name=f"run_{i}", table=df)
run.metadata.add("load_ratio", 0.1, unit="-")
experiment.add_data(run)
# Export entire hierarchy to Excel (each object becomes a folder with .xlsx file)
experiment.to_folder("fatigue_data", dtype="xlsx")
Creating Portable ZIP Archives
from sdata.tools import zip_folder, unzip_folder
# Export and compress
experiment.to_folder("temp_export", dtype="json")
zip_folder("temp_export", "portable_experiment.zip")
# Later: extract and restore
unzip_folder("portable_experiment.zip", extract_to="restored")
full_tree = sdata.Data.from_folder("restored")
SQLite Backend for Large Trees
For datasets too large for file-based export, to_sqlite() and from_sqlite() serialize the entire object tree to an SQLite database, preserving the group hierarchy in relational tables.
experiment.to_sqlite("experiment.db")
database_restored = sdata.Data.from_sqlite("experiment.db")
Summary
- Advanced sdata serialization handles complex nested structures through the
Data.grouphierarchy and deterministic SUUID linking. - Export methods in
sdata/data.pysupport JSON (to_json), CSV (to_csv), XLSX (to_xlsx), and ZIP (to_folder+zip_folder) formats. - The
MetadataandAttributesystem insdata/metadata.pypreserves physical units and data types across all export formats. - SHA-3 checksums verify round-trip integrity between serialization and deserialization.
- Recursive folder export maintains parent-child relationships, enabling reconstruction of complete experimental trees from ZIP archives or SQLite databases.
Frequently Asked Questions
How does sdata handle nested Data objects during export?
When calling Data.to_folder(), the method recursively traverses the Data.group dictionary and creates subdirectories for each child object. Each subdirectory contains the child's metadata and table files. The from_folder() class method reverses this process, reconstructing the object tree and restoring original UUID/SUUID relationships as implemented in sdata/data.py.
What is the difference between UUID and SUUID in sdata?
Standard UUIDs provide random unique identifiers, while SUUIDs (Super UUIDs) defined in sdata/suuid.py are deterministic hashes built from the class name, object name, and parent UUID. SUUIDs guarantee reproducible identifiers across export/import cycles, enabling stable linking in complex nested data trees without central coordination.
Can sdata preserve physical units during CSV export?
Yes. When using to_csv(), the metadata block (prefixed with #;) contains the unit attribute for each parameter. The Metadata class in sdata/metadata.py serializes all Attribute properties—including units, data types, and descriptions—ensuring physical context survives the export process and is available after Data.from_csv() reconstruction.
How do I verify data integrity after serialization?
Compare the sha3_256 property values before and after serialization. This property calculates a SHA-3-256 hash over the JSON representation of metadata, the DataFrame content, and the description text. Matching checksums confirm that to_json/from_json, to_csv/from_csv, or other round-trip operations preserved data integrity exactly as implemented in the source code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →