# How sdata Handles Hierarchical Data Structures: Parent-Child Relationships Explained

> Discover how sdata manages hierarchical data and parent-child relationships. Learn about object metadata, SUUID accessors, and efficient data navigation in this technical explanation.

- Repository: [lepy/sdata](https://github.com/lepy/sdata)
- Tags: deep-dive
- Published: 2026-03-05

---

**`sdata` implements hierarchical tree-like relationships by storing parent identifiers in object metadata and providing accessor methods that return `SUUID` objects for navigation.**

The `lepy/sdata` library provides a robust framework for managing **hierarchical data structures** through explicit parent-child relationships embedded in object metadata. By storing lightweight parent references (SUUID snames or raw UUIDs) and maintaining in-memory child collections, `sdata` enables complex data organization while ensuring relationships persist through serialization and storage operations.

## Storing Parent References in Metadata

The foundation of `sdata` hierarchical handling lies in the `Base` class defined in [`sdata/base.py`](https://github.com/lepy/sdata/blob/main/sdata/base.py). This class establishes a standardized mechanism for tracking parent objects using metadata keys.

### The SDATA_PARENT_SNAME Constant

In [`sdata/base.py`](https://github.com/lepy/sdata/blob/main/sdata/base.py), the `Base` class defines the metadata key `SDATA_PARENT_SNAME = "_sdata_parent_sname"`. During initialization via `Base.__init__()`, the constructor accepts an optional `parent` argument. When provided, the parent object's SUUID is extracted using `SUUID.from_obj(parent)`, and the sname (short string identifier) is persisted to metadata:

```python
self.metadata.add(
    self.SDATA_PARENT_SNAME, parent_sname, dtype="str",
    description="sname of the parent"
)

```

If no parent is supplied, an empty string is stored, effectively marking the object as a root node.

### Parent SUUID Extraction

The initialization process converts the parent object to its SUUID representation before storage. This ensures that parent references remain lightweight and consistent across the object hierarchy, using the short sname format rather than full object serialization.

## Accessing Parent Objects

Once parent references are stored in metadata, `sdata` provides intuitive accessors for retrieving parent relationships without manual metadata parsing.

### The get_parent() Method and parent Property

The `Base` class implements both a `get_parent()` method and a `parent` property (located in [`sdata/base.py`](https://github.com/lepy/sdata/blob/main/sdata/base.py)) that reconstruct the parent `SUUID` from stored metadata:

```python
sname = self.metadata.get(self.SDATA_PARENT_SNAME).value
return SUUID.from_suuid_sname(sname) if sname else None

```

This approach returns a proper `SUUID` object when a parent exists, or `None` for root-level objects. The dual interface (property and method) accommodates different coding patterns while maintaining consistent underlying logic.

## Specialized Handling in the Data Class

While the `Base` class uses sname-based parent tracking, the `Data` class in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py) implements a variation using raw UUIDs for specific persistence requirements.

### UUID-Based Parent Metadata

The `Data` class defines `SDATA_PARENT = "!sdata_parent"` as its metadata key. By default, initialization adds an empty parent reference:

```python
self.metadata.add(self.SDATA_PARENT, "", ...)

```

This raw UUID approach differs from the `Base` class sname storage, providing direct identifier access for vault operations and database persistence.

### Parent Link Propagation During Copy Operations

When duplicating data objects via `Data.copy()`, the library automatically establishes parent relationships between the original and copied instances. The copy method records the source object's SUUID as the new instance's parent:

```python
data.metadata.add(self.SDATA_PARENT, self.suuid)

```

This automatic lineage tracking ensures that data provenance and versioning hierarchies are maintained without manual intervention.

## Persisting Hierarchies to Storage

Parent-child relationships survive serialization through explicit handling in `sdata` vault backends, ensuring hierarchical integrity across storage formats.

### SQLite Vault Implementation

In [`sdata/iolib/vault.py`](https://github.com/lepy/sdata/blob/main/sdata/iolib/vault.py), the `VaultSqliteIndex` class manages hierarchical persistence through database operations. When updating metadata via `VaultSqliteIndex.update_from_metadata()`, the parent UUID is stored in the `!sdata_parent` column. Additionally, the system creates explicit graph edges representing the relationship:

```python
self.db.connect_nodes(puid, uid, {"con_type": "parent"})

```

This graph structure enables efficient querying of parent-child relationships within the SQLite index.

### HDF5 Vault Support

The `Hdf5Vault` implementation similarly preserves parent links during serialization, maintaining the `!sdata_parent` metadata attribute within HDF5 attributes or datasets. Both vault implementations ensure that loaded objects retain their original hierarchical positions.

## Working with Child Objects

Beyond parent tracking, `sdata` maintains bidirectional hierarchical navigation through in-memory child collections.

### The Internal Group Dictionary

The `Base` class maintains an internal `group` dictionary storing child objects keyed by their UUIDs. The class exposes standard dictionary-style accessors including `keys()`, `values()`, and `items()` for iterating over children.

Children are added to this collection through `add_data()` (or `add_blob()` in specialized subclasses), which inserts the child object under its unique identifier. This architecture provides O(1) access to immediate children while maintaining the parent reference in the child's own metadata.

## Complete Working Examples

The following examples demonstrate practical usage of `sdata` hierarchical features:

```python
from sdata import Base, Data, sdata_factory
from sdata.iolib.vault import FileSystemVault

# Example 1: Creating parent-child relationships

parent = Base(name="Project", project=None)
child = Base(name="Experiment", parent=parent)

# Verify parent storage

assert child.metadata.get(Base.SDATA_PARENT_SNAME).value == parent.sname
print("Parent sname:", child.metadata.get(Base.SDATA_PARENT_SNAME).value)

# Example 2: Accessing parent via accessor

parent_suuid = child.get_parent()
print("Parent SUUID:", parent_suuid.sname)

# Example 3: Data copying with parent links

orig = Data(name="raw", uuid="38b26864e7794f5182d38459bab85842")
copy = orig.copy(name="raw_copy")
print("Copy's parent:", copy.metadata.get(Data.SDATA_PARENT).value)

# Example 4: Persistent storage with hierarchy

vault = FileSystemVault(rootpath="/tmp/myvault")
vault.dump_blob(orig)
vault.dump_blob(copy)

# Query parent edges in SQLite

parent_edges = vault.index.db.find_node(
    copy.metadata.get(Data.SDATA_PARENT).value
)
print("Parent edges:", parent_edges)

```

## Summary

- **Metadata-based storage**: `sdata` stores parent references using `SDATA_PARENT_SNAME` (sname format) in `Base` objects and `SDATA_PARENT` (UUID format) in `Data` objects.
- **Easy navigation**: The `get_parent()` method and `parent` property return `SUUID` objects for traversing upward in the hierarchy.
- **Automatic lineage**: The `Data.copy()` method automatically sets the source object as the parent of the new copy.
- **Persistent relationships**: Vault backends in [`sdata/iolib/vault.py`](https://github.com/lepy/sdata/blob/main/sdata/iolib/vault.py) store parent links in SQLite and HDF5 formats, creating explicit graph edges for querying.
- **Child management**: Objects maintain internal `group` dictionaries accessible via `keys()`, `values()`, and `items()` for downward hierarchy traversal.

## Frequently Asked Questions

### How does sdata store parent references?

`sdata` stores parent references in object metadata using distinct strategies: the `Base` class uses the `SDATA_PARENT_SNAME` key to store the parent's short SUUID name (sname), while the `Data` class uses `SDATA_PARENT` to store the raw UUID. Both approaches maintain lightweight references that avoid circular object dependencies.

### What is the difference between Base and Data parent handling?

The `Base` class in [`sdata/base.py`](https://github.com/lepy/sdata/blob/main/sdata/base.py) stores parent snames (short identifiers) via `SDATA_PARENT_SNAME`, suitable for general object hierarchies. The `Data` class in [`sdata/data.py`](https://github.com/lepy/sdata/blob/main/sdata/data.py) stores full UUIDs via `SDATA_PARENT`, optimized for vault persistence and data provenance tracking. Additionally, `Data` automatically establishes parent links during `copy()` operations, while `Base` requires explicit parent specification during initialization.

### Can parent relationships survive serialization?

Yes, parent relationships persist through serialization in `sdata`. The vault implementations in [`sdata/iolib/vault.py`](https://github.com/lepy/sdata/blob/main/sdata/iolib/vault.py) explicitly handle parent metadata: `VaultSqliteIndex` stores parent UUIDs in dedicated columns and creates graph edges via `connect_nodes()`, while `Hdf5Vault` preserves the `!sdata_parent` attribute. When objects are loaded from these stores, the parent metadata is restored, maintaining the hierarchical structure.

### How do I retrieve children of a specific object?

Child objects are accessed through the `Base` class dictionary interface methods: `keys()` returns child UUIDs, `values()` returns the child objects themselves, and `items()` returns UUID-object pairs. Children are added via `add_data()` (or `add_blob()` for binary data), which inserts them into the internal `group` dictionary indexed by their UUIDs.