Implementing Data Versioning and Change Management Strategies with sdata

sdata provides a dual-layer versioning system that combines object-level metadata tracking with SQLite schema versioning, enabling robust change management through automated hooks and transactional migrations.

The sdata library embeds versioning capabilities directly into every data object and storage backend, allowing you to track provenance, manage schema evolution, and audit changes without external dependencies. By leveraging the Base class metadata system alongside the JSONSQLiteStore backend, you can implement comprehensive data versioning and change management strategies with sdata that satisfy rigorous scientific and production requirements.

Object-Level Versioning with Base Metadata

Every object derived from Base automatically captures the library version used at creation time. In sdata/base.py, the SDATA_VERSION constant (defined as "_sdata_version" on line 36) is injected into the object's metadata during initialization.


# sdata/base.py – lines 106-108

self.metadata.add(
    self.SDATA_VERSION, __version__, dtype="str",
    description="sdata package version", required=True
)

This metadata field persists through serialization, providing immutable provenance tracking. When you instantiate any Base object, the current sdata.__version__ is recorded automatically.

from sdata import Base

# Automatically records sdata version in metadata

obj = Base(name="experiment_01")
print(obj.metadata.get("_sdata_version").value)  # → "1.0.0"

Store-Level Schema Versioning

The JSONSQLiteStore class provides a mutable schema versioning mechanism using SQLite's user_version pragma. This enables you to version the database schema independently from individual data objects.

According to the source code in sdata/iolib/jsonsqlitestore.py, you can query and modify the schema version using dedicated methods:


# sdata/iolib/jsonsqlitestore.py – lines 475-480

def get_version(self) -> int:
    return self.conn.execute("PRAGMA user_version").fetchone()[0]

# lines 488-500

def set_version(self, version: int) -> None:
    self.conn.execute(f"PRAGMA user_version = {version}")

When you initialize a new store, the schema version defaults to 0. You can increment this value to track schema iterations across deployments.

from sdata.iolib.jsonsqlitestore import JSONSQLiteStore

store = JSONSQLiteStore("mydata.sqlite")
print(store.get_version())  # → 0

# Bump version after structural changes

store.set_version(1)

Schema Migrations and Evolution

Database schema evolution is handled through the migrate() method, which executes arbitrary SQL scripts within a single transaction. This approach allows you to add tables, indexes, or columns while maintaining version control.

The implementation in sdata/iolib/jsonsqlitestore.py (lines 501-512) uses executescript() for atomic migration execution:

def migrate(self, migration_sql: str) -> None:
    self.conn.executescript(migration_sql)

To implement a migration safely, combine the schema change with a version bump:

migration_sql = """
CREATE TABLE IF NOT EXISTS experiment_log (
    id INTEGER PRIMARY KEY,
    timestamp TEXT DEFAULT CURRENT_TIMESTAMP,
    action TEXT NOT NULL
);
"""

store.migrate(migration_sql)
store.set_version(2)

Change Tracking and Audit Hooks

sdata enables real-time change management through configurable hooks that fire on insert, update, and delete operations. These hooks allow you to implement audit trails, validation logic, or external notifications.

The hook system is implemented in sdata/iolib/jsonsqlitestore.py at three critical points:

  • Insert hook (lines 198-200): Invoked after store.insert() completes
  • Update hook (lines 113-115): Invoked after store.update() completes
  • Delete hook (lines 44-46): Invoked after store.delete() completes

Each hook receives the record ID and relevant payload:

def audit_log(event_type, record_id, data=None):
    print(f"[AUDIT] {event_type}: id={record_id}, data={data}")

store.hooks = {
    "on_insert": lambda rid, obj: audit_log("INSERT", rid, obj),
    "on_update": lambda rid, obj: audit_log("UPDATE", rid, obj),
    "on_delete": lambda rid: audit_log("DELETE", rid)
}

# Operations now trigger audit callbacks

rid = store.insert({"sample": "A"})
store.update(rid, {"sample": "B"})

Transactional Safety for Version Control

All mutating operations support atomic transactions through the transaction() context manager (lines 68-88 in sdata/iolib/jsonsqlitestore.py). This ensures that schema version bumps and data modifications succeed or fail together, preventing orphaned states.

with store.transaction():
    # Insert new experimental data

    rid = store.insert({"measurement": 42.0})
    
    # Only increment version if insert succeeds

    current_version = store.get_version()
    store.set_version(current_version + 1)
    
    # If any exception occurs here, both operations roll back

Summary

Implementing data versioning and change management strategies with sdata relies on three core architectural components:

  • Immutable object metadata: Every Base instance stores the creating library version in _sdata_version, providing automatic provenance tracking
  • Mutable schema versioning: JSONSQLiteStore exposes get_version() and set_version() methods to track database schema iterations via SQLite's user_version pragma
  • Hook-based auditing: The hooks dictionary supports on_insert, on_update, and on_delete callbacks for real-time change monitoring

Combined with the transaction() context manager for atomic operations and the migrate() method for schema evolution, these features provide a complete versioning toolkit for scientific data management.

Frequently Asked Questions

How does sdata track versions at the object level?

The Base class in sdata/base.py automatically injects a _sdata_version metadata attribute during instantiation (lines 106-108). This field stores the current sdata.__version__ string and persists through all serialization operations, creating an immutable audit trail of the library version used to create each data object.

What is the difference between object-level and store-level versioning?

Object-level versioning tracks the sdata library version used to create individual data objects through the _sdata_version metadata field. Store-level versioning tracks the database schema version using SQLite's user_version pragma, managed via JSONSQLiteStore.get_version() and set_version(). The former captures data provenance, while the latter enables schema migration workflows.

How can I implement audit logging for data changes?

Register callback functions in the store.hooks dictionary with keys on_insert, on_update, and on_delete. According to sdata/iolib/jsonsqlitestore.py, these hooks receive the record ID and modified object (for insert/update) or just the record ID (for delete), allowing you to write audit logs to external systems or secondary tables after each mutation.

How do I safely migrate schema versions without data loss?

Wrap schema modifications and version increments in a with store.transaction(): block. This context manager (lines 68-88) ensures atomic execution: use store.migrate() to execute SQL schema changes, then store.set_version() to bump the version number. If either operation fails, SQLite rolls back both changes, maintaining data integrity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →