Implementing Data Versioning and Change Management Strategies with sdata
sdata provides a dual-layer versioning system that combines object-level metadata tracking with SQLite schema versioning, enabling robust change management through automated hooks and transactional migrations.
The sdata library embeds versioning capabilities directly into every data object and storage backend, allowing you to track provenance, manage schema evolution, and audit changes without external dependencies. By leveraging the Base class metadata system alongside the JSONSQLiteStore backend, you can implement comprehensive data versioning and change management strategies with sdata that satisfy rigorous scientific and production requirements.
Object-Level Versioning with Base Metadata
Every object derived from Base automatically captures the library version used at creation time. In sdata/base.py, the SDATA_VERSION constant (defined as "_sdata_version" on line 36) is injected into the object's metadata during initialization.
# sdata/base.py – lines 106-108
self.metadata.add(
self.SDATA_VERSION, __version__, dtype="str",
description="sdata package version", required=True
)
This metadata field persists through serialization, providing immutable provenance tracking. When you instantiate any Base object, the current sdata.__version__ is recorded automatically.
from sdata import Base
# Automatically records sdata version in metadata
obj = Base(name="experiment_01")
print(obj.metadata.get("_sdata_version").value) # → "1.0.0"
Store-Level Schema Versioning
The JSONSQLiteStore class provides a mutable schema versioning mechanism using SQLite's user_version pragma. This enables you to version the database schema independently from individual data objects.
According to the source code in sdata/iolib/jsonsqlitestore.py, you can query and modify the schema version using dedicated methods:
# sdata/iolib/jsonsqlitestore.py – lines 475-480
def get_version(self) -> int:
return self.conn.execute("PRAGMA user_version").fetchone()[0]
# lines 488-500
def set_version(self, version: int) -> None:
self.conn.execute(f"PRAGMA user_version = {version}")
When you initialize a new store, the schema version defaults to 0. You can increment this value to track schema iterations across deployments.
from sdata.iolib.jsonsqlitestore import JSONSQLiteStore
store = JSONSQLiteStore("mydata.sqlite")
print(store.get_version()) # → 0
# Bump version after structural changes
store.set_version(1)
Schema Migrations and Evolution
Database schema evolution is handled through the migrate() method, which executes arbitrary SQL scripts within a single transaction. This approach allows you to add tables, indexes, or columns while maintaining version control.
The implementation in sdata/iolib/jsonsqlitestore.py (lines 501-512) uses executescript() for atomic migration execution:
def migrate(self, migration_sql: str) -> None:
self.conn.executescript(migration_sql)
To implement a migration safely, combine the schema change with a version bump:
migration_sql = """
CREATE TABLE IF NOT EXISTS experiment_log (
id INTEGER PRIMARY KEY,
timestamp TEXT DEFAULT CURRENT_TIMESTAMP,
action TEXT NOT NULL
);
"""
store.migrate(migration_sql)
store.set_version(2)
Change Tracking and Audit Hooks
sdata enables real-time change management through configurable hooks that fire on insert, update, and delete operations. These hooks allow you to implement audit trails, validation logic, or external notifications.
The hook system is implemented in sdata/iolib/jsonsqlitestore.py at three critical points:
- Insert hook (lines 198-200): Invoked after
store.insert()completes - Update hook (lines 113-115): Invoked after
store.update()completes - Delete hook (lines 44-46): Invoked after
store.delete()completes
Each hook receives the record ID and relevant payload:
def audit_log(event_type, record_id, data=None):
print(f"[AUDIT] {event_type}: id={record_id}, data={data}")
store.hooks = {
"on_insert": lambda rid, obj: audit_log("INSERT", rid, obj),
"on_update": lambda rid, obj: audit_log("UPDATE", rid, obj),
"on_delete": lambda rid: audit_log("DELETE", rid)
}
# Operations now trigger audit callbacks
rid = store.insert({"sample": "A"})
store.update(rid, {"sample": "B"})
Transactional Safety for Version Control
All mutating operations support atomic transactions through the transaction() context manager (lines 68-88 in sdata/iolib/jsonsqlitestore.py). This ensures that schema version bumps and data modifications succeed or fail together, preventing orphaned states.
with store.transaction():
# Insert new experimental data
rid = store.insert({"measurement": 42.0})
# Only increment version if insert succeeds
current_version = store.get_version()
store.set_version(current_version + 1)
# If any exception occurs here, both operations roll back
Summary
Implementing data versioning and change management strategies with sdata relies on three core architectural components:
- Immutable object metadata: Every
Baseinstance stores the creating library version in_sdata_version, providing automatic provenance tracking - Mutable schema versioning:
JSONSQLiteStoreexposesget_version()andset_version()methods to track database schema iterations via SQLite'suser_versionpragma - Hook-based auditing: The
hooksdictionary supportson_insert,on_update, andon_deletecallbacks for real-time change monitoring
Combined with the transaction() context manager for atomic operations and the migrate() method for schema evolution, these features provide a complete versioning toolkit for scientific data management.
Frequently Asked Questions
How does sdata track versions at the object level?
The Base class in sdata/base.py automatically injects a _sdata_version metadata attribute during instantiation (lines 106-108). This field stores the current sdata.__version__ string and persists through all serialization operations, creating an immutable audit trail of the library version used to create each data object.
What is the difference between object-level and store-level versioning?
Object-level versioning tracks the sdata library version used to create individual data objects through the _sdata_version metadata field. Store-level versioning tracks the database schema version using SQLite's user_version pragma, managed via JSONSQLiteStore.get_version() and set_version(). The former captures data provenance, while the latter enables schema migration workflows.
How can I implement audit logging for data changes?
Register callback functions in the store.hooks dictionary with keys on_insert, on_update, and on_delete. According to sdata/iolib/jsonsqlitestore.py, these hooks receive the record ID and modified object (for insert/update) or just the record ID (for delete), allowing you to write audit logs to external systems or secondary tables after each mutation.
How do I safely migrate schema versions without data loss?
Wrap schema modifications and version increments in a with store.transaction(): block. This context manager (lines 68-88) ensures atomic execution: use store.migrate() to execute SQL schema changes, then store.set_version() to bump the version number. If either operation fails, SQLite rolls back both changes, maintaining data integrity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →