Extending sdata with Custom IOLib Storage Backends: Architecture and Implementation Guide
Developers can extend sdata with custom storage backends by subclassing the Database abstract base class in sdata/iolib/db.py and implementing four core CRUD methods—insert, read, update, and delete—enabling seamless integration with any persistence layer from in-memory dictionaries to cloud object stores.
The sdata library provides a modular I/O abstraction layer (sdata.iolib) designed to decouple data-class objects from their underlying persistence mechanism. By adhering to a minimal abstract database API, you can implement custom IOLib storage backends without modifying existing business logic or consumer code. This guide explores the architecture of the storage layer and provides a complete implementation roadmap for creating bespoke backends.
Understanding the sdata IOLib Architecture
The storage architecture in sdata follows a layered design that separates domain objects from concrete storage implementations. At the foundation lies the abstract database API, which enforces a uniform contract across all backends.
The Database Abstract Base Class
The Database class in [sdata/iolib/db.py](https://github.com/lepy/sdata/blob/master/sdata/iolib/db.py) defines the contract that every storage backend must fulfill:
from abc import ABC, abstractmethod
class Database(ABC):
@abstractmethod
def insert(self, data): ...
@abstractmethod
def read(self, identifier): ...
@abstractmethod
def update(self, identifier, data): ...
@abstractmethod
def delete(self, identifier): ...
All built-in stores inherit from this class. The signatures are deliberately minimal, making it straightforward to implement alternative storage engines such as PostgreSQL, Redis, or cloud object stores. When a store is instantiated, it automatically creates the underlying schema if needed (see _ensure_table in the SQLite stores).
Built-in Backend Implementations
sdata ships with several concrete implementations that demonstrate best practices for different use cases:
JSONSQLiteStore([sdata/iolib/jsonsqlitestore.py](https://github.com/lepy/sdata/blob/master/sdata/iolib/jsonsqlitestore.py)): Stores JSON payloads in SQLite with optional Z-lib compression, dynamic index creation, pagination, TTL support, and hooks.JSON1SQLiteStore([sdata/iolib/json1sqlitestore.py](https://github.com/lepy/sdata/blob/master/sdata/iolib/json1sqlitestore.py)): Leverages the SQLite JSON1 extension for native JSON functions, generated columns for_sdata_*attributes, automatic index generation, bulk operations, and advanced query helpers likefind_byandreplace_path.Vault([sdata/iolib/vault.py](https://github.com/lepy/sdata/blob/master/sdata/iolib/vault.py)): A hierarchical, version-aware key/value store backed by either the local filesystem or a SQLite index.FlatHDFDataStore([sdata/iolib/hdf.py](https://github.com/lepy/sdata/blob/master/sdata/iolib/hdf.py)): HDF5-based persistence optimized for large array-like scientific data.
Each implementation exposes the same core CRUD methods plus utility functions such as exists, count, fetch_page, and transaction.
Implementing a Custom Storage Backend
Creating a custom backend involves subclassing Database and providing concrete implementations for the abstract methods. The following checklist ensures your implementation integrates seamlessly with the sdata ecosystem.
Step-by-Step Implementation Checklist
- Create a new module under
sdata/iolib/(e.g.,mybackend.py). - Subclass
Databaseand implement the four abstract methods. - Add optional helpers for indexing, bulk operations, or transaction support.
- Provide a context-manager
transactionusingtry/exceptfor rollback capability. - Register the class in
sdata/iolib/__init__.pyfor direct exposure. - Write unit tests under
tests/iolib/mirroring existing store test patterns.
Minimal Example: In-Memory Dictionary Store
The following example demonstrates a minimal backend using Python dictionaries, useful for prototyping or testing:
# sdata/iolib/mydictstore.py
from sdata.iolib.db import Database
from typing import Any, Dict, Optional
import uuid
from contextlib import contextmanager
class DictStore(Database):
"""Very simple in-memory store useful for prototyping or testing."""
def __init__(self) -> None:
self._data: Dict[uuid.UUID, Dict[str, Any]] = {}
def insert(self, data: Dict[str, Any]) -> uuid.UUID:
uid = uuid.uuid4()
self._data[uid] = data
return uid
def read(self, identifier: uuid.UUID) -> Optional[Dict[str, Any]]:
return self._data.get(identifier)
def update(self, identifier: uuid.UUID, data: Dict[str, Any]) -> None:
if identifier not in self._data:
raise KeyError(f"No record with id {identifier}")
self._data[identifier] = data
def delete(self, identifier: uuid.UUID) -> None:
self._data.pop(identifier, None)
@contextmanager
def transaction(self):
"""Optional transaction helper with rollback capability."""
snapshot = self._data.copy()
try:
yield
except Exception:
self._data = snapshot
raise
To expose this backend, add from .mydictstore import DictStore to sdata/iolib/__init__.py.
Advanced Example: Redis Backend
For production scenarios requiring distributed storage, you might implement a Redis backend:
# sdata/iolib/redisstore.py
import uuid
import json
import redis
from sdata.iolib.db import Database
class RedisStore(Database):
def __init__(self, url="redis://localhost:6379/0"):
self.r = redis.StrictRedis.from_url(url)
def _key(self, uid: uuid.UUID) -> str:
return f"sdata:{uid}"
def insert(self, data: dict) -> uuid.UUID:
uid = uuid.uuid4()
self.r.set(self._key(uid), json.dumps(data))
return uid
def read(self, identifier: uuid.UUID):
raw = self.r.get(self._key(identifier))
return json.loads(raw) if raw else None
def update(self, identifier: uuid.UUID, data: dict):
if not self.r.exists(self._key(identifier)):
raise KeyError(f"No record {identifier}")
self.r.set(self._key(identifier), json.dumps(data))
def delete(self, identifier: uuid.UUID):
self.r.delete(self._key(identifier))
Usage remains identical across all backends:
from sdata.iolib.redisstore import RedisStore
store = RedisStore()
uid = store.insert({"msg": "hello", "cnt": 1})
print(store.read(uid))
store.update(uid, {"msg": "world", "cnt": 2})
store.delete(uid)
Working with Built-in Stores
Before implementing a custom backend, review the built-in stores for patterns regarding schema initialization, indexing, and transaction handling.
CRUD Operations with JSON1SQLiteStore
The JSON1SQLiteStore demonstrates advanced features like generated columns and JSON path manipulation:
from sdata.iolib.json1sqlitestore import JSON1SQLiteStore
# Initialize (creates SQLite file and generated columns automatically)
store = JSON1SQLiteStore("mydata.db", index_keys=["user"])
# Insert returns the SQLite rowid
uid = store.insert({
"_sdata_class": "Person",
"_sdata_name": "bob",
"_sdata_suuid": "suuid-01",
"age": 30,
"tags": ["engineer", "python"]
})
# Retrieve record
record = store.get(uid)
# Update field
record["age"] = 31
store.update(uid, record)
# Query by generated column (utilizes index)
people = store.find_by("_sdata_name", "bob")
# Modify nested array element in-place using JSON path
store.replace_path(uid, "$.tags[0]", "senior engineer")
# Backup and restore
store.backup("backup.db")
store.restore("backup.db")
Key Source Files for Backend Development
The following files contain essential reference implementations for extending sdata:
sdata/iolib/db.py: Contains the abstractDatabaseclass defining the CRUD contract.sdata/iolib/__init__.py: Houses thePIDclass and public re-exports.sdata/iolib/jsonsqlitestore.py: Reference for compression, TTL, and hook implementations.sdata/iolib/json1sqlitestore.py: Demonstrates generated columns, bulk operations, and JSON1 path operations.sdata/iolib/vault.py: Example of hierarchical storage with versioning.tests/iolib/test_jsonsqlitestore.py: Test suite validating the backend contract.
Summary
- Minimal Contract: The
Databaseabstract base class requires only four methods (insert,read,update,delete) to create a functional backend. - Plug-and-Play Architecture: Custom backends integrate seamlessly because all consumers interact with the abstract API, not concrete implementations.
- Schema Initialization: Follow the pattern in existing stores by implementing
_ensure_tableor equivalent methods to handle automatic schema creation. - Transaction Support: Optional but recommended—implement a context-manager
transactionmethod with rollback capability for data integrity. - Testing: Mirror the test structure in
tests/iolib/to ensure your backend adheres to the expected contract.
Frequently Asked Questions
What methods are required to implement a custom sdata storage backend?
You must implement the four abstract methods defined in sdata/iolib/db.py: insert(self, data), read(self, identifier), update(self, identifier, data), and delete(self, identifier). These methods handle the basic CRUD operations. You may also add optional utility methods like exists, count, or transaction to match the feature set of built-in stores.
How do I add transaction support to a custom storage backend?
Implement a transaction method as a Python context manager that captures the state before operations and restores it if an exception occurs. For example, in an in-memory store, copy the internal dictionary before the yield statement and restore it in the except block. For database-backed stores, use the underlying connection's native transaction capabilities.
Can I use sdata with existing database systems like PostgreSQL or MongoDB?
Yes. By subclassing Database and mapping the CRUD methods to your target database's API, you can create a backend for any storage system. The Redis example above demonstrates this pattern for key-value stores, while SQLite implementations show SQL-based approaches. The public API remains identical regardless of the underlying technology.
Where should I place my custom backend code within the sdata project?
Create a new module file in the sdata/iolib/ directory (e.g., sdata/iolib/mycustomstore.py). To make it available as sdata.iolib.MyCustomStore, import the class in sdata/iolib/__init__.py. Place corresponding unit tests in tests/iolib/test_mycustomstore.py following the patterns in existing test files like test_jsonsqlitestore.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →