Provenance Tracking with W3C PROV-O in Semantica: A Complete Implementation Guide
Semantica delivers a production-grade, W3C PROV-O-compliant provenance engine featuring dual storage backends, cryptographic hash-chain integrity, and native RDF export capabilities for full data lineage auditability.
The open-source Semantica framework (semantica-agi/semantica) implements a comprehensive provenance tracking system built entirely around the W3C PROV-O standard. This guide covers the three-layer architecture—from underlying data schemas to high-level management APIs—demonstrating how to implement provenance tracking with W3C PROV-O in Semantica for both research prototypes and production workloads.
Architecture Overview: Data Model, Storage, and Manager
Semantica’s provenance engine separates concerns across three distinct layers, each mapped to specific source files.
The PROV-O Data Model in schemas.py
The foundation resides in semantica/provenance/schemas.py, where the ProvenanceEntry dataclass maps directly to W3C PROV-O concepts:
prov:Entity→entity_id,entity_typeprov:Activity→activity_id,activity_started_at_time,activity_ended_at_timeprov:Agent→agent_id,agent_type,is_automated- Relationships →
parent_entity_id,derived_from_id,used_entities,previous_version_id - Integrity Metadata →
timestamp,checksum,invalidated
The schema also provides lightweight wrappers including SourceReference, PropertySource, AgentRecord, and ActivityRecord, each offering to_dict and from_dict helpers for serialization interoperability.
Pluggable Storage Backends in storage.py
The semantica/provenance/storage.py module defines an abstract ProvenanceStorage interface with two concrete implementations:
-
InMemoryStorage(lines 91–150): A dictionary-backed store ideal for unit tests and small workloads. It maintains entries in a Python dict, implements a hash chain for integrity verification, and provides BFS lineage tracing. -
SQLiteStorage(lines 91–1040): A production-grade backend featuring a fully W3C PROV-O-compliant schema, indexed lookups for fast lineage and descendant queries, and transactional safety.
Both expose identical APIs (store, retrieve, retrieve_all, trace_lineage, clear, transaction), enabling seamless swapping between development and production environments.
High-Level Orchestration in manager.py
The ProvenanceManager in semantica/provenance/manager.py serves as the façade, consolidating tracking, querying, and export operations. It normalizes agent and activity kwargs via _resolve_agent_kwargs and _resolve_activity_kwargs (lines 24–38), wraps operations in transaction savepoints, and handles batch optimizations.
Configuring Storage Backends
In-Memory Storage for Development
For rapid prototyping and testing, instantiate InMemoryStorage directly:
from semantica.provenance.storage import InMemoryStorage
from semantica.provenance.manager import ProvenanceManager
storage = InMemoryStorage()
prov_mgr = ProvenanceManager(storage=storage)
SQLite Storage for Production
For persistent, concurrent workloads, use SQLiteStorage with a file path:
from semantica.provenance.storage import SQLiteStorage
storage = SQLiteStorage("provenance.db")
prov_mgr = ProvenanceManager(storage=storage)
The SQLiteStorage implementation provides ACID guarantees through a transaction context manager, ensuring atomic provenance writes even under concurrent access.
Recording Provenance with the Manager API
Tracking Individual Entities
Use track_entity to record entity creation with full agent and activity attribution. The method accepts AgentRecord and ActivityRecord objects along with a role parameter specifying the relationship:
from semantica.provenance.schemas import AgentRecord, ActivityRecord
agent = AgentRecord(
id="reviewer_jane",
agent_type="person",
is_automated=False
)
activity = ActivityRecord(
id="ner_extraction",
activity_type="nlp_process",
started_at_time="2026-08-18T12:00:00",
ended_at_time="2026-08-18T12:00:02"
)
entry = prov_mgr.track_entity(
entity_id="entity_123",
source="DOI:10.1371/journal.pone.0023601",
metadata={"confidence": 0.94},
agent=agent,
activity=activity,
role="generator"
)
Batch Ingestion for Scale
For high-throughput scenarios, track_entities_batch and track_chunks_batch partition payloads into configurable chunks (default 1,000 entries) to mitigate SQLite lock contention and reduce transaction overhead:
entities = [
{"id": f"entity_{i}", "metadata": {"confidence": 0.9 - i * 0.01}}
for i in range(2500)
]
tracked = prov_mgr.track_entities_batch(entities, source="dataset_v1")
print(f"Tracked {tracked} entities")
Querying Lineage and Graph Traversal
Retrieve complete ancestry chains using get_lineage, which aggregates metadata from all ancestors, validates the hash-chain integrity, and returns a rich summary including source_documents and merged metadata:
lineage = prov_mgr.get_lineage("entity_123")
print(lineage["source_documents"]) # List of upstream source IDs
print(lineage["metadata"]) # Merged metadata from full ancestry
print(lineage["integrity_verified"]) # Boolean hash-chain status
For downstream analysis, get_descendants performs a reverse graph walk to identify all entities derived from a given source.
Exporting W3C PROV-O RDF
The export_prov method (lines 302–327 in manager.py) serializes the entire provenance store to standards-compliant RDF. It accepts a format parameter (turtle, jsonld, n-triples) and uses a configurable base URI (default https://semantica.dev/ns#), binding the prov namespace and a custom ex prefix:
# Export as Turtle
ttl = prov_mgr.export_prov(format="turtle")
with open("provenance.ttl", "w") as f:
f.write(ttl)
# Export as JSON-LD
jsonld = prov_mgr.export_prov(format="jsonld")
Advanced Integrity and Versioning Features
Cryptographic Hash-Chain Verification
Every stored entry receives a SHA-256 checksum computed via compute_checksum, linking to the previous_checksum of the prior entry. The manager automatically verifies this chain during lineage queries, setting integrity_verified to False if tampering is detected.
Invalidation and Tombstones
The invalidate method (lines 110–149) implements soft deletion through tombstones. It archives the pre-invalidated entry, creates a new entry marked as invalidated, and preserves the hash chain to maintain audit continuity:
prov_mgr.invalidate("entity_123", reason="data_retracted")
Versioning vs. Derivation Semantics
Semantica distinguishes between corrections and genuine derivations:
previous_version_id: Records an updated version of the same fact (correction).derived_from_id: Records a new entity genuinely derived from a different source.
Both relationships are exposed via the revision_history API, enabling precise temporal analysis of how data evolves versus how it transforms.
Summary
- Three-layer architecture separates data modeling (
schemas.py), persistence (storage.py), and orchestration (manager.py) for maintainable provenance implementations. - Dual storage backends allow seamless transition from
InMemoryStorage(development/testing) toSQLiteStorage(production) without API changes. - Batch operations default to 1,000-item chunks to optimize database performance while maintaining per-item savepoints for isolated rollback.
- Native W3C PROV-O export generates Turtle, JSON-LD, or N-Triples with proper namespace binding for immediate interoperability with RDF tools.
- SHA-256 hash chains provide tamper-evident audit trails, automatically verified during lineage reconstruction.
Frequently Asked Questions
What storage backends does Semantica support for provenance tracking?
Semantica provides two implementations of the abstract ProvenanceStorage interface: InMemoryStorage for ephemeral, high-speed development use, and SQLiteStorage for persistent, transactional production workloads. Both support identical lineage tracing and integrity features.
How does Semantica ensure the integrity of provenance records?
The engine computes a SHA-256 checksum for every ProvenanceEntry and links it to the previous entry's checksum, forming a cryptographic hash chain. During get_lineage queries, the manager validates this chain and reports integrity_verified status, detecting any unauthorized modifications.
Can I export provenance data to standard RDF formats?
Yes. The ProvenanceManager.export_prov() method supports W3C PROV-O serialization to Turtle, JSON-LD, and N-Triples formats. The export uses a configurable base URI and standard prov namespace prefixes for compatibility with external provenance tools.
What is the difference between previous_version_id and derived_from_id?
previous_version_id indicates a correction or update to the same entity (versioning), while derived_from_id indicates that a new entity was created by genuinely transforming or deriving from a different source entity. This distinction allows precise tracking of data evolution versus data transformation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →