# Provenance Tracking with W3C PROV-O in Semantica: A Complete Implementation Guide

> Implement W3C PROV-O provenance tracking with Semantica. Explore dual storage, cryptographic integrity, and RDF export for complete data lineage auditability. Get the complete guide.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Semantica delivers a production-grade, W3C PROV-O-compliant provenance engine featuring dual storage backends, cryptographic hash-chain integrity, and native RDF export capabilities for full data lineage auditability.**

The open-source Semantica framework (`semantica-agi/semantica`) implements a comprehensive provenance tracking system built entirely around the W3C PROV-O standard. This guide covers the three-layer architecture—from underlying data schemas to high-level management APIs—demonstrating how to implement provenance tracking with W3C PROV-O in Semantica for both research prototypes and production workloads.

## Architecture Overview: Data Model, Storage, and Manager

Semantica’s provenance engine separates concerns across three distinct layers, each mapped to specific source files.

### The PROV-O Data Model in schemas.py

The foundation resides in [`semantica/provenance/schemas.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/schemas.py), where the `ProvenanceEntry` dataclass maps directly to W3C PROV-O concepts:

- **`prov:Entity`** → `entity_id`, `entity_type`
- **`prov:Activity`** → `activity_id`, `activity_started_at_time`, `activity_ended_at_time`
- **`prov:Agent`** → `agent_id`, `agent_type`, `is_automated`
- **Relationships** → `parent_entity_id`, `derived_from_id`, `used_entities`, `previous_version_id`
- **Integrity Metadata** → `timestamp`, `checksum`, `invalidated`

The schema also provides lightweight wrappers including `SourceReference`, `PropertySource`, `AgentRecord`, and `ActivityRecord`, each offering `to_dict` and `from_dict` helpers for serialization interoperability.

### Pluggable Storage Backends in storage.py

The [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py) module defines an abstract `ProvenanceStorage` interface with two concrete implementations:

- **`InMemoryStorage`** (lines 91–150): A dictionary-backed store ideal for unit tests and small workloads. It maintains entries in a Python dict, implements a hash chain for integrity verification, and provides BFS lineage tracing.

- **`SQLiteStorage`** (lines 91–1040): A production-grade backend featuring a fully W3C PROV-O-compliant schema, indexed lookups for fast lineage and descendant queries, and transactional safety.

Both expose identical APIs (`store`, `retrieve`, `retrieve_all`, `trace_lineage`, `clear`, `transaction`), enabling seamless swapping between development and production environments.

### High-Level Orchestration in manager.py

The `ProvenanceManager` in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) serves as the façade, consolidating tracking, querying, and export operations. It normalizes agent and activity kwargs via `_resolve_agent_kwargs` and `_resolve_activity_kwargs` (lines 24–38), wraps operations in transaction savepoints, and handles batch optimizations.

## Configuring Storage Backends

### In-Memory Storage for Development

For rapid prototyping and testing, instantiate `InMemoryStorage` directly:

```python
from semantica.provenance.storage import InMemoryStorage
from semantica.provenance.manager import ProvenanceManager

storage = InMemoryStorage()
prov_mgr = ProvenanceManager(storage=storage)

```

### SQLite Storage for Production

For persistent, concurrent workloads, use `SQLiteStorage` with a file path:

```python
from semantica.provenance.storage import SQLiteStorage

storage = SQLiteStorage("provenance.db")
prov_mgr = ProvenanceManager(storage=storage)

```

The `SQLiteStorage` implementation provides ACID guarantees through a `transaction` context manager, ensuring atomic provenance writes even under concurrent access.

## Recording Provenance with the Manager API

### Tracking Individual Entities

Use `track_entity` to record entity creation with full agent and activity attribution. The method accepts `AgentRecord` and `ActivityRecord` objects along with a `role` parameter specifying the relationship:

```python
from semantica.provenance.schemas import AgentRecord, ActivityRecord

agent = AgentRecord(
    id="reviewer_jane",
    agent_type="person",
    is_automated=False
)

activity = ActivityRecord(
    id="ner_extraction",
    activity_type="nlp_process",
    started_at_time="2026-08-18T12:00:00",
    ended_at_time="2026-08-18T12:00:02"
)

entry = prov_mgr.track_entity(
    entity_id="entity_123",
    source="DOI:10.1371/journal.pone.0023601",
    metadata={"confidence": 0.94},
    agent=agent,
    activity=activity,
    role="generator"
)

```

### Batch Ingestion for Scale

For high-throughput scenarios, `track_entities_batch` and `track_chunks_batch` partition payloads into configurable chunks (default **1,000** entries) to mitigate SQLite lock contention and reduce transaction overhead:

```python
entities = [
    {"id": f"entity_{i}", "metadata": {"confidence": 0.9 - i * 0.01}}
    for i in range(2500)
]

tracked = prov_mgr.track_entities_batch(entities, source="dataset_v1")
print(f"Tracked {tracked} entities")

```

## Querying Lineage and Graph Traversal

Retrieve complete ancestry chains using `get_lineage`, which aggregates metadata from all ancestors, validates the hash-chain integrity, and returns a rich summary including `source_documents` and merged metadata:

```python
lineage = prov_mgr.get_lineage("entity_123")
print(lineage["source_documents"])  # List of upstream source IDs

print(lineage["metadata"])          # Merged metadata from full ancestry

print(lineage["integrity_verified"])  # Boolean hash-chain status

```

For downstream analysis, `get_descendants` performs a reverse graph walk to identify all entities derived from a given source.

## Exporting W3C PROV-O RDF

The `export_prov` method (lines 302–327 in [`manager.py`](https://github.com/semantica-agi/semantica/blob/main/manager.py)) serializes the entire provenance store to standards-compliant RDF. It accepts a `format` parameter (`turtle`, `jsonld`, `n-triples`) and uses a configurable base URI (default `https://semantica.dev/ns#`), binding the `prov` namespace and a custom `ex` prefix:

```python

# Export as Turtle

ttl = prov_mgr.export_prov(format="turtle")
with open("provenance.ttl", "w") as f:
    f.write(ttl)

# Export as JSON-LD

jsonld = prov_mgr.export_prov(format="jsonld")

```

## Advanced Integrity and Versioning Features

### Cryptographic Hash-Chain Verification

Every stored entry receives a SHA-256 checksum computed via `compute_checksum`, linking to the `previous_checksum` of the prior entry. The manager automatically verifies this chain during lineage queries, setting `integrity_verified` to `False` if tampering is detected.

### Invalidation and Tombstones

The `invalidate` method (lines 110–149) implements soft deletion through tombstones. It archives the pre-invalidated entry, creates a new entry marked as `invalidated`, and preserves the hash chain to maintain audit continuity:

```python
prov_mgr.invalidate("entity_123", reason="data_retracted")

```

### Versioning vs. Derivation Semantics

Semantica distinguishes between corrections and genuine derivations:

- **`previous_version_id`**: Records an updated version of the same fact (correction).
- **`derived_from_id`**: Records a new entity genuinely derived from a different source.

Both relationships are exposed via the `revision_history` API, enabling precise temporal analysis of how data evolves versus how it transforms.

## Summary

- **Three-layer architecture** separates data modeling ([`schemas.py`](https://github.com/semantica-agi/semantica/blob/main/schemas.py)), persistence ([`storage.py`](https://github.com/semantica-agi/semantica/blob/main/storage.py)), and orchestration ([`manager.py`](https://github.com/semantica-agi/semantica/blob/main/manager.py)) for maintainable provenance implementations.
- **Dual storage backends** allow seamless transition from `InMemoryStorage` (development/testing) to `SQLiteStorage` (production) without API changes.
- **Batch operations** default to 1,000-item chunks to optimize database performance while maintaining per-item savepoints for isolated rollback.
- **Native W3C PROV-O export** generates Turtle, JSON-LD, or N-Triples with proper namespace binding for immediate interoperability with RDF tools.
- **SHA-256 hash chains** provide tamper-evident audit trails, automatically verified during lineage reconstruction.

## Frequently Asked Questions

### What storage backends does Semantica support for provenance tracking?

Semantica provides two implementations of the abstract `ProvenanceStorage` interface: `InMemoryStorage` for ephemeral, high-speed development use, and `SQLiteStorage` for persistent, transactional production workloads. Both support identical lineage tracing and integrity features.

### How does Semantica ensure the integrity of provenance records?

The engine computes a SHA-256 checksum for every `ProvenanceEntry` and links it to the previous entry's checksum, forming a cryptographic hash chain. During `get_lineage` queries, the manager validates this chain and reports `integrity_verified` status, detecting any unauthorized modifications.

### Can I export provenance data to standard RDF formats?

Yes. The `ProvenanceManager.export_prov()` method supports W3C PROV-O serialization to Turtle, JSON-LD, and N-Triples formats. The export uses a configurable base URI and standard `prov` namespace prefixes for compatibility with external provenance tools.

### What is the difference between `previous_version_id` and `derived_from_id`?

`previous_version_id` indicates a correction or update to the same entity (versioning), while `derived_from_id` indicates that a new entity was created by genuinely transforming or deriving from a different source entity. This distinction allows precise tracking of data evolution versus data transformation.