# Provenance in Semantica Knowledge Graph: Immutable Data Lineage and Source Tracking

> Discover how Semantica's Knowledge Graph uses provenance for immutable data lineage and source tracking. Securely record origin, transformations, and integrity with SHA-256 hash chains and PROV-O.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-12

---

**Provenance in Semantica's Knowledge Graph serves as an immutable audit backbone that cryptographically records the origin, transformation history, and integrity of every entity, relationship, and property using SHA-256 hash chains and W3C PROV-O standards.**

The semantica-agi/semantica repository implements a comprehensive provenance layer that ensures every piece of knowledge remains traceable, verifiable, and defensible. By attaching cryptographic checksums and version-chaining entries, the system creates tamper-evident trails that satisfy stringent regulatory requirements while supporting forensic analysis and conflict resolution.

## Architectural Foundation of Provenance Tracking

At the core of Semantica's provenance system is the `ProvenanceManager` class, located in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py). This unified API consolidates knowledge graph provenance, chunk provenance, and source-property provenance into a single interface, eliminating duplication across extraction, reasoning, and visualization components.

The architecture relies on `SQLiteStorage` (defined in [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py)) configured with Write-Ahead Logging (WAL) mode for durability. Each `ProvenanceEntry` persisted to storage includes a SHA-256 checksum that incorporates the previous entry's checksum, forming an immutable hash chain within the `_save_entry` method. This cryptographic linking ensures that any post-write tampering is immediately detectable, satisfying requirements such as FDA 21 CFR Part 11 and Basel III.

Data classes in [`semantica/provenance/schemas.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/schemas.py) map directly to W3C PROV-O terms, defining structures like `ProvenanceEntry`, `SourceReference`, and lineage metadata that standardize how origin information is captured and queried.

## Recording Entity Origins and Transformations

When raw data enters the pipeline, the system captures comprehensive metadata about its source and processing context. The `track_entity()` method records the originating document, timestamp, activity identifier, and optional confidence scores.

```python
from semantica.provenance import ProvenanceManager

prov = ProvenanceManager(storage_path="kg_provenance.db")

entry = prov.track_entity(
    entity_id="cve-2024-3400",
    source="NVD_feed_2024-04-12",
    metadata={"cvss_score": 10.0, "vector": "AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H"},
    confidence=0.98,
    entity_type="vulnerability",
    activity_id="nvd_ingestion",
    source_location="CVE‑2024‑3400 JSON record",
)

```

For bulk operations, `track_entities_batch()` and `track_chunks_batch()` propagate typed parameters such as `agent_id` and `activity_id` into provenance records, ensuring that high-volume ingest pipelines maintain complete audit trails without performance degradation.

## Multi-Source Property Attribution

Complex knowledge graphs often encounter conflicting information about the same property from different sources. The `track_property_source()` method creates separate provenance records for each source, enabling precise conflict detection and resolution.

```python
from semantica.provenance.schemas import SourceReference

nvd_ref = SourceReference(
    document="NVD_feed_2024-04-12",
    section="cvssMetricV31",
    confidence=0.98,
    metadata={"publisher": "NIST"}
)

commercial_ref = SourceReference(
    document="commercial_feed_2024-04-12",
    section="cvss_assessment",
    confidence=0.91,
    metadata={"publisher": "ThreatFeed-Co"}
)

prov.track_property_source(
    entity_id="cve-2024-3400",
    property_name="cvss_score",
    value=10.0,
    source=nvd_ref
)

prov.track_property_source(
    entity_id="cve-2024-3400",
    property_name="cvss_score",
    value=9.8,
    source=commercial_ref
)

```

This granular tracking stores each source's confidence level and document location, allowing downstream systems to apply source-priority rules or merge strategies based on verifiable provenance data.

## Querying Data Lineage and Descendants

The provenance system supports sophisticated lineage queries that reconstruct how facts evolved over time. The `get_lineage()` method in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) walks upstream ancestors via `parent_entity_id` and `used_entities` relationships, aggregating all source documents and transformation history.

```python
lineage = prov.get_lineage("cve-2024-3400")

print("Sources seen:", lineage["source_documents"])
print("History depth:", lineage["entity_count"])

for entry in lineage["lineage_chain"]:
    print(f"[{entry['timestamp'][:19]}] agent={entry['agent_id']} source={entry['source_document']}")

```

For downstream impact analysis, `trace_descendants()` identifies all entities derived from a specific source, while `get_all_sources()` provides flattened lists of origin documents filtered by depth or time range. These capabilities answer critical compliance questions such as "how did this fact evolve?" and "which entities depend on this corrupted source?"

## Cryptographic Integrity and Version Control

Each provenance entry maintains version integrity through `previous_version_id` and `derived_from_id` fields. When `track_entity()` is called with an existing `entity_id`, the system automatically archives the previous entry and links the new version into the chain. The SHA-256 checksum calculation in `_save_entry` incorporates the previous entry's hash, creating a blockchain-like verification mechanism that detects any unauthorized modifications to the audit trail.

## Standards-Compliant Export

Semantica implements full W3C PROV-O compatibility for interoperability with external audit systems. The `export_prov()` method serializes the entire provenance store or specific subgraphs into RDF formats.

```python
rdf_turtle = prov.export_prov(format="turtle")

with open("kg_provenance.ttl", "w") as f:
    f.write(rdf_turtle)

```

The export uses the default base URI `https://semantica.dev/ns#` (as defined in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) lines 382-388), generating standard PROV-O RDF that regulators and third-party verification tools can consume directly. The system also provides `audit_log()` and CLI commands in the `provenance` group for generating human-readable tamper-evident reports.

## Summary

- **ProvenanceManager** in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) provides a unified API for tracking entities, relationships, chunks, and properties with full lineage.
- **Cryptographic verification** uses SHA-256 hash chains linking `ProvenanceEntry` records to prevent tampering and ensure regulatory compliance.
- **Multi-source tracking** via `track_property_source()` and `SourceReference` enables conflict resolution by recording competing claims from different origins.
- **Lineage queries** through `get_lineage()` and `trace_descendants()` support forensic analysis and impact assessment via `parent_entity_id` and `used_entities` traversal.
- **Standards compliance** is achieved through W3C PROV-O RDF export using `export_prov()`, supporting audit requirements for FDA 21 CFR Part 11 and Basel III.

## Frequently Asked Questions

### What is the primary role of provenance in Semantica's Knowledge Graph?

Provenance serves as the immutable audit backbone that records where every entity, relationship, and property originates, how it transforms, and who performed each operation. It enables complete traceability and verifiability required for regulatory compliance, forensic analysis, and trustworthy AI pipelines by maintaining cryptographic hash chains of all data operations.

### How does Semantica verify the integrity of provenance records?

The system generates SHA-256 checksums for each `ProvenanceEntry` that incorporate the previous entry's checksum, forming an unbroken hash chain within the `_save_entry` method in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py). This cryptographic linking ensures that any post-write modification to the SQLite storage would invalidate the chain, making tampering immediately detectable during integrity audits.

### Can Semantica export provenance data for external regulatory audits?

Yes. The `export_prov()` method serializes provenance graphs into W3C PROV-O RDF format (Turtle, XML, or JSON-LD), using the default namespace `https://semantica.dev/ns#`. This standard-compliant output can be consumed directly by regulatory tools, auditors, and third-party verification systems to validate data lineage and processing history without accessing the internal database.

### How does batch processing maintain provenance without performance loss?

The `track_entities_batch()` and `track_chunks_batch()` methods in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) efficiently propagate provenance parameters such as `agent_id`, `activity_id`, and `source` into bulk records. These operations use optimized SQLite transactions with WAL mode to ensure that high-volume ingest pipelines generate complete audit trails without creating storage bottlenecks or duplicate entries.