Provenance in Semantica Knowledge Graph: Immutable Data Lineage and Source Tracking
Provenance in Semantica's Knowledge Graph serves as an immutable audit backbone that cryptographically records the origin, transformation history, and integrity of every entity, relationship, and property using SHA-256 hash chains and W3C PROV-O standards.
The semantica-agi/semantica repository implements a comprehensive provenance layer that ensures every piece of knowledge remains traceable, verifiable, and defensible. By attaching cryptographic checksums and version-chaining entries, the system creates tamper-evident trails that satisfy stringent regulatory requirements while supporting forensic analysis and conflict resolution.
Architectural Foundation of Provenance Tracking
At the core of Semantica's provenance system is the ProvenanceManager class, located in semantica/provenance/manager.py. This unified API consolidates knowledge graph provenance, chunk provenance, and source-property provenance into a single interface, eliminating duplication across extraction, reasoning, and visualization components.
The architecture relies on SQLiteStorage (defined in semantica/provenance/storage.py) configured with Write-Ahead Logging (WAL) mode for durability. Each ProvenanceEntry persisted to storage includes a SHA-256 checksum that incorporates the previous entry's checksum, forming an immutable hash chain within the _save_entry method. This cryptographic linking ensures that any post-write tampering is immediately detectable, satisfying requirements such as FDA 21 CFR Part 11 and Basel III.
Data classes in semantica/provenance/schemas.py map directly to W3C PROV-O terms, defining structures like ProvenanceEntry, SourceReference, and lineage metadata that standardize how origin information is captured and queried.
Recording Entity Origins and Transformations
When raw data enters the pipeline, the system captures comprehensive metadata about its source and processing context. The track_entity() method records the originating document, timestamp, activity identifier, and optional confidence scores.
from semantica.provenance import ProvenanceManager
prov = ProvenanceManager(storage_path="kg_provenance.db")
entry = prov.track_entity(
entity_id="cve-2024-3400",
source="NVD_feed_2024-04-12",
metadata={"cvss_score": 10.0, "vector": "AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H"},
confidence=0.98,
entity_type="vulnerability",
activity_id="nvd_ingestion",
source_location="CVE‑2024‑3400 JSON record",
)
For bulk operations, track_entities_batch() and track_chunks_batch() propagate typed parameters such as agent_id and activity_id into provenance records, ensuring that high-volume ingest pipelines maintain complete audit trails without performance degradation.
Multi-Source Property Attribution
Complex knowledge graphs often encounter conflicting information about the same property from different sources. The track_property_source() method creates separate provenance records for each source, enabling precise conflict detection and resolution.
from semantica.provenance.schemas import SourceReference
nvd_ref = SourceReference(
document="NVD_feed_2024-04-12",
section="cvssMetricV31",
confidence=0.98,
metadata={"publisher": "NIST"}
)
commercial_ref = SourceReference(
document="commercial_feed_2024-04-12",
section="cvss_assessment",
confidence=0.91,
metadata={"publisher": "ThreatFeed-Co"}
)
prov.track_property_source(
entity_id="cve-2024-3400",
property_name="cvss_score",
value=10.0,
source=nvd_ref
)
prov.track_property_source(
entity_id="cve-2024-3400",
property_name="cvss_score",
value=9.8,
source=commercial_ref
)
This granular tracking stores each source's confidence level and document location, allowing downstream systems to apply source-priority rules or merge strategies based on verifiable provenance data.
Querying Data Lineage and Descendants
The provenance system supports sophisticated lineage queries that reconstruct how facts evolved over time. The get_lineage() method in semantica/provenance/manager.py walks upstream ancestors via parent_entity_id and used_entities relationships, aggregating all source documents and transformation history.
lineage = prov.get_lineage("cve-2024-3400")
print("Sources seen:", lineage["source_documents"])
print("History depth:", lineage["entity_count"])
for entry in lineage["lineage_chain"]:
print(f"[{entry['timestamp'][:19]}] agent={entry['agent_id']} source={entry['source_document']}")
For downstream impact analysis, trace_descendants() identifies all entities derived from a specific source, while get_all_sources() provides flattened lists of origin documents filtered by depth or time range. These capabilities answer critical compliance questions such as "how did this fact evolve?" and "which entities depend on this corrupted source?"
Cryptographic Integrity and Version Control
Each provenance entry maintains version integrity through previous_version_id and derived_from_id fields. When track_entity() is called with an existing entity_id, the system automatically archives the previous entry and links the new version into the chain. The SHA-256 checksum calculation in _save_entry incorporates the previous entry's hash, creating a blockchain-like verification mechanism that detects any unauthorized modifications to the audit trail.
Standards-Compliant Export
Semantica implements full W3C PROV-O compatibility for interoperability with external audit systems. The export_prov() method serializes the entire provenance store or specific subgraphs into RDF formats.
rdf_turtle = prov.export_prov(format="turtle")
with open("kg_provenance.ttl", "w") as f:
f.write(rdf_turtle)
The export uses the default base URI https://semantica.dev/ns# (as defined in semantica/provenance/manager.py lines 382-388), generating standard PROV-O RDF that regulators and third-party verification tools can consume directly. The system also provides audit_log() and CLI commands in the provenance group for generating human-readable tamper-evident reports.
Summary
- ProvenanceManager in
semantica/provenance/manager.pyprovides a unified API for tracking entities, relationships, chunks, and properties with full lineage. - Cryptographic verification uses SHA-256 hash chains linking
ProvenanceEntryrecords to prevent tampering and ensure regulatory compliance. - Multi-source tracking via
track_property_source()andSourceReferenceenables conflict resolution by recording competing claims from different origins. - Lineage queries through
get_lineage()andtrace_descendants()support forensic analysis and impact assessment viaparent_entity_idandused_entitiestraversal. - Standards compliance is achieved through W3C PROV-O RDF export using
export_prov(), supporting audit requirements for FDA 21 CFR Part 11 and Basel III.
Frequently Asked Questions
What is the primary role of provenance in Semantica's Knowledge Graph?
Provenance serves as the immutable audit backbone that records where every entity, relationship, and property originates, how it transforms, and who performed each operation. It enables complete traceability and verifiability required for regulatory compliance, forensic analysis, and trustworthy AI pipelines by maintaining cryptographic hash chains of all data operations.
How does Semantica verify the integrity of provenance records?
The system generates SHA-256 checksums for each ProvenanceEntry that incorporate the previous entry's checksum, forming an unbroken hash chain within the _save_entry method in semantica/provenance/manager.py. This cryptographic linking ensures that any post-write modification to the SQLite storage would invalidate the chain, making tampering immediately detectable during integrity audits.
Can Semantica export provenance data for external regulatory audits?
Yes. The export_prov() method serializes provenance graphs into W3C PROV-O RDF format (Turtle, XML, or JSON-LD), using the default namespace https://semantica.dev/ns#. This standard-compliant output can be consumed directly by regulatory tools, auditors, and third-party verification systems to validate data lineage and processing history without accessing the internal database.
How does batch processing maintain provenance without performance loss?
The track_entities_batch() and track_chunks_batch() methods in semantica/provenance/manager.py efficiently propagate provenance parameters such as agent_id, activity_id, and source into bulk records. These operations use optimized SQLite transactions with WAL mode to ensure that high-volume ingest pipelines generate complete audit trails without creating storage bottlenecks or duplicate entries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →