How Semantica Tracks Provenance in the Knowledge Graph
Semantica tracks provenance in the knowledge graph by maintaining an immutable audit trail that records the source origin, UTC timestamp, and metadata for every entity, enabling complete lineage tracing from any graph element back to its original document or API call.
The semantica-agi/semantica repository implements a robust provenance subsystem that ensures every node, edge, and fact inserted into the knowledge graph carries verifiable origin data. By mapping entity identifiers to structured provenance entries, the system supports compliance requirements, temporal reasoning, and debugging of knowledge extraction pipelines.
Core Provenance Architecture
Semantica provides two primary interfaces for provenance tracking: the deprecated ProvenanceTracker class and the modern ProvenanceManager. Both components rely on the semantica/provenance/storage.py backends to persist records, but the manager offers a higher-level API that integrates directly with the knowledge graph construction pipeline.
The Legacy ProvenanceTracker Class
Located in semantica/kg/provenance_tracker.py, the ProvenanceTracker class maintains an in-memory dictionary (self._records) that maps entity IDs to lists of provenance entries. Each entry stores the source (file path, URL, or upstream identifier), a recorded_at timestamp generated via datetime.now(timezone.utc), and optional caller-supplied metadata such as extractor type or author identity.
from semantica.kg.provenance_tracker import ProvenanceTracker
tracker = ProvenanceTracker()
tracker.track_entity(
entity_id="E1",
source="https://github.com/example/repo",
metadata={"type": "github_repo", "author": "alice"},
)
sources = tracker.get_all_sources("E1")
print(sources)
# [{'source': 'https://github.com/example/repo',
# 'recorded_at': '2024-03-15T12:34:56.789012+00:00',
# 'type': 'github_repo', 'author': 'alice'}]
The Modern ProvenanceManager
The preferred interface resides in semantica/provenance/manager.py. The ProvenanceManager wraps storage backends (defined in semantica/provenance/storage.py) such as InMemoryStorage or SQLiteStorage, providing durability guarantees while exposing methods like revision_history() and export_audit_log().
from semantica.provenance.manager import ProvenanceManager
prov = ProvenanceManager()
prov.track_entity("E42", "https://arxiv.org/abs/2101.00001")
history = prov.revision_history("E42")
print(history)
Provenance Data Structure
Every provenance entry follows a strict schema defined in semantica/provenance/schemas.py. The system captures three essential fields to ensure comprehensive auditability:
- source: The authoritative origin of the entity (document filename, URL, or database ID)
- recorded_at: An ISO 8601 UTC timestamp generated using
datetime.now(timezone.utc) - metadata: An optional dictionary for domain-specific context (extractor version, confidence scores, author)
This structure enables temporal reasoning queries that determine what was true at specific points in time, as well as traceability back to the exact document or API call that generated a fact.
Automatic Provenance Injection During Graph Construction
When building the knowledge graph, the GraphBuilderWithProvenance (implemented in semantica/kg/graph_builder.py) automatically invokes the provenance tracker upon entity creation. This ensures no node or edge enters the graph without associated lineage metadata.
# Simplified excerpt from semantica/kg/graph_builder.py
if self.provenance:
self.provenance_tracker.track_entity(
entity_id=node.id,
source=source_id,
metadata={"extractor": extractor_name},
)
This automated injection guarantees that downstream consumers can trace any fact to its originating source document, supporting use cases such as regulatory compliance verification and source conflict resolution.
Querying and Auditing Provenance Records
The provenance subsystem exposes several query interfaces for retrieving lineage information and generating compliance reports.
Retrieving Source Lineage
Call get_all_sources(entity_id) to fetch the complete list of provenance entries for a specific entity. This returns chronologically ordered records showing every source that contributed to the entity's current state.
sources = prov.get_all_sources("E42")
for record in sources:
print(f"Source: {record['source']}, Recorded: {record['recorded_at']}")
Temporal Queries and Audit Exports
The query_recorded_between(start, end) method returns all records whose timestamps fall within a specified UTC interval, supporting "what-was-true-when" analyses. For compliance workflows, export_audit_log() serializes provenance data for a set of entity IDs as JSON or CSV.
from datetime import datetime, timezone
start = datetime(2024, 1, 1, tzinfo=timezone.utc)
end = datetime(2024, 12, 31, tzinfo=timezone.utc)
records = prov.query_recorded_between(start, end)
# Export for external auditing
json_log = prov.export_audit_log(["E42", "E99"], format="json")
Storage Backends and Durability
The semantica/provenance/storage.py module defines pluggable storage backends that guarantee durability while optimizing for lookup performance. Available implementations include:
- InMemoryStorage: High-speed volatile storage suitable for testing and short-lived pipelines
- SQLiteStorage: Persistent disk-based storage for production knowledge graphs requiring durability across restarts
These backends power the ProvenanceManager, ensuring that provenance records survive system restarts while maintaining fast retrieval times for graph visualization tools and CLI queries.
Summary
- Semantica tracks provenance in the knowledge graph using a dedicated subsystem that records source, timestamp, and metadata for every entity.
- The ProvenanceManager (and legacy ProvenanceTracker) in
semantica/provenance/manager.pyandsemantica/kg/provenance_tracker.pyprovide the core API for recording lineage. - Automatic injection during graph construction ensures every node and edge carries origin metadata without manual intervention.
- Temporal queries via
query_recorded_between()and audit exports viaexport_audit_log()support compliance and debugging workflows. - Pluggable storage backends in
semantica/provenance/storage.pyoffer both in-memory speed and persistent durability.
Frequently Asked Questions
What is the difference between ProvenanceTracker and ProvenanceManager?
The ProvenanceTracker class in semantica/kg/provenance_tracker.py is the legacy implementation that maintains records in a simple in-memory dictionary. The ProvenanceManager in semantica/provenance/manager.py is the modern, preferred interface that abstracts storage backends (SQLite, in-memory) and provides additional methods like revision_history() and export_audit_log(). New projects should use ProvenanceManager for better durability and API stability.
How do I export a provenance audit log in Semantica?
Use the export_audit_log() method on a ProvenanceManager instance. Pass a list of entity IDs and specify the format ("json" or "csv"). This method serializes the complete provenance records—including source, timestamps, and metadata—into a portable format suitable for external compliance audits or regulatory reporting.
Can I query provenance records by time range?
Yes. The ProvenanceManager.query_recorded_between(start_datetime, end_datetime) method accepts timezone-aware datetime objects (recommended: datetime.now(timezone.utc)) and returns all provenance entries recorded within that interval. This enables temporal reasoning queries to determine the state of knowledge at specific historical moments.
How does Semantica automatically track provenance when building a knowledge graph?
During graph construction, the GraphBuilderWithProvenance (in semantica/kg/graph_builder.py) automatically calls track_entity() for every node and edge it creates, passing the source document ID and extractor metadata. This ensures the knowledge graph builder populates provenance data transparently without requiring manual instrumentation from developers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →