How to Track Source Provenance for Entities in Semantica: A Complete Guide
Semantica records the origin of every knowledge-graph entity through the ProvenanceManager class in semantica/provenance/manager.py, which stores source identifiers, confidence scores, and full lineage chains with cryptographic integrity checks.
The semantica-agi/semantica repository provides a robust provenance tracking system that allows you to track source provenance for entities in Semantica with fine-grained metadata and audit capabilities. Whether you are ingesting scientific papers, database records, or curated annotations, the provenance manager creates an immutable trail linking each entity to its original source document.
Initializing the ProvenanceManager
The ProvenanceManager class serves as the primary interface for all provenance operations in semantica/provenance/manager.py. By default, it uses an in-memory store suitable for development or testing. For production workloads, instantiate the manager with a storage_path to enable persistent SQLite storage.
from semantica.provenance.manager import ProvenanceManager
# In-memory storage (default)
prov = ProvenanceManager()
# Persistent SQLite storage
prov = ProvenanceManager(storage_path="provenance.db")
The manager implements the same public API previously provided by the older kg.ProvenanceTracker, but adds support for batch operations, richer metadata schemas, and cryptographic integrity verification.
Tracking Individual Entity Provenance
To track source provenance for a single entity, call the track_entity method defined at lines 66‑78 of semantica/provenance/manager.py. This method requires an entity_id and source parameter, with optional metadata and typed keyword arguments that populate dedicated database columns.
Key parameters:
entity_id– The knowledge graph identifier to tracksource– A stable source identifier (DOI, URL, or file path)metadata– Free-form dictionary for arbitrary contextconfidence,agent_id,activity_id– Typed kwargs stored in dedicated columns
from semantica.provenance.manager import ProvenanceManager
prov = ProvenanceManager()
prov.track_entity(
entity_id="E123",
source="doi:10.1038/nature12345",
metadata={"label": "Gene", "confidence": 0.98},
confidence=0.98,
agent_id="curation_bot",
activity_id="entity_ingest"
)
The implementation separates structured provenance data from flexible metadata, enabling efficient querying while preserving extensibility.
Querying Provenance and Lineage
The manager provides two primary methods to retrieve provenance information from semantica/provenance/manager.py.
get_all_sources(entity_id) (lines 90‑100) returns a flat list of source records containing the document reference, location, timestamp, and confidence score.
get_lineage(entity_id) (lines 753‑770) constructs a comprehensive lineage dictionary that includes the full ancestry chain, aggregated metadata across all sources, and integrity verification status.
# Retrieve all source records
sources = prov.get_all_sources("E123")
print("Sources:", sources)
# Retrieve rich lineage with ancestry chain
lineage = prov.get_lineage("E123")
print("Lineage chain length:", lineage["entity_count"])
print("Aggregated metadata:", lineage["metadata"])
Batch Tracking for High-Volume Ingestion
When you need to track source provenance for thousands of entities simultaneously, use track_entities_batch (lines 60‑70). The method automatically partitions the input into 1,000‑item transactions and distinguishes between typed kwargs and free-form metadata for each record.
entities = [
{"id": "E001", "metadata": {"type": "Disease"}},
{"id": "E002", "metadata": {"type": "Drug"}},
# ... thousands more ...
]
count = prov.track_entities_batch(entities, source="github.com/example/repo")
print(f"Tracked {count} entities")
This approach minimizes database overhead while maintaining the same provenance guarantees as individual tracking calls.
Legacy API and Migration Path
The deprecated kg.ProvenanceTracker class in semantica/kg/provenance_tracker.py (lines 17‑33) remains importable for backward compatibility but simply forwards calls to ProvenanceManager. New implementations should import directly from semantica.provenance.manager to avoid deprecation warnings and access the full feature set including cryptographic integrity checks.
# Deprecated approach (avoid)
from semantica.kg.provenance_tracker import ProvenanceTracker
# Recommended approach
from semantica.provenance.manager import ProvenanceManager
Summary
- ProvenanceManager in
semantica/provenance/manager.pyis the unified interface for tracking entity origins with cryptographic integrity. - Use
track_entityfor single entities andtrack_entities_batchfor bulk ingestion (automatically chunked into 1,000‑item transactions). - Query provenance with
get_all_sourcesfor flat lists orget_lineagefor complete ancestry chains with aggregated metadata. - Supply
storage_pathto enable persistent SQLite storage instead of the default in-memory backend. - The legacy
kg.ProvenanceTrackeris deprecated but forwards to the new manager for backward compatibility.
Frequently Asked Questions
What storage backends does ProvenanceManager support?
According to the source code in semantica/provenance/storage.py, the manager supports two storage backends: InMemoryStorage for ephemeral testing and SQLiteStorage for persistent production databases. Pass a filepath to the storage_path parameter to activate SQLite persistence.
How does batch tracking handle transaction failures?
The track_entities_batch method splits large entity lists into 1,000‑item chunks as implemented in semantica/provenance/manager.py lines 60‑70. Each chunk executes as a separate transaction, allowing partial success. If a chunk fails, previous successful commits remain persisted, and you receive the count of successfully tracked entities as the return value.
What is the difference between metadata and typed kwargs in track_entity?
The track_entity method accepts a metadata dictionary for arbitrary JSON‑serializable data stored as a blob, while typed kwargs like confidence, agent_id, and activity_id are extracted and stored in dedicated database columns for indexed querying. This dual approach balances schema flexibility with query performance.
Is the legacy ProvenanceTracker still maintained?
The kg.ProvenanceTracker in semantica/kg/provenance_tracker.py is marked deprecated but remains functional by forwarding all method calls to ProvenanceManager. While maintained for backward compatibility, new code should migrate to the unified manager API to access batch operations, cryptographic checks, and future enhancements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →