# How to Track Source Provenance for Entities in Semantica: A Complete Guide

> Learn how to track source provenance for entities in Semantica. Discover how the ProvenanceManager class ensures cryptographic integrity, confidence scores, and full lineage chains for your knowledge graph data.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-13

---

**Semantica records the origin of every knowledge-graph entity through the `ProvenanceManager` class in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py), which stores source identifiers, confidence scores, and full lineage chains with cryptographic integrity checks.**

The `semantica-agi/semantica` repository provides a robust provenance tracking system that allows you to track source provenance for entities in Semantica with fine-grained metadata and audit capabilities. Whether you are ingesting scientific papers, database records, or curated annotations, the provenance manager creates an immutable trail linking each entity to its original source document.

## Initializing the ProvenanceManager

The `ProvenanceManager` class serves as the primary interface for all provenance operations in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py). By default, it uses an in-memory store suitable for development or testing. For production workloads, instantiate the manager with a `storage_path` to enable persistent SQLite storage.

```python
from semantica.provenance.manager import ProvenanceManager

# In-memory storage (default)

prov = ProvenanceManager()

# Persistent SQLite storage

prov = ProvenanceManager(storage_path="provenance.db")

```

The manager implements the same public API previously provided by the older `kg.ProvenanceTracker`, but adds support for batch operations, richer metadata schemas, and cryptographic integrity verification.

## Tracking Individual Entity Provenance

To track source provenance for a single entity, call the `track_entity` method defined at lines 66‑78 of [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py). This method requires an `entity_id` and `source` parameter, with optional metadata and typed keyword arguments that populate dedicated database columns.

**Key parameters:**
- `entity_id` – The knowledge graph identifier to track
- `source` – A stable source identifier (DOI, URL, or file path)
- `metadata` – Free-form dictionary for arbitrary context
- `confidence`, `agent_id`, `activity_id` – Typed kwargs stored in dedicated columns

```python
from semantica.provenance.manager import ProvenanceManager

prov = ProvenanceManager()
prov.track_entity(
    entity_id="E123",
    source="doi:10.1038/nature12345",
    metadata={"label": "Gene", "confidence": 0.98},
    confidence=0.98,
    agent_id="curation_bot",
    activity_id="entity_ingest"
)

```

The implementation separates structured provenance data from flexible metadata, enabling efficient querying while preserving extensibility.

## Querying Provenance and Lineage

The manager provides two primary methods to retrieve provenance information from [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py).

**`get_all_sources(entity_id)`** (lines 90‑100) returns a flat list of source records containing the document reference, location, timestamp, and confidence score.

**`get_lineage(entity_id)`** (lines 753‑770) constructs a comprehensive lineage dictionary that includes the full ancestry chain, aggregated metadata across all sources, and integrity verification status.

```python

# Retrieve all source records

sources = prov.get_all_sources("E123")
print("Sources:", sources)

# Retrieve rich lineage with ancestry chain

lineage = prov.get_lineage("E123")
print("Lineage chain length:", lineage["entity_count"])
print("Aggregated metadata:", lineage["metadata"])

```

## Batch Tracking for High-Volume Ingestion

When you need to track source provenance for thousands of entities simultaneously, use `track_entities_batch` (lines 60‑70). The method automatically partitions the input into 1,000‑item transactions and distinguishes between typed kwargs and free-form metadata for each record.

```python
entities = [
    {"id": "E001", "metadata": {"type": "Disease"}},
    {"id": "E002", "metadata": {"type": "Drug"}},
    # ... thousands more ...

]

count = prov.track_entities_batch(entities, source="github.com/example/repo")
print(f"Tracked {count} entities")

```

This approach minimizes database overhead while maintaining the same provenance guarantees as individual tracking calls.

## Legacy API and Migration Path

The deprecated `kg.ProvenanceTracker` class in [`semantica/kg/provenance_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/provenance_tracker.py) (lines 17‑33) remains importable for backward compatibility but simply forwards calls to `ProvenanceManager`. New implementations should import directly from `semantica.provenance.manager` to avoid deprecation warnings and access the full feature set including cryptographic integrity checks.

```python

# Deprecated approach (avoid)

from semantica.kg.provenance_tracker import ProvenanceTracker

# Recommended approach

from semantica.provenance.manager import ProvenanceManager

```

## Summary

- **ProvenanceManager** in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) is the unified interface for tracking entity origins with cryptographic integrity.
- Use `track_entity` for single entities and `track_entities_batch` for bulk ingestion (automatically chunked into 1,000‑item transactions).
- Query provenance with `get_all_sources` for flat lists or `get_lineage` for complete ancestry chains with aggregated metadata.
- Supply `storage_path` to enable persistent SQLite storage instead of the default in-memory backend.
- The legacy `kg.ProvenanceTracker` is deprecated but forwards to the new manager for backward compatibility.

## Frequently Asked Questions

### What storage backends does ProvenanceManager support?

According to the source code in [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py), the manager supports two storage backends: `InMemoryStorage` for ephemeral testing and `SQLiteStorage` for persistent production databases. Pass a filepath to the `storage_path` parameter to activate SQLite persistence.

### How does batch tracking handle transaction failures?

The `track_entities_batch` method splits large entity lists into 1,000‑item chunks as implemented in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) lines 60‑70. Each chunk executes as a separate transaction, allowing partial success. If a chunk fails, previous successful commits remain persisted, and you receive the count of successfully tracked entities as the return value.

### What is the difference between metadata and typed kwargs in track_entity?

The `track_entity` method accepts a `metadata` dictionary for arbitrary JSON‑serializable data stored as a blob, while typed kwargs like `confidence`, `agent_id`, and `activity_id` are extracted and stored in dedicated database columns for indexed querying. This dual approach balances schema flexibility with query performance.

### Is the legacy ProvenanceTracker still maintained?

The `kg.ProvenanceTracker` in [`semantica/kg/provenance_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/provenance_tracker.py) is marked deprecated but remains functional by forwarding all method calls to `ProvenanceManager`. While maintained for backward compatibility, new code should migrate to the unified manager API to access batch operations, cryptographic checks, and future enhancements.