# How Semantica Tracks Provenance in the Knowledge Graph

> Semantica tracks knowledge graph provenance with an immutable audit trail recording source, timestamp, and metadata for complete lineage tracing from any graph element.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-10

---

**Semantica tracks provenance in the knowledge graph by maintaining an immutable audit trail that records the source origin, UTC timestamp, and metadata for every entity, enabling complete lineage tracing from any graph element back to its original document or API call.**

The `semantica-agi/semantica` repository implements a robust provenance subsystem that ensures every node, edge, and fact inserted into the knowledge graph carries verifiable origin data. By mapping entity identifiers to structured provenance entries, the system supports compliance requirements, temporal reasoning, and debugging of knowledge extraction pipelines.

## Core Provenance Architecture

Semantica provides two primary interfaces for provenance tracking: the deprecated `ProvenanceTracker` class and the modern `ProvenanceManager`. Both components rely on the [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py) backends to persist records, but the manager offers a higher-level API that integrates directly with the knowledge graph construction pipeline.

### The Legacy ProvenanceTracker Class

Located in [`semantica/kg/provenance_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/provenance_tracker.py), the `ProvenanceTracker` class maintains an in-memory dictionary (`self._records`) that maps entity IDs to lists of provenance entries. Each entry stores the **source** (file path, URL, or upstream identifier), a **recorded_at** timestamp generated via `datetime.now(timezone.utc)`, and optional caller-supplied **metadata** such as extractor type or author identity.

```python
from semantica.kg.provenance_tracker import ProvenanceTracker

tracker = ProvenanceTracker()
tracker.track_entity(
    entity_id="E1",
    source="https://github.com/example/repo",
    metadata={"type": "github_repo", "author": "alice"},
)
sources = tracker.get_all_sources("E1")
print(sources)

# [{'source': 'https://github.com/example/repo',

#   'recorded_at': '2024-03-15T12:34:56.789012+00:00',

#   'type': 'github_repo', 'author': 'alice'}]

```

### The Modern ProvenanceManager

The preferred interface resides in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py). The `ProvenanceManager` wraps storage backends (defined in [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py)) such as `InMemoryStorage` or `SQLiteStorage`, providing durability guarantees while exposing methods like `revision_history()` and `export_audit_log()`.

```python
from semantica.provenance.manager import ProvenanceManager

prov = ProvenanceManager()
prov.track_entity("E42", "https://arxiv.org/abs/2101.00001")
history = prov.revision_history("E42")
print(history)

```

## Provenance Data Structure

Every provenance entry follows a strict schema defined in [`semantica/provenance/schemas.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/schemas.py). The system captures three essential fields to ensure comprehensive **auditability**:

- **source**: The authoritative origin of the entity (document filename, URL, or database ID)
- **recorded_at**: An ISO 8601 UTC timestamp generated using `datetime.now(timezone.utc)`
- **metadata**: An optional dictionary for domain-specific context (extractor version, confidence scores, author)

This structure enables **temporal reasoning** queries that determine what was true at specific points in time, as well as **traceability** back to the exact document or API call that generated a fact.

## Automatic Provenance Injection During Graph Construction

When building the knowledge graph, the `GraphBuilderWithProvenance` (implemented in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py)) automatically invokes the provenance tracker upon entity creation. This ensures no node or edge enters the graph without associated lineage metadata.

```python

# Simplified excerpt from semantica/kg/graph_builder.py

if self.provenance:
    self.provenance_tracker.track_entity(
        entity_id=node.id,
        source=source_id,
        metadata={"extractor": extractor_name},
    )

```

This automated injection guarantees that downstream consumers can trace any fact to its originating source document, supporting use cases such as regulatory compliance verification and source conflict resolution.

## Querying and Auditing Provenance Records

The provenance subsystem exposes several query interfaces for retrieving lineage information and generating compliance reports.

### Retrieving Source Lineage

Call `get_all_sources(entity_id)` to fetch the complete list of provenance entries for a specific entity. This returns chronologically ordered records showing every source that contributed to the entity's current state.

```python
sources = prov.get_all_sources("E42")
for record in sources:
    print(f"Source: {record['source']}, Recorded: {record['recorded_at']}")

```

### Temporal Queries and Audit Exports

The `query_recorded_between(start, end)` method returns all records whose timestamps fall within a specified UTC interval, supporting "what-was-true-when" analyses. For compliance workflows, `export_audit_log()` serializes provenance data for a set of entity IDs as JSON or CSV.

```python
from datetime import datetime, timezone

start = datetime(2024, 1, 1, tzinfo=timezone.utc)
end = datetime(2024, 12, 31, tzinfo=timezone.utc)
records = prov.query_recorded_between(start, end)

# Export for external auditing

json_log = prov.export_audit_log(["E42", "E99"], format="json")

```

## Storage Backends and Durability

The [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py) module defines pluggable storage backends that guarantee durability while optimizing for lookup performance. Available implementations include:

- **InMemoryStorage**: High-speed volatile storage suitable for testing and short-lived pipelines
- **SQLiteStorage**: Persistent disk-based storage for production knowledge graphs requiring durability across restarts

These backends power the `ProvenanceManager`, ensuring that provenance records survive system restarts while maintaining fast retrieval times for graph visualization tools and CLI queries.

## Summary

- **Semantica tracks provenance in the knowledge graph** using a dedicated subsystem that records source, timestamp, and metadata for every entity.
- The **ProvenanceManager** (and legacy **ProvenanceTracker**) in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) and [`semantica/kg/provenance_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/provenance_tracker.py) provide the core API for recording lineage.
- **Automatic injection** during graph construction ensures every node and edge carries origin metadata without manual intervention.
- **Temporal queries** via `query_recorded_between()` and **audit exports** via `export_audit_log()` support compliance and debugging workflows.
- **Pluggable storage backends** in [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py) offer both in-memory speed and persistent durability.

## Frequently Asked Questions

### What is the difference between ProvenanceTracker and ProvenanceManager?

The `ProvenanceTracker` class in [`semantica/kg/provenance_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/provenance_tracker.py) is the legacy implementation that maintains records in a simple in-memory dictionary. The `ProvenanceManager` in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) is the modern, preferred interface that abstracts storage backends (SQLite, in-memory) and provides additional methods like `revision_history()` and `export_audit_log()`. New projects should use `ProvenanceManager` for better durability and API stability.

### How do I export a provenance audit log in Semantica?

Use the `export_audit_log()` method on a `ProvenanceManager` instance. Pass a list of entity IDs and specify the format ("json" or "csv"). This method serializes the complete provenance records—including source, timestamps, and metadata—into a portable format suitable for external compliance audits or regulatory reporting.

### Can I query provenance records by time range?

Yes. The `ProvenanceManager.query_recorded_between(start_datetime, end_datetime)` method accepts timezone-aware datetime objects (recommended: `datetime.now(timezone.utc)`) and returns all provenance entries recorded within that interval. This enables temporal reasoning queries to determine the state of knowledge at specific historical moments.

### How does Semantica automatically track provenance when building a knowledge graph?

During graph construction, the `GraphBuilderWithProvenance` (in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py)) automatically calls `track_entity()` for every node and edge it creates, passing the source document ID and extractor metadata. This ensures the knowledge graph builder populates provenance data transparently without requiring manual instrumentation from developers.