# How Semantica Handles Provenance Tracking for AI Decisions: A Technical Deep Dive

> Discover how Semantica handles provenance tracking for AI decisions, capturing complete data lineage from ingestion to reasoning with pluggable storage. Explore the technical deep dive.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-13

---

**Semantica implements provenance tracking through a dedicated Provenance Manager that captures the complete lineage of data across ingestion, embedding generation, semantic extraction, and reasoning workflows, with pluggable SQLite or in-memory storage backends.**

The `semantica-agi/semantica` framework provides a comprehensive provenance tracking subsystem designed for complete auditability of AI decisions. This system records the origin and transformation history of every knowledge graph entity and fact, ensuring developers can trace any AI-generated output back to its source data and model artifacts.

## Architecture of the Provenance Subsystem

### The Provenance Manager

At the core of the implementation is the **Provenance Manager**, located in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py). This class supersedes the legacy `ProvenanceTracker` found in [`semantica/kg/provenance_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/provenance_tracker.py), offering a consistent API for recording data lineage throughout the AI pipeline.

The manager exposes five primary methods:

- `track_entity(entity_id, source, metadata=None)` – Registers that a KG entity originated from a specific source (file, URL, or API call). Optional metadata can include model versions, prompts, or confidence scores.
- `track_fact(fact_id, source, metadata=None)` – Records provenance for KG facts (triples) using the same mechanism as entities.
- `get_all_sources(entity_id)` – Retrieves the complete list of provenance records attached to a specific entity.
- `query_recorded_between(start, end)` – Filters records by timestamp for audit windows and compliance checks.
- `export_audit_log(fact_ids, format="json")` – Serializes provenance data to JSON or CSV for external consumption.

### Storage Backends

The manager delegates persistence to **Provenance Storage** implementations defined in [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py). The framework provides two built-in backends:

- **SQLiteStorage** – Persistent storage suitable for production workloads, maintaining durability across sessions.
- **InMemoryStorage** – High-speed storage designed for unit tests and transient workflows.

Each record is automatically stamped with an aware UTC datetime (`datetime.now(timezone.utc)`) via the schemas defined in [`semantica/provenance/schemas.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/schemas.py), ensuring reliable event ordering and version history computation.

## The Provenance Lifecycle in AI Workflows

Provenance data flows through the framework in five distinct stages:

1. **Data Ingestion** – When the ingest module reads a document, it calls `track_entity` with the document URI and parser metadata (e.g., `{"parser": "pdfminer", "parser_version": "2024.3"}`).

2. **Embedding Generation** – The embeddings layer attaches model provenance by recording the model name, version, and prompt information used to generate vector representations.

3. **Semantic Extraction** – Extractors such as NER and relation extraction components wrap their logic with provenance wrappers (see [`semantica/semantic_extract/semantic_extract_provenance.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/semantic_extract_provenance.py)). These automatically propagate original source metadata while appending the extractor name and configuration.

4. **Knowledge Graph Construction** – As triples are created, `track_fact` registers the source of each relationship, whether derived from an arXiv paper, GitHub repository, or external API.

5. **Reasoning and Decision Making** – The reasoning engine queries provenance records for any fact used in inference, enabling traceability of AI decisions back to original data sources and model artifacts.

## Configuration and Backward Compatibility

The design isolates provenance handling from business logic through a configuration flag. Developers enable tracking by setting `provenance=True` when initializing components.

When disabled (`provenance=False`), all provenance calls become no-ops, preserving backward compatibility with existing codebases. This pattern was established in the legacy `ProvenanceTracker` implementation and maintained in the current manager, ensuring zero overhead when audit trails are not required.

## Working with Provenance in Code

The following examples demonstrate typical usage patterns for recording and retrieving lineage information:

```python

# Initialize the manager (defaults to SQLite storage in user config)

from semantica.provenance.manager import ProvenanceManager

prov = ProvenanceManager()

# 1. Ingest a document and record its provenance

doc_id = "doc-123"
prov.track_entity(
    entity_id=doc_id,
    source="s3://my-bucket/articles/article1.pdf",
    metadata={"parser": "pdfminer", "parser_version": "2024.3"},
)

# 2. Generate embeddings and attach model provenance

emb_id = "emb-456"
prov.track_entity(
    entity_id=emb_id,
    source=doc_id,
    metadata={"model": "openai/text-embedding-3-large", "model_version": "v1.2"},
)

# 3. Run a NER extractor that automatically propagates provenance

from semantica.semantic_extract.semantic_extract_provenance import NERExtractorWithProvenance

ner = NERExtractorWithProvenance(provenance=True)
entities = ner.extract(text="Alice went to Paris.", source=doc_id)

# 4. Query provenance for a specific entity

source_records = prov.get_all_sources(entity_id="Alice")
for rec in source_records:
    print(rec)

```

## Summary

- **Semantica** implements provenance tracking through a centralized **Provenance Manager** that records the complete lineage of entities and facts across AI workflows.
- The system supports both **SQLite** (production) and **in-memory** (testing) storage backends via [`semantica/provenance/storage.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/storage.py).
- Five core methods—`track_entity`, `track_fact`, `get_all_sources`, `query_recorded_between`, and `export_audit_log`—provide complete audit capabilities.
- Provenance wraps automatically through the extraction pipeline using classes like `NERExtractorWithProvenance`, ensuring metadata propagation without manual instrumentation.
- The subsystem can be disabled entirely via the `provenance` flag, causing all tracking calls to become no-ops for backward compatibility.

## Frequently Asked Questions

### What is the difference between ProvenanceManager and ProvenanceTracker?

`ProvenanceTracker` in [`semantica/kg/provenance_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/provenance_tracker.py) represents the legacy implementation that established the original API patterns. `ProvenanceManager` in [`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py) is the current implementation that supersedes it, offering improved storage abstraction and timezone-aware timestamp handling. The legacy tracker is maintained for backward compatibility but new code should use the manager.

### Can provenance tracking be disabled for performance?

Yes. Setting `provenance=False` when initializing components causes all provenance calls to become no-ops, eliminating any runtime overhead. This design ensures that production systems can disable auditing when maximum performance is required without modifying application code or removing provenance calls.

### How does Semantica ensure accurate timestamps across time zones?

The framework uses timezone-aware UTC datetimes (`datetime.now(timezone.utc)`) for all provenance records, as implemented in the storage layer. This ensures reliable event ordering and prevents ambiguity when computing version histories across distributed systems or different geographic regions.

### What file formats are supported for exporting audit logs?

The `export_audit_log` method supports **JSON** and **CSV** formats via the `format` parameter. This allows integration with external compliance tools and data governance platforms that require standardized audit trail exports for regulatory review.