How Semantica Handles Provenance Tracking for AI Decisions: A Technical Deep Dive

Semantica implements provenance tracking through a dedicated Provenance Manager that captures the complete lineage of data across ingestion, embedding generation, semantic extraction, and reasoning workflows, with pluggable SQLite or in-memory storage backends.

The semantica-agi/semantica framework provides a comprehensive provenance tracking subsystem designed for complete auditability of AI decisions. This system records the origin and transformation history of every knowledge graph entity and fact, ensuring developers can trace any AI-generated output back to its source data and model artifacts.

Architecture of the Provenance Subsystem

The Provenance Manager

At the core of the implementation is the Provenance Manager, located in semantica/provenance/manager.py. This class supersedes the legacy ProvenanceTracker found in semantica/kg/provenance_tracker.py, offering a consistent API for recording data lineage throughout the AI pipeline.

The manager exposes five primary methods:

  • track_entity(entity_id, source, metadata=None) – Registers that a KG entity originated from a specific source (file, URL, or API call). Optional metadata can include model versions, prompts, or confidence scores.
  • track_fact(fact_id, source, metadata=None) – Records provenance for KG facts (triples) using the same mechanism as entities.
  • get_all_sources(entity_id) – Retrieves the complete list of provenance records attached to a specific entity.
  • query_recorded_between(start, end) – Filters records by timestamp for audit windows and compliance checks.
  • export_audit_log(fact_ids, format="json") – Serializes provenance data to JSON or CSV for external consumption.

Storage Backends

The manager delegates persistence to Provenance Storage implementations defined in semantica/provenance/storage.py. The framework provides two built-in backends:

  • SQLiteStorage – Persistent storage suitable for production workloads, maintaining durability across sessions.
  • InMemoryStorage – High-speed storage designed for unit tests and transient workflows.

Each record is automatically stamped with an aware UTC datetime (datetime.now(timezone.utc)) via the schemas defined in semantica/provenance/schemas.py, ensuring reliable event ordering and version history computation.

The Provenance Lifecycle in AI Workflows

Provenance data flows through the framework in five distinct stages:

  1. Data Ingestion – When the ingest module reads a document, it calls track_entity with the document URI and parser metadata (e.g., {"parser": "pdfminer", "parser_version": "2024.3"}).

  2. Embedding Generation – The embeddings layer attaches model provenance by recording the model name, version, and prompt information used to generate vector representations.

  3. Semantic Extraction – Extractors such as NER and relation extraction components wrap their logic with provenance wrappers (see semantica/semantic_extract/semantic_extract_provenance.py). These automatically propagate original source metadata while appending the extractor name and configuration.

  4. Knowledge Graph Construction – As triples are created, track_fact registers the source of each relationship, whether derived from an arXiv paper, GitHub repository, or external API.

  5. Reasoning and Decision Making – The reasoning engine queries provenance records for any fact used in inference, enabling traceability of AI decisions back to original data sources and model artifacts.

Configuration and Backward Compatibility

The design isolates provenance handling from business logic through a configuration flag. Developers enable tracking by setting provenance=True when initializing components.

When disabled (provenance=False), all provenance calls become no-ops, preserving backward compatibility with existing codebases. This pattern was established in the legacy ProvenanceTracker implementation and maintained in the current manager, ensuring zero overhead when audit trails are not required.

Working with Provenance in Code

The following examples demonstrate typical usage patterns for recording and retrieving lineage information:


# Initialize the manager (defaults to SQLite storage in user config)

from semantica.provenance.manager import ProvenanceManager

prov = ProvenanceManager()

# 1. Ingest a document and record its provenance

doc_id = "doc-123"
prov.track_entity(
    entity_id=doc_id,
    source="s3://my-bucket/articles/article1.pdf",
    metadata={"parser": "pdfminer", "parser_version": "2024.3"},
)

# 2. Generate embeddings and attach model provenance

emb_id = "emb-456"
prov.track_entity(
    entity_id=emb_id,
    source=doc_id,
    metadata={"model": "openai/text-embedding-3-large", "model_version": "v1.2"},
)

# 3. Run a NER extractor that automatically propagates provenance

from semantica.semantic_extract.semantic_extract_provenance import NERExtractorWithProvenance

ner = NERExtractorWithProvenance(provenance=True)
entities = ner.extract(text="Alice went to Paris.", source=doc_id)

# 4. Query provenance for a specific entity

source_records = prov.get_all_sources(entity_id="Alice")
for rec in source_records:
    print(rec)

Summary

  • Semantica implements provenance tracking through a centralized Provenance Manager that records the complete lineage of entities and facts across AI workflows.
  • The system supports both SQLite (production) and in-memory (testing) storage backends via semantica/provenance/storage.py.
  • Five core methods—track_entity, track_fact, get_all_sources, query_recorded_between, and export_audit_log—provide complete audit capabilities.
  • Provenance wraps automatically through the extraction pipeline using classes like NERExtractorWithProvenance, ensuring metadata propagation without manual instrumentation.
  • The subsystem can be disabled entirely via the provenance flag, causing all tracking calls to become no-ops for backward compatibility.

Frequently Asked Questions

What is the difference between ProvenanceManager and ProvenanceTracker?

ProvenanceTracker in semantica/kg/provenance_tracker.py represents the legacy implementation that established the original API patterns. ProvenanceManager in semantica/provenance/manager.py is the current implementation that supersedes it, offering improved storage abstraction and timezone-aware timestamp handling. The legacy tracker is maintained for backward compatibility but new code should use the manager.

Can provenance tracking be disabled for performance?

Yes. Setting provenance=False when initializing components causes all provenance calls to become no-ops, eliminating any runtime overhead. This design ensures that production systems can disable auditing when maximum performance is required without modifying application code or removing provenance calls.

How does Semantica ensure accurate timestamps across time zones?

The framework uses timezone-aware UTC datetimes (datetime.now(timezone.utc)) for all provenance records, as implemented in the storage layer. This ensures reliable event ordering and prevents ambiguity when computing version histories across distributed systems or different geographic regions.

What file formats are supported for exporting audit logs?

The export_audit_log method supports JSON and CSV formats via the format parameter. This allows integration with external compliance tools and data governance platforms that require standardized audit trail exports for regulatory review.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →