# Triplet Extraction Handling Temporal and Provenance Metadata in Semantica AGI

> Discover how Semantica AGI's modular pipeline extracts triplets from text, enriching them with temporal and provenance metadata for robust knowledge representation.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-11

---

**The Semantica AGI framework implements a modular four-stage pipeline that extracts subject-predicate-object triplets from natural language and enriches them with temporal bounds and provenance metadata through immutable versioned snapshots.**

The `semantica-agi/semantica` repository provides a production-ready architecture for triplet extraction handling temporal and provenance metadata. The system separates concerns across extraction, temporal binding, version management, and storage layers, enabling robust knowledge graph construction with full historical traceability.

## Stage 1: Core Triplet Extraction

The pipeline begins in [`semantica/semantic_extract/triplet_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/triplet_extractor.py), where the **`TripletExtractor`** class parses raw text into structured **`Triplet`** objects. This module uses spaCy or configurable LLM backends to identify entities and relationships, emitting dataclass instances containing `subject`, `predicate`, and `object` fields.

The extractor operates agnostically regarding temporal data, producing clean triplets that downstream components enrich. This separation ensures that the core parsing logic remains lightweight while allowing optional metadata attachment without breaking the extraction contract.

## Stage 2: Temporal Enrichment

Once raw triplets exist, the system associates them with temporal information through [`semantica/kg/temporal_model.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/temporal_model.py). The **`TemporalBound`** class wraps timestamps and temporal intervals, supported by helper functions **`parse_temporal_value`** and **`serialize_temporal_bound`**.

These utilities normalize diverse input formats—including ISO 8601 strings, Python `datetime` objects, and relative temporal descriptors—into a consistent internal representation. When the extractor detects temporal cues in source text or when users supply explicit timestamps, the framework attaches a `TemporalBound` instance to the triplet's optional `temporal` attribute without altering the core subject-predicate-object structure.

## Stage 3: Provenance Tracking and Versioning

The **`TemporalVersionManager`** class in [`semantica/kg/temporal_query.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/temporal_query.py) implements provenance tracking by creating immutable snapshots of the knowledge graph. Each snapshot captures both **transaction time** (when the data was ingested) and **valid time** (when the fact is asserted to hold true).

This distinction enables temporal database semantics, allowing the system to answer questions like "What did we know on Monday?" versus "What was true on Monday?" The manager serializes these snapshots using **OWL-Time-compatible RDF**, ensuring that temporal instants and intervals conform to W3C standards for downstream interoperability.

## Stage 4: Persistent Storage

The enriched triplets persist through [`semantica/triplet_store/triplet_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/triplet_store/triplet_store.py), which abstracts multiple graph backends including Blazegraph, RDF4J, and Apache Jena. Concrete implementations such as **`BlazegraphStore`** (defined in [`semantica/triplet_store/blazegraph_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/triplet_store/blazegraph_store.py)) perform SPARQL sanitization to prevent injection attacks when inserting subjects or predicates containing IRIs.

The `TripletStore` class handles addition, deletion, and querying operations across backend variants, ensuring that temporal and provenance metadata remains intact throughout the persistence lifecycle.

## End-to-End Implementation Example

The following workflow demonstrates the complete pipeline from extraction to visualization:

```python

# 1. Extract raw triplets from unstructured text

from semantica.semantic_extract.triplet_extractor import TripletExtractor

text = """
Alice met Bob on 2022-05-01. Later that day she emailed him about the project.
"""
extractor = TripletExtractor()
raw_triplets = extractor.extract(text)  # Returns List[Triplet]

# 2. Enrich with temporal metadata

from semantica.kg.temporal_model import TemporalBound, parse_temporal_value

timestamp = parse_temporal_value("2022-05-01")
for triplet in raw_triplets:
    triplet.temporal = TemporalBound(value=timestamp)

# 3. Create versioned snapshot with provenance

from semantica.kg.temporal_query import TemporalVersionManager

tv_manager = TemporalVersionManager()
snapshot = tv_manager.create_version(
    graph={"triplets": raw_triplets},
    timestamp="2022-05-01T12:00:00Z",
    version_label="v1.0"
)

# snapshot now contains transaction_time and valid_time metadata

# 4. Persist to graph store

from semantica.triplet_store.triplet_store import TripletStore

store = TripletStore(backend="blazegraph")
store.add_triplets(snapshot["triplets"])

# 5. Visualize temporal evolution (optional)

from semantica.visualization.temporal_visualizer import TemporalVisualizer

vis = TemporalVisualizer()
vis.visualize_timeline(snapshot)  # Generates interactive Plotly chart

```

## Orchestration via CLI

The framework wires all components together through the **`semantica temporal`** CLI sub-command. This interface orchestrates extraction, temporal enrichment, version creation, and optional visualization—including timeline charts, metric evolution graphs, and snapshot comparisons—without requiring manual pipeline construction.

## Summary

- **Modular extraction** in [`triplet_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/triplet_extractor.py) generates clean `Triplet` objects agnostic of metadata.
- **Temporal binding** via `TemporalBound` and `parse_temporal_value` in [`temporal_model.py`](https://github.com/semantica-agi/semantica/blob/main/temporal_model.py) normalizes diverse time formats.
- **Provenance versioning** through `TemporalVersionManager` maintains immutable snapshots with transaction and valid time semantics.
- **Secure persistence** across multiple graph backends includes SPARQL injection prevention in backend-specific store implementations.
- **OWL-Time compliance** ensures exported temporal metadata integrates with standard semantic web toolchains.

## Frequently Asked Questions

### How does the framework handle different temporal input formats?

The **`parse_temporal_value`** utility in [`semantica/kg/temporal_model.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/temporal_model.py) accepts ISO 8601 strings, Python `datetime` objects, and relative descriptors, normalizing them into consistent `TemporalBound` instances. This abstraction allows extractors to capture temporal information from heterogeneous sources without preprocessing.

### What distinguishes transaction time from valid time in Semantica AGI?

**Transaction time** records when the system ingested a fact into the knowledge graph, while **valid time** indicates when the fact actually holds true in the real world. The `TemporalVersionManager` tracks both dimensions independently, enabling queries that reconstruct historical knowledge states or verify temporal accuracy of assertions.

### Which graph databases support the enriched triplet storage?

The `TripletStore` abstraction supports **Blazegraph**, **RDF4J**, **Apache Jena**, and other SPARQL-compliant backends through configurable driver classes. Each backend implementation in `semantica/triplet_store/` handles connection management and query sanitization specific to that database's requirements.

### How does the system prevent SPARQL injection when persisting triplets?

Backend implementations such as **`BlazegraphStore`** sanitize all IRI and literal inputs before constructing SPARQL INSERT statements. The sanitization routines escape special characters and validate URI schemes, ensuring that malicious payloads in subject or predicate fields cannot alter query semantics.