Triplet Extraction Handling Temporal and Provenance Metadata in Semantica AGI

The Semantica AGI framework implements a modular four-stage pipeline that extracts subject-predicate-object triplets from natural language and enriches them with temporal bounds and provenance metadata through immutable versioned snapshots.

The semantica-agi/semantica repository provides a production-ready architecture for triplet extraction handling temporal and provenance metadata. The system separates concerns across extraction, temporal binding, version management, and storage layers, enabling robust knowledge graph construction with full historical traceability.

Stage 1: Core Triplet Extraction

The pipeline begins in semantica/semantic_extract/triplet_extractor.py, where the TripletExtractor class parses raw text into structured Triplet objects. This module uses spaCy or configurable LLM backends to identify entities and relationships, emitting dataclass instances containing subject, predicate, and object fields.

The extractor operates agnostically regarding temporal data, producing clean triplets that downstream components enrich. This separation ensures that the core parsing logic remains lightweight while allowing optional metadata attachment without breaking the extraction contract.

Stage 2: Temporal Enrichment

Once raw triplets exist, the system associates them with temporal information through semantica/kg/temporal_model.py. The TemporalBound class wraps timestamps and temporal intervals, supported by helper functions parse_temporal_value and serialize_temporal_bound.

These utilities normalize diverse input formats—including ISO 8601 strings, Python datetime objects, and relative temporal descriptors—into a consistent internal representation. When the extractor detects temporal cues in source text or when users supply explicit timestamps, the framework attaches a TemporalBound instance to the triplet's optional temporal attribute without altering the core subject-predicate-object structure.

Stage 3: Provenance Tracking and Versioning

The TemporalVersionManager class in semantica/kg/temporal_query.py implements provenance tracking by creating immutable snapshots of the knowledge graph. Each snapshot captures both transaction time (when the data was ingested) and valid time (when the fact is asserted to hold true).

This distinction enables temporal database semantics, allowing the system to answer questions like "What did we know on Monday?" versus "What was true on Monday?" The manager serializes these snapshots using OWL-Time-compatible RDF, ensuring that temporal instants and intervals conform to W3C standards for downstream interoperability.

Stage 4: Persistent Storage

The enriched triplets persist through semantica/triplet_store/triplet_store.py, which abstracts multiple graph backends including Blazegraph, RDF4J, and Apache Jena. Concrete implementations such as BlazegraphStore (defined in semantica/triplet_store/blazegraph_store.py) perform SPARQL sanitization to prevent injection attacks when inserting subjects or predicates containing IRIs.

The TripletStore class handles addition, deletion, and querying operations across backend variants, ensuring that temporal and provenance metadata remains intact throughout the persistence lifecycle.

End-to-End Implementation Example

The following workflow demonstrates the complete pipeline from extraction to visualization:


# 1. Extract raw triplets from unstructured text

from semantica.semantic_extract.triplet_extractor import TripletExtractor

text = """
Alice met Bob on 2022-05-01. Later that day she emailed him about the project.
"""
extractor = TripletExtractor()
raw_triplets = extractor.extract(text)  # Returns List[Triplet]

# 2. Enrich with temporal metadata

from semantica.kg.temporal_model import TemporalBound, parse_temporal_value

timestamp = parse_temporal_value("2022-05-01")
for triplet in raw_triplets:
    triplet.temporal = TemporalBound(value=timestamp)

# 3. Create versioned snapshot with provenance

from semantica.kg.temporal_query import TemporalVersionManager

tv_manager = TemporalVersionManager()
snapshot = tv_manager.create_version(
    graph={"triplets": raw_triplets},
    timestamp="2022-05-01T12:00:00Z",
    version_label="v1.0"
)

# snapshot now contains transaction_time and valid_time metadata

# 4. Persist to graph store

from semantica.triplet_store.triplet_store import TripletStore

store = TripletStore(backend="blazegraph")
store.add_triplets(snapshot["triplets"])

# 5. Visualize temporal evolution (optional)

from semantica.visualization.temporal_visualizer import TemporalVisualizer

vis = TemporalVisualizer()
vis.visualize_timeline(snapshot)  # Generates interactive Plotly chart

Orchestration via CLI

The framework wires all components together through the semantica temporal CLI sub-command. This interface orchestrates extraction, temporal enrichment, version creation, and optional visualization—including timeline charts, metric evolution graphs, and snapshot comparisons—without requiring manual pipeline construction.

Summary

  • Modular extraction in triplet_extractor.py generates clean Triplet objects agnostic of metadata.
  • Temporal binding via TemporalBound and parse_temporal_value in temporal_model.py normalizes diverse time formats.
  • Provenance versioning through TemporalVersionManager maintains immutable snapshots with transaction and valid time semantics.
  • Secure persistence across multiple graph backends includes SPARQL injection prevention in backend-specific store implementations.
  • OWL-Time compliance ensures exported temporal metadata integrates with standard semantic web toolchains.

Frequently Asked Questions

How does the framework handle different temporal input formats?

The parse_temporal_value utility in semantica/kg/temporal_model.py accepts ISO 8601 strings, Python datetime objects, and relative descriptors, normalizing them into consistent TemporalBound instances. This abstraction allows extractors to capture temporal information from heterogeneous sources without preprocessing.

What distinguishes transaction time from valid time in Semantica AGI?

Transaction time records when the system ingested a fact into the knowledge graph, while valid time indicates when the fact actually holds true in the real world. The TemporalVersionManager tracks both dimensions independently, enabling queries that reconstruct historical knowledge states or verify temporal accuracy of assertions.

Which graph databases support the enriched triplet storage?

The TripletStore abstraction supports Blazegraph, RDF4J, Apache Jena, and other SPARQL-compliant backends through configurable driver classes. Each backend implementation in semantica/triplet_store/ handles connection management and query sanitization specific to that database's requirements.

How does the system prevent SPARQL injection when persisting triplets?

Backend implementations such as BlazegraphStore sanitize all IRI and literal inputs before constructing SPARQL INSERT statements. The sanitization routines escape special characters and validate URI schemes, ensuring that malicious payloads in subject or predicate fields cannot alter query semantics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →