Triplet Extraction Handling Temporal and Provenance Metadata in Semantica AGI
The Semantica AGI framework implements a modular four-stage pipeline that extracts subject-predicate-object triplets from natural language and enriches them with temporal bounds and provenance metadata through immutable versioned snapshots.
The semantica-agi/semantica repository provides a production-ready architecture for triplet extraction handling temporal and provenance metadata. The system separates concerns across extraction, temporal binding, version management, and storage layers, enabling robust knowledge graph construction with full historical traceability.
Stage 1: Core Triplet Extraction
The pipeline begins in semantica/semantic_extract/triplet_extractor.py, where the TripletExtractor class parses raw text into structured Triplet objects. This module uses spaCy or configurable LLM backends to identify entities and relationships, emitting dataclass instances containing subject, predicate, and object fields.
The extractor operates agnostically regarding temporal data, producing clean triplets that downstream components enrich. This separation ensures that the core parsing logic remains lightweight while allowing optional metadata attachment without breaking the extraction contract.
Stage 2: Temporal Enrichment
Once raw triplets exist, the system associates them with temporal information through semantica/kg/temporal_model.py. The TemporalBound class wraps timestamps and temporal intervals, supported by helper functions parse_temporal_value and serialize_temporal_bound.
These utilities normalize diverse input formats—including ISO 8601 strings, Python datetime objects, and relative temporal descriptors—into a consistent internal representation. When the extractor detects temporal cues in source text or when users supply explicit timestamps, the framework attaches a TemporalBound instance to the triplet's optional temporal attribute without altering the core subject-predicate-object structure.
Stage 3: Provenance Tracking and Versioning
The TemporalVersionManager class in semantica/kg/temporal_query.py implements provenance tracking by creating immutable snapshots of the knowledge graph. Each snapshot captures both transaction time (when the data was ingested) and valid time (when the fact is asserted to hold true).
This distinction enables temporal database semantics, allowing the system to answer questions like "What did we know on Monday?" versus "What was true on Monday?" The manager serializes these snapshots using OWL-Time-compatible RDF, ensuring that temporal instants and intervals conform to W3C standards for downstream interoperability.
Stage 4: Persistent Storage
The enriched triplets persist through semantica/triplet_store/triplet_store.py, which abstracts multiple graph backends including Blazegraph, RDF4J, and Apache Jena. Concrete implementations such as BlazegraphStore (defined in semantica/triplet_store/blazegraph_store.py) perform SPARQL sanitization to prevent injection attacks when inserting subjects or predicates containing IRIs.
The TripletStore class handles addition, deletion, and querying operations across backend variants, ensuring that temporal and provenance metadata remains intact throughout the persistence lifecycle.
End-to-End Implementation Example
The following workflow demonstrates the complete pipeline from extraction to visualization:
# 1. Extract raw triplets from unstructured text
from semantica.semantic_extract.triplet_extractor import TripletExtractor
text = """
Alice met Bob on 2022-05-01. Later that day she emailed him about the project.
"""
extractor = TripletExtractor()
raw_triplets = extractor.extract(text) # Returns List[Triplet]
# 2. Enrich with temporal metadata
from semantica.kg.temporal_model import TemporalBound, parse_temporal_value
timestamp = parse_temporal_value("2022-05-01")
for triplet in raw_triplets:
triplet.temporal = TemporalBound(value=timestamp)
# 3. Create versioned snapshot with provenance
from semantica.kg.temporal_query import TemporalVersionManager
tv_manager = TemporalVersionManager()
snapshot = tv_manager.create_version(
graph={"triplets": raw_triplets},
timestamp="2022-05-01T12:00:00Z",
version_label="v1.0"
)
# snapshot now contains transaction_time and valid_time metadata
# 4. Persist to graph store
from semantica.triplet_store.triplet_store import TripletStore
store = TripletStore(backend="blazegraph")
store.add_triplets(snapshot["triplets"])
# 5. Visualize temporal evolution (optional)
from semantica.visualization.temporal_visualizer import TemporalVisualizer
vis = TemporalVisualizer()
vis.visualize_timeline(snapshot) # Generates interactive Plotly chart
Orchestration via CLI
The framework wires all components together through the semantica temporal CLI sub-command. This interface orchestrates extraction, temporal enrichment, version creation, and optional visualization—including timeline charts, metric evolution graphs, and snapshot comparisons—without requiring manual pipeline construction.
Summary
- Modular extraction in
triplet_extractor.pygenerates cleanTripletobjects agnostic of metadata. - Temporal binding via
TemporalBoundandparse_temporal_valueintemporal_model.pynormalizes diverse time formats. - Provenance versioning through
TemporalVersionManagermaintains immutable snapshots with transaction and valid time semantics. - Secure persistence across multiple graph backends includes SPARQL injection prevention in backend-specific store implementations.
- OWL-Time compliance ensures exported temporal metadata integrates with standard semantic web toolchains.
Frequently Asked Questions
How does the framework handle different temporal input formats?
The parse_temporal_value utility in semantica/kg/temporal_model.py accepts ISO 8601 strings, Python datetime objects, and relative descriptors, normalizing them into consistent TemporalBound instances. This abstraction allows extractors to capture temporal information from heterogeneous sources without preprocessing.
What distinguishes transaction time from valid time in Semantica AGI?
Transaction time records when the system ingested a fact into the knowledge graph, while valid time indicates when the fact actually holds true in the real world. The TemporalVersionManager tracks both dimensions independently, enabling queries that reconstruct historical knowledge states or verify temporal accuracy of assertions.
Which graph databases support the enriched triplet storage?
The TripletStore abstraction supports Blazegraph, RDF4J, Apache Jena, and other SPARQL-compliant backends through configurable driver classes. Each backend implementation in semantica/triplet_store/ handles connection management and query sanitization specific to that database's requirements.
How does the system prevent SPARQL injection when persisting triplets?
Backend implementations such as BlazegraphStore sanitize all IRI and literal inputs before constructing SPARQL INSERT statements. The sanitization routines escape special characters and validate URI schemes, ensuring that malicious payloads in subject or predicate fields cannot alter query semantics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →