Semantica Semantic Extraction Capabilities: A Complete Guide to Knowledge Graph Pipeline

Semantica offers nine core semantic extraction capabilities ranging from named-entity recognition and relation extraction to event detection and provenance tracking, all orchestrated through a modular Python pipeline that converts raw text into queryable RDF-style knowledge graphs.

The open-source Semantica library (semantica-agi/semantica) provides a full-stack semantic extraction system designed for production knowledge graph construction. Its architecture separates each extraction concern into dedicated modules under semantica.semantic_extract, allowing developers to run complete pipelines or invoke individual extractors as needed.

Core Semantic Extraction Capabilities

Semantica's extraction layer addresses every stage of the knowledge graph construction pipeline. Each capability is implemented as a focused Python class with swappable backends.

Named-Entity Recognition (NER)

The NERExtractor class in semantica/semantic_extract/ner_extractor.py detects and classifies entities including people, organizations, locations, and dates. It supports three provider modes:

  • Pattern-based: Rule-driven regex matching
  • spaCy: Dependency parsing with linguistic features
  • LLM providers: OpenAI, Ollama, or custom implementations
from semantica.semantic_extract import NERExtractor

ner = NERExtractor(method="pattern")  # or "ollama", "openai"

entities = ner.extract_entities("Apple announced iPhone 15 in California.")

# → [Entity(text='Apple'), Entity(text='iPhone 15'), Entity(text='California')]

Relation Extraction

The RelationExtractor in semantica/semantic_extract/relation_extractor.py identifies semantic relationships between entities. It implements:

  • Pattern-based extraction: Predefined linguistic templates
  • Co-occurrence heuristics: Proximity-based relationship scoring
  • Dependency-tree parsing: Grammatical structure analysis

The module includes graceful degradation when spaCy dependencies are unavailable.

Triplet Extraction

TripletExtractor (semantica/semantic_extract/triplet_extractor.py) converts entities and relations into RDF-style (subject, predicate, object) triples. Key features include:

  • Rule-based generation as fallback when LLM methods are disabled
  • Validation and normalization of extracted triples
  • Direct compatibility with graph database import formats

Event Detection

The EventDetector class in semantica/semantic_extract/event_detector.py extracts narrative events with temporal semantics. It identifies:

  • Actions and their agents
  • Timestamps and temporal expressions
  • Event participants and their roles

This enables time-aware graph queries and temporal reasoning over document collections.

Coreference Resolution

CoreferenceResolver (semantica/semantic_extract/coreference_resolver.py) links pronouns and aliases to canonical entity mentions. This ensures:

  • Consistent node IDs across the knowledge graph
  • Reduced entity fragmentation
  • Accurate relationship mapping for referred entities

Semantic Role Analysis

The SemanticAnalyzer in semantica/semantic_extract/semantic_analyzer.py performs semantic role labeling, assigning functional roles such as:

  • Agent: The doer of an action
  • Patient: The entity affected by an action
  • Instrument: The means by which an action is performed

Semantic Network Construction

SemanticNetworkExtractor (semantica/semantic_extract/semantic_network_extractor.py) serves as the high-level orchestrator. It assembles:

  • Extracted nodes (entities, events)
  • Edges (relations, roles)
  • Provenance metadata

The output is a coherent knowledge graph ready for persistence in Neo4j, Amazon Neptune, or other RDF stores.

Provenance Tracking

Every extractor integrates auditability features capturing:

  • Source document references
  • Extraction timestamps
  • Method metadata and provider configurations

This supports reproducibility and compliance requirements. The progress_tracker pattern used throughout the test suite demonstrates consistent provenance implementation.

Pluggable Provider Architecture

The provider abstraction layer in semantica/semantic_extract/providers.py enables backend swapping without pipeline changes. Supported providers include:

Provider Use Case
PatternProvider Fast, deterministic extraction without external dependencies
SpacyProvider Linguistically-informed analysis
OllamaProvider Local LLM inference
OpenAIProvider Cloud-based language model access
Custom implementations Domain-specific extraction logic

Running a Complete Extraction Pipeline

The following example demonstrates end-to-end semantic extraction using Semantica's orchestrated pipeline:

from semantica.semantic_extract import (
    NERExtractor,
    RelationExtractor,
    TripletExtractor,
    EventDetector,
    SemanticNetworkExtractor,
)

document = """Apple announced the new iPhone 15 in California. 
              Tim Cook said it will launch in September."""

# Individual extractor approach

ner = NERExtractor(method="pattern")
entities = ner.extract_entities(document)

rel_extractor = RelationExtractor(method="dependency")
relations = rel_extractor.extract_relations(document, entities)

triplets = TripletExtractor().extract_triplets(document, entities, relations)

events = EventDetector().extract_events(document, entities)

# Unified pipeline approach

network_extractor = SemanticNetworkExtractor()
semantic_graph = network_extractor.build(document)

# Contains nodes, edges, provenance metadata ready for export

Switching to LLM-backed extraction requires only parameter changes:

ner = NERExtractor(method="ollama")      # Local LLM via Ollama

rel_extractor = RelationExtractor(method="openai")  # OpenAI API

Module Reference and Source Files

File Path Purpose
semantica/semantic_extract/ner_extractor.py Entity detection with provider dispatch
semantica/semantic_extract/relation_extractor.py Relationship identification strategies
semantica/semantic_extract/triplet_extractor.py RDF triple generation and validation
semantica/semantic_extract/event_detector.py Temporal event extraction
semantica/semantic_extract/coreference_resolver.py Pronoun and alias resolution
semantica/semantic_extract/semantic_analyzer.py Semantic role assignment
semantica/semantic_extract/semantic_network_extractor.py Pipeline orchestration and graph assembly
semantica/semantic_extract/providers.py Backend abstraction layer
semantica/semantic_extract/methods.py Core algorithm implementations

Summary

  • Semantica provides nine specialized semantic extraction capabilities through modular Python classes in semantica.semantic_extract
  • Named-entity recognition, relation extraction, and triplet generation form the core entity-relationship pipeline
  • Event detection and semantic role analysis add temporal and functional depth to extracted knowledge
  • Coreference resolution and provenance tracking ensure graph quality and auditability
  • Pluggable providers allow swapping between rule-based, spaCy, and LLM backends without pipeline rewrites
  • SemanticNetworkExtractor orchestrates complete pipelines into export-ready knowledge graphs

Frequently Asked Questions

How does Semantica handle extraction when LLM services are unavailable?

Semantica implements graceful degradation through rule-based fallbacks in every extractor class. The TripletExtractor automatically switches to rule-based generation when LLM providers are disabled, and RelationExtractor continues operating with pattern-matching when spaCy dependencies are missing. This design ensures production reliability regardless of external service availability.

Can I use only specific extractors without running the full pipeline?

Yes. Each extractor in the semantica.semantic_extract package operates as an independent class with its own extract_*() methods. The test suite demonstrates isolated usage patterns, and you can import NERExtractor, EventDetector, or any single component without invoking SemanticNetworkExtractor.

What graph databases are compatible with Semantica's output?

Semantica generates RDF-compatible knowledge graphs with standardized node-edge structures. According to the source implementation in semantic_network_extractor.py, the output integrates directly with Neo4j, Amazon Neptune, and other RDF stores. The provenance metadata included in each graph element supports database-specific import requirements.

How do I add a custom extraction provider to Semantica?

Custom providers implement the interface defined in semantica/semantic_extract/providers.py. You register your provider class and reference it via the method parameter in any extractor (e.g., NERExtractor(method="my_custom")). The provider abstraction layer handles initialization and method dispatch automatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →