Semantica Semantic Extraction Capabilities: A Complete Guide to Knowledge Graph Pipeline
Semantica offers nine core semantic extraction capabilities ranging from named-entity recognition and relation extraction to event detection and provenance tracking, all orchestrated through a modular Python pipeline that converts raw text into queryable RDF-style knowledge graphs.
The open-source Semantica library (semantica-agi/semantica) provides a full-stack semantic extraction system designed for production knowledge graph construction. Its architecture separates each extraction concern into dedicated modules under semantica.semantic_extract, allowing developers to run complete pipelines or invoke individual extractors as needed.
Core Semantic Extraction Capabilities
Semantica's extraction layer addresses every stage of the knowledge graph construction pipeline. Each capability is implemented as a focused Python class with swappable backends.
Named-Entity Recognition (NER)
The NERExtractor class in semantica/semantic_extract/ner_extractor.py detects and classifies entities including people, organizations, locations, and dates. It supports three provider modes:
- Pattern-based: Rule-driven regex matching
- spaCy: Dependency parsing with linguistic features
- LLM providers: OpenAI, Ollama, or custom implementations
from semantica.semantic_extract import NERExtractor
ner = NERExtractor(method="pattern") # or "ollama", "openai"
entities = ner.extract_entities("Apple announced iPhone 15 in California.")
# → [Entity(text='Apple'), Entity(text='iPhone 15'), Entity(text='California')]
Relation Extraction
The RelationExtractor in semantica/semantic_extract/relation_extractor.py identifies semantic relationships between entities. It implements:
- Pattern-based extraction: Predefined linguistic templates
- Co-occurrence heuristics: Proximity-based relationship scoring
- Dependency-tree parsing: Grammatical structure analysis
The module includes graceful degradation when spaCy dependencies are unavailable.
Triplet Extraction
TripletExtractor (semantica/semantic_extract/triplet_extractor.py) converts entities and relations into RDF-style (subject, predicate, object) triples. Key features include:
- Rule-based generation as fallback when LLM methods are disabled
- Validation and normalization of extracted triples
- Direct compatibility with graph database import formats
Event Detection
The EventDetector class in semantica/semantic_extract/event_detector.py extracts narrative events with temporal semantics. It identifies:
- Actions and their agents
- Timestamps and temporal expressions
- Event participants and their roles
This enables time-aware graph queries and temporal reasoning over document collections.
Coreference Resolution
CoreferenceResolver (semantica/semantic_extract/coreference_resolver.py) links pronouns and aliases to canonical entity mentions. This ensures:
- Consistent node IDs across the knowledge graph
- Reduced entity fragmentation
- Accurate relationship mapping for referred entities
Semantic Role Analysis
The SemanticAnalyzer in semantica/semantic_extract/semantic_analyzer.py performs semantic role labeling, assigning functional roles such as:
- Agent: The doer of an action
- Patient: The entity affected by an action
- Instrument: The means by which an action is performed
Semantic Network Construction
SemanticNetworkExtractor (semantica/semantic_extract/semantic_network_extractor.py) serves as the high-level orchestrator. It assembles:
- Extracted nodes (entities, events)
- Edges (relations, roles)
- Provenance metadata
The output is a coherent knowledge graph ready for persistence in Neo4j, Amazon Neptune, or other RDF stores.
Provenance Tracking
Every extractor integrates auditability features capturing:
- Source document references
- Extraction timestamps
- Method metadata and provider configurations
This supports reproducibility and compliance requirements. The progress_tracker pattern used throughout the test suite demonstrates consistent provenance implementation.
Pluggable Provider Architecture
The provider abstraction layer in semantica/semantic_extract/providers.py enables backend swapping without pipeline changes. Supported providers include:
| Provider | Use Case |
|---|---|
PatternProvider |
Fast, deterministic extraction without external dependencies |
SpacyProvider |
Linguistically-informed analysis |
OllamaProvider |
Local LLM inference |
OpenAIProvider |
Cloud-based language model access |
| Custom implementations | Domain-specific extraction logic |
Running a Complete Extraction Pipeline
The following example demonstrates end-to-end semantic extraction using Semantica's orchestrated pipeline:
from semantica.semantic_extract import (
NERExtractor,
RelationExtractor,
TripletExtractor,
EventDetector,
SemanticNetworkExtractor,
)
document = """Apple announced the new iPhone 15 in California.
Tim Cook said it will launch in September."""
# Individual extractor approach
ner = NERExtractor(method="pattern")
entities = ner.extract_entities(document)
rel_extractor = RelationExtractor(method="dependency")
relations = rel_extractor.extract_relations(document, entities)
triplets = TripletExtractor().extract_triplets(document, entities, relations)
events = EventDetector().extract_events(document, entities)
# Unified pipeline approach
network_extractor = SemanticNetworkExtractor()
semantic_graph = network_extractor.build(document)
# Contains nodes, edges, provenance metadata ready for export
Switching to LLM-backed extraction requires only parameter changes:
ner = NERExtractor(method="ollama") # Local LLM via Ollama
rel_extractor = RelationExtractor(method="openai") # OpenAI API
Module Reference and Source Files
| File Path | Purpose |
|---|---|
semantica/semantic_extract/ner_extractor.py |
Entity detection with provider dispatch |
semantica/semantic_extract/relation_extractor.py |
Relationship identification strategies |
semantica/semantic_extract/triplet_extractor.py |
RDF triple generation and validation |
semantica/semantic_extract/event_detector.py |
Temporal event extraction |
semantica/semantic_extract/coreference_resolver.py |
Pronoun and alias resolution |
semantica/semantic_extract/semantic_analyzer.py |
Semantic role assignment |
semantica/semantic_extract/semantic_network_extractor.py |
Pipeline orchestration and graph assembly |
semantica/semantic_extract/providers.py |
Backend abstraction layer |
semantica/semantic_extract/methods.py |
Core algorithm implementations |
Summary
- Semantica provides nine specialized semantic extraction capabilities through modular Python classes in
semantica.semantic_extract - Named-entity recognition, relation extraction, and triplet generation form the core entity-relationship pipeline
- Event detection and semantic role analysis add temporal and functional depth to extracted knowledge
- Coreference resolution and provenance tracking ensure graph quality and auditability
- Pluggable providers allow swapping between rule-based, spaCy, and LLM backends without pipeline rewrites
- SemanticNetworkExtractor orchestrates complete pipelines into export-ready knowledge graphs
Frequently Asked Questions
How does Semantica handle extraction when LLM services are unavailable?
Semantica implements graceful degradation through rule-based fallbacks in every extractor class. The TripletExtractor automatically switches to rule-based generation when LLM providers are disabled, and RelationExtractor continues operating with pattern-matching when spaCy dependencies are missing. This design ensures production reliability regardless of external service availability.
Can I use only specific extractors without running the full pipeline?
Yes. Each extractor in the semantica.semantic_extract package operates as an independent class with its own extract_*() methods. The test suite demonstrates isolated usage patterns, and you can import NERExtractor, EventDetector, or any single component without invoking SemanticNetworkExtractor.
What graph databases are compatible with Semantica's output?
Semantica generates RDF-compatible knowledge graphs with standardized node-edge structures. According to the source implementation in semantic_network_extractor.py, the output integrates directly with Neo4j, Amazon Neptune, and other RDF stores. The provenance metadata included in each graph element supports database-specific import requirements.
How do I add a custom extraction provider to Semantica?
Custom providers implement the interface defined in semantica/semantic_extract/providers.py. You register your provider class and reference it via the method parameter in any extractor (e.g., NERExtractor(method="my_custom")). The provider abstraction layer handles initialization and method dispatch automatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →