# Semantica Semantic Extraction Capabilities: A Complete Guide to Knowledge Graph Pipeline

> Semantica provides nine core semantic extraction capabilities, including NER and relation extraction, to build queryable knowledge graphs with its modular Python pipeline. Discover its full potential.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-06

---

**Semantica offers nine core semantic extraction capabilities ranging from named-entity recognition and relation extraction to event detection and provenance tracking, all orchestrated through a modular Python pipeline that converts raw text into queryable RDF-style knowledge graphs.**

The open-source **Semantica** library (`semantica-agi/semantica`) provides a full-stack semantic extraction system designed for production knowledge graph construction. Its architecture separates each extraction concern into dedicated modules under `semantica.semantic_extract`, allowing developers to run complete pipelines or invoke individual extractors as needed.

## Core Semantic Extraction Capabilities

Semantica's extraction layer addresses every stage of the **knowledge graph construction pipeline**. Each capability is implemented as a focused Python class with swappable backends.

### Named-Entity Recognition (NER)

The **NERExtractor** class in [`semantica/semantic_extract/ner_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/ner_extractor.py) detects and classifies entities including people, organizations, locations, and dates. It supports three provider modes:

- **Pattern-based**: Rule-driven regex matching
- **spaCy**: Dependency parsing with linguistic features
- **LLM providers**: OpenAI, Ollama, or custom implementations

```python
from semantica.semantic_extract import NERExtractor

ner = NERExtractor(method="pattern")  # or "ollama", "openai"

entities = ner.extract_entities("Apple announced iPhone 15 in California.")

# → [Entity(text='Apple'), Entity(text='iPhone 15'), Entity(text='California')]

```

### Relation Extraction

The **RelationExtractor** in [`semantica/semantic_extract/relation_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/relation_extractor.py) identifies semantic relationships between entities. It implements:

- **Pattern-based extraction**: Predefined linguistic templates
- **Co-occurrence heuristics**: Proximity-based relationship scoring
- **Dependency-tree parsing**: Grammatical structure analysis

The module includes graceful degradation when spaCy dependencies are unavailable.

### Triplet Extraction

**TripletExtractor** ([`semantica/semantic_extract/triplet_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/triplet_extractor.py)) converts entities and relations into **RDF-style (subject, predicate, object) triples**. Key features include:

- Rule-based generation as fallback when LLM methods are disabled
- Validation and normalization of extracted triples
- Direct compatibility with graph database import formats

### Event Detection

The **EventDetector** class in [`semantica/semantic_extract/event_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/event_detector.py) extracts narrative events with temporal semantics. It identifies:

- Actions and their agents
- Timestamps and temporal expressions
- Event participants and their roles

This enables time-aware graph queries and temporal reasoning over document collections.

### Coreference Resolution

**CoreferenceResolver** ([`semantica/semantic_extract/coreference_resolver.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/coreference_resolver.py)) links pronouns and aliases to canonical entity mentions. This ensures:

- Consistent node IDs across the knowledge graph
- Reduced entity fragmentation
- Accurate relationship mapping for referred entities

### Semantic Role Analysis

The **SemanticAnalyzer** in [`semantica/semantic_extract/semantic_analyzer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/semantic_analyzer.py) performs **semantic role labeling**, assigning functional roles such as:

- **Agent**: The doer of an action
- **Patient**: The entity affected by an action
- **Instrument**: The means by which an action is performed

### Semantic Network Construction

**SemanticNetworkExtractor** ([`semantica/semantic_extract/semantic_network_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/semantic_network_extractor.py)) serves as the **high-level orchestrator**. It assembles:

- Extracted nodes (entities, events)
- Edges (relations, roles)
- Provenance metadata

The output is a coherent knowledge graph ready for persistence in **Neo4j**, **Amazon Neptune**, or other RDF stores.

### Provenance Tracking

Every extractor integrates **auditability features** capturing:

- Source document references
- Extraction timestamps
- Method metadata and provider configurations

This supports reproducibility and compliance requirements. The `progress_tracker` pattern used throughout the test suite demonstrates consistent provenance implementation.

### Pluggable Provider Architecture

The **provider abstraction layer** in [`semantica/semantic_extract/providers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/providers.py) enables backend swapping without pipeline changes. Supported providers include:

| Provider | Use Case |
|----------|----------|
| `PatternProvider` | Fast, deterministic extraction without external dependencies |
| `SpacyProvider` | Linguistically-informed analysis |
| `OllamaProvider` | Local LLM inference |
| `OpenAIProvider` | Cloud-based language model access |
| Custom implementations | Domain-specific extraction logic |

## Running a Complete Extraction Pipeline

The following example demonstrates end-to-end semantic extraction using Semantica's orchestrated pipeline:

```python
from semantica.semantic_extract import (
    NERExtractor,
    RelationExtractor,
    TripletExtractor,
    EventDetector,
    SemanticNetworkExtractor,
)

document = """Apple announced the new iPhone 15 in California. 
              Tim Cook said it will launch in September."""

# Individual extractor approach

ner = NERExtractor(method="pattern")
entities = ner.extract_entities(document)

rel_extractor = RelationExtractor(method="dependency")
relations = rel_extractor.extract_relations(document, entities)

triplets = TripletExtractor().extract_triplets(document, entities, relations)

events = EventDetector().extract_events(document, entities)

# Unified pipeline approach

network_extractor = SemanticNetworkExtractor()
semantic_graph = network_extractor.build(document)

# Contains nodes, edges, provenance metadata ready for export

```

Switching to LLM-backed extraction requires only parameter changes:

```python
ner = NERExtractor(method="ollama")      # Local LLM via Ollama

rel_extractor = RelationExtractor(method="openai")  # OpenAI API

```

## Module Reference and Source Files

| File Path | Purpose |
|-----------|---------|
| [`semantica/semantic_extract/ner_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/ner_extractor.py) | Entity detection with provider dispatch |
| [`semantica/semantic_extract/relation_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/relation_extractor.py) | Relationship identification strategies |
| [`semantica/semantic_extract/triplet_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/triplet_extractor.py) | RDF triple generation and validation |
| [`semantica/semantic_extract/event_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/event_detector.py) | Temporal event extraction |
| [`semantica/semantic_extract/coreference_resolver.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/coreference_resolver.py) | Pronoun and alias resolution |
| [`semantica/semantic_extract/semantic_analyzer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/semantic_analyzer.py) | Semantic role assignment |
| [`semantica/semantic_extract/semantic_network_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/semantic_network_extractor.py) | Pipeline orchestration and graph assembly |
| [`semantica/semantic_extract/providers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/providers.py) | Backend abstraction layer |
| [`semantica/semantic_extract/methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/methods.py) | Core algorithm implementations |

## Summary

- **Semantica provides nine specialized semantic extraction capabilities** through modular Python classes in `semantica.semantic_extract`
- **Named-entity recognition, relation extraction, and triplet generation** form the core entity-relationship pipeline
- **Event detection and semantic role analysis** add temporal and functional depth to extracted knowledge
- **Coreference resolution and provenance tracking** ensure graph quality and auditability
- **Pluggable providers** allow swapping between rule-based, spaCy, and LLM backends without pipeline rewrites
- **SemanticNetworkExtractor** orchestrates complete pipelines into export-ready knowledge graphs

## Frequently Asked Questions

### How does Semantica handle extraction when LLM services are unavailable?

Semantica implements **graceful degradation** through rule-based fallbacks in every extractor class. The `TripletExtractor` automatically switches to rule-based generation when LLM providers are disabled, and `RelationExtractor` continues operating with pattern-matching when spaCy dependencies are missing. This design ensures production reliability regardless of external service availability.

### Can I use only specific extractors without running the full pipeline?

Yes. Each extractor in the `semantica.semantic_extract` package operates as an **independent class** with its own `extract_*()` methods. The test suite demonstrates isolated usage patterns, and you can import `NERExtractor`, `EventDetector`, or any single component without invoking `SemanticNetworkExtractor`.

### What graph databases are compatible with Semantica's output?

Semantica generates **RDF-compatible knowledge graphs** with standardized node-edge structures. According to the source implementation in [`semantic_network_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantic_network_extractor.py), the output integrates directly with **Neo4j**, **Amazon Neptune**, and other RDF stores. The provenance metadata included in each graph element supports database-specific import requirements.

### How do I add a custom extraction provider to Semantica?

Custom providers implement the interface defined in [`semantica/semantic_extract/providers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/providers.py). You register your provider class and reference it via the `method` parameter in any extractor (e.g., `NERExtractor(method="my_custom")`). The provider abstraction layer handles initialization and method dispatch automatically.