# How Semantica Performs Named Entity Recognition (NER): A Deep Dive into the Multi-Method Pipeline

> Discover how Semantica performs Named Entity Recognition (NER) using a powerful multi-method pipeline. Explore pattern, regex, rules, spaCy, HuggingFace, and LLM extraction with intelligent fallback and scoring.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-06

---

**Semantica performs Named Entity Recognition through the `NERExtractor` class, which orchestrates six extraction methods—pattern, regex, rules, ML (spaCy), HuggingFace, and LLM—with automatic fallback chains, optional ensemble voting, and intelligent confidence scoring.**

Named Entity Recognition is a core capability of the Semantica framework, designed for flexibility in resource-constrained and production environments. The NER pipeline in `semantica-agi/semantica` adapts to available dependencies, gracefully degrades when advanced models fail, and can combine multiple extraction signals for improved accuracy. This article examines the architecture, method implementations, and practical configuration options based on the source code.

---

## Core Architecture: The NERExtractor Orchestrator

The `NERExtractor` class in [`semantica/semantic_extract/ner_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/ner_extractor.py) serves as the central dispatcher for all NER operations. It abstracts method selection, execution order, and result combination behind a unified interface.

### Constructor and Method Selection

The constructor accepts a `method` parameter that controls extraction strategy:

- **Single method**: `"pattern"`, `"regex"`, `"rules"`, `"ml"`, `"huggingface"`, or `"llm"`
- **Fallback chain**: A prioritized list like `["llm", "ml", "pattern"]`

When `"ml"` is selected, the extractor validates spaCy availability once and caches the loaded model via `methods.load_spacy_model()`. This prevents expensive reloading across multiple calls.

### The extract_entities() Execution Flow

The `extract_entities()` method implements the core orchestration logic:

1. Resolves the effective method list from `method` parameter
2. Filters out unusable methods (e.g., spaCy import failures)
3. Iterates through methods in priority order
4. Retrieves the concrete function via `methods.get_entity_method(method_name)`
5. Merges user-provided options and applies weighted confidence scoring when `entity_types` are specified
6. Stops at first success, or continues for ensemble voting

If **all methods fail**, the pipeline executes `_extract_fallback()`—a regex capturing capitalized words as a last-resort heuristic. Optional `_post_process_entities()` cleans boundaries and standardizes metadata.

---

## Six NER Methods Implemented in Semantica

All method implementations reside in [`semantica/semantic_extract/methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/methods.py). Each follows a consistent signature for plug-and-play compatibility with the orchestrator.

### Pattern-Based Extraction (extract_entities_pattern)

Lightweight regex patterns for common entity types:

- **PERSON**: Capitalized name sequences
- **ORG**: Organization indicators (Inc., Corp., Ltd.)
- **GPE**: Geographic locations
- **DATE**: Numeric date formats

Best for: Rapid prototyping, environments without ML dependencies.

### Regex-Based Extraction (extract_entities_regex)

Extends pattern matching with customizable rules. Default patterns cover PERSON, ORG, GPE, DATE, MONEY, and PERCENT. Users inject custom patterns via the `patterns` parameter.

### Rule-Based Extraction (extract_entities_rules)

Simple linguistic heuristic: capitalized sentence-initial words are classified as **PERSON**. Minimal computational overhead, useful as a baseline.

### ML-Based Extraction (extract_entities_ml)

Loads spaCy models (default: `en_core_web_sm`) and extracts `doc.ents`. Implementation includes:

- Graceful degradation to pattern extraction if spaCy unavailable
- Model caching via `_spacy_model_cache`
- Direct mapping of spaCy labels to Semantica's `Entity` type

### HuggingFace Extraction (extract_entities_huggingface)

Uses `HuggingFaceModelLoader` for transformer-based token classification:

- Auto-detects IOB tag formats from pipeline outputs
- Performs manual aggregation when needed
- Supports any `token-classification` pipeline model

Default model: `dslim/bert-base-NER` (configurable).

### LLM-Based Extraction (extract_entities_llm)

Structured prompting with JSON output validation:

- Constructs entity extraction prompts with Pydantic schema validation
- Calls providers via `create_provider()` (OpenAI, Gemini, Groq, etc.)
- Implements **input chunking** for oversized documents
- Caches results to avoid repeat API calls

---

## Advanced Features: Confidence Scoring and Ensemble Voting

### Weighted Confidence Calculation

`calculate_weighted_confidence()` blends raw method confidence with similarity to user-specified `entity_types`. The similarity calculation in `find_best_match_index()` uses:

- Fast substring matching
- Synonym expansion
- Embedding similarity
- Fuzzy string matching

This allows prioritization of semantically relevant entities even when confidence scores are moderate.

### Ensemble Voting (_vote_entities)

When `ensemble_voting=True`, the extractor:

1. Collects entities from **all successful methods**
2. Applies majority-vote aggregation
3. Filters by average confidence threshold (default: 0.5)

This multi-signal approach reduces false positives from any single method.

---

## Practical Code Examples

### Basic ML Extraction

```python
from semantica.semantic_extract import NERExtractor

text = "Apple Inc. was founded by Steve Jobs in 1976."

extractor = NERExtractor()  # defaults to "ml"

entities = extractor.extract_entities(text)
print(entities)

# → [Entity(text='Apple Inc.', label='ORG', ...),

#    Entity(text='Steve Jobs', label='PERSON', ...),

#    Entity(text='1976', label='DATE', ...)]

```

### Fallback Chain Configuration

```python

# Try LLM first, fall back to spaCy, then pattern

extractor = NERExtractor(
    method=["llm", "ml", "pattern"],
    ensemble_voting=False
)
entities = extractor.extract_entities(text)

# Returns first successful method's output

```

### Ensemble Voting

```python

# Combine LLM and spaCy signals

extractor = NERExtractor(
    method=["llm", "ml"],
    ensemble_voting=True
)
entities = extractor.extract_entities(text)

# Entities after majority voting across both methods

```

### Custom Regex Patterns

```python
custom_patterns = {
    "PRODUCT": r"\b([A-Z][a-zA-Z]+(?:\s+[A-Z][a-zA-Z]+)*)\b"
}
extractor = NERExtractor(
    method="regex",
    patterns=custom_patterns
)
entities = extractor.extract_entities("I love using OpenAI GPT-4.")

# Detects "OpenAI GPT-4" as PRODUCT

```

### HuggingFace Integration

```python
extractor = NERExtractor(
    method="huggingface",
    huggingface_model="dslim/bert-base-NER"
)
entities = extractor.extract_entities(text)

```

---

## Key Source Files and Components

| File | Purpose | Key Functions/Classes |
|------|---------|----------------------|
| [`semantica/semantic_extract/ner_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/ner_extractor.py) | High-level orchestrator | `NERExtractor`, `extract_entities()`, `_vote_entities()` |
| [`semantica/semantic_extract/methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/methods.py) | Method implementations | `extract_entities_pattern()`, `extract_entities_ml()`, `get_entity_method()`, `load_spacy_model()` |
| [`semantica/semantic_extract/types.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/types.py) | Core data structures | `Entity`, `Relation`, `Triplet` dataclasses |
| [`semantica/semantic_extract/providers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/providers.py) | LLM provider factory | `create_provider()` |
| [`semantica/semantic_extract/cache.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/cache.py) | Process-level caching | `_spacy_model_cache`, `_embedder_cache` |

---

## Summary

Semantica's Named Entity Recognition implementation offers production-ready flexibility through these key design decisions:

- **Modular architecture**: Six pluggable extraction methods from regex to LLMs
- **Intelligent fallback**: Automatic degradation when preferred methods fail
- **Performance optimization**: Process-level caching for spaCy models and HuggingFace pipelines
- **Quality controls**: Weighted confidence scoring and optional ensemble voting
- **Operational resilience**: Regex fallback ensures entity extraction even in complete failure scenarios

The `NERExtractor` class abstracts complexity while exposing fine-grained control through method chains, custom patterns, and confidence thresholds.

---

## Frequently Asked Questions

### What NER methods does Semantica support?

Semantica supports six extraction methods: **pattern** (lightweight regexes), **regex** (customizable patterns), **rules** (linguistic heuristics), **ml** (spaCy models), **huggingface** (transformer token classification), and **llm** (structured prompting with any provider). These are configured via the `method` parameter in `NERExtractor`.

### How does Semantica handle missing dependencies like spaCy?

The `extract_entities()` method filters unusable methods before execution. If spaCy fails to import or the model fails to load, that method is skipped. If **all methods fail**, the pipeline executes `_extract_fallback()`—a regex capturing capitalized words. This ensures entity extraction continues regardless of environment constraints.

### Can Semantica combine multiple NER methods?

Yes. Set `ensemble_voting=True` to enable `_vote_entities()`, which aggregates results from all successful methods and applies majority voting with a confidence threshold (default 0.5). Without voting, methods are tried in order until one succeeds, creating a **fallback chain** behavior.

### How do I use a custom HuggingFace model for NER?

Pass `method="huggingface"` with `huggingface_model` set to your model identifier:

```python
extractor = NERExtractor(
    method="huggingface",
    huggingface_model="your-username/your-ner-model"
)

```

The implementation in `extract_entities_huggingface()` auto-detects IOB formats and handles manual aggregation when needed.