How Semantica Performs Named Entity Recognition (NER): A Deep Dive into the Multi-Method Pipeline

Semantica performs Named Entity Recognition through the NERExtractor class, which orchestrates six extraction methods—pattern, regex, rules, ML (spaCy), HuggingFace, and LLM—with automatic fallback chains, optional ensemble voting, and intelligent confidence scoring.

Named Entity Recognition is a core capability of the Semantica framework, designed for flexibility in resource-constrained and production environments. The NER pipeline in semantica-agi/semantica adapts to available dependencies, gracefully degrades when advanced models fail, and can combine multiple extraction signals for improved accuracy. This article examines the architecture, method implementations, and practical configuration options based on the source code.


Core Architecture: The NERExtractor Orchestrator

The NERExtractor class in semantica/semantic_extract/ner_extractor.py serves as the central dispatcher for all NER operations. It abstracts method selection, execution order, and result combination behind a unified interface.

Constructor and Method Selection

The constructor accepts a method parameter that controls extraction strategy:

  • Single method: "pattern", "regex", "rules", "ml", "huggingface", or "llm"
  • Fallback chain: A prioritized list like ["llm", "ml", "pattern"]

When "ml" is selected, the extractor validates spaCy availability once and caches the loaded model via methods.load_spacy_model(). This prevents expensive reloading across multiple calls.

The extract_entities() Execution Flow

The extract_entities() method implements the core orchestration logic:

  1. Resolves the effective method list from method parameter
  2. Filters out unusable methods (e.g., spaCy import failures)
  3. Iterates through methods in priority order
  4. Retrieves the concrete function via methods.get_entity_method(method_name)
  5. Merges user-provided options and applies weighted confidence scoring when entity_types are specified
  6. Stops at first success, or continues for ensemble voting

If all methods fail, the pipeline executes _extract_fallback()—a regex capturing capitalized words as a last-resort heuristic. Optional _post_process_entities() cleans boundaries and standardizes metadata.


Six NER Methods Implemented in Semantica

All method implementations reside in semantica/semantic_extract/methods.py. Each follows a consistent signature for plug-and-play compatibility with the orchestrator.

Pattern-Based Extraction (extract_entities_pattern)

Lightweight regex patterns for common entity types:

  • PERSON: Capitalized name sequences
  • ORG: Organization indicators (Inc., Corp., Ltd.)
  • GPE: Geographic locations
  • DATE: Numeric date formats

Best for: Rapid prototyping, environments without ML dependencies.

Regex-Based Extraction (extract_entities_regex)

Extends pattern matching with customizable rules. Default patterns cover PERSON, ORG, GPE, DATE, MONEY, and PERCENT. Users inject custom patterns via the patterns parameter.

Rule-Based Extraction (extract_entities_rules)

Simple linguistic heuristic: capitalized sentence-initial words are classified as PERSON. Minimal computational overhead, useful as a baseline.

ML-Based Extraction (extract_entities_ml)

Loads spaCy models (default: en_core_web_sm) and extracts doc.ents. Implementation includes:

  • Graceful degradation to pattern extraction if spaCy unavailable
  • Model caching via _spacy_model_cache
  • Direct mapping of spaCy labels to Semantica's Entity type

HuggingFace Extraction (extract_entities_huggingface)

Uses HuggingFaceModelLoader for transformer-based token classification:

  • Auto-detects IOB tag formats from pipeline outputs
  • Performs manual aggregation when needed
  • Supports any token-classification pipeline model

Default model: dslim/bert-base-NER (configurable).

LLM-Based Extraction (extract_entities_llm)

Structured prompting with JSON output validation:

  • Constructs entity extraction prompts with Pydantic schema validation
  • Calls providers via create_provider() (OpenAI, Gemini, Groq, etc.)
  • Implements input chunking for oversized documents
  • Caches results to avoid repeat API calls

Advanced Features: Confidence Scoring and Ensemble Voting

Weighted Confidence Calculation

calculate_weighted_confidence() blends raw method confidence with similarity to user-specified entity_types. The similarity calculation in find_best_match_index() uses:

  • Fast substring matching
  • Synonym expansion
  • Embedding similarity
  • Fuzzy string matching

This allows prioritization of semantically relevant entities even when confidence scores are moderate.

Ensemble Voting (_vote_entities)

When ensemble_voting=True, the extractor:

  1. Collects entities from all successful methods
  2. Applies majority-vote aggregation
  3. Filters by average confidence threshold (default: 0.5)

This multi-signal approach reduces false positives from any single method.


Practical Code Examples

Basic ML Extraction

from semantica.semantic_extract import NERExtractor

text = "Apple Inc. was founded by Steve Jobs in 1976."

extractor = NERExtractor()  # defaults to "ml"

entities = extractor.extract_entities(text)
print(entities)

# → [Entity(text='Apple Inc.', label='ORG', ...),

#    Entity(text='Steve Jobs', label='PERSON', ...),

#    Entity(text='1976', label='DATE', ...)]

Fallback Chain Configuration


# Try LLM first, fall back to spaCy, then pattern

extractor = NERExtractor(
    method=["llm", "ml", "pattern"],
    ensemble_voting=False
)
entities = extractor.extract_entities(text)

# Returns first successful method's output

Ensemble Voting


# Combine LLM and spaCy signals

extractor = NERExtractor(
    method=["llm", "ml"],
    ensemble_voting=True
)
entities = extractor.extract_entities(text)

# Entities after majority voting across both methods

Custom Regex Patterns

custom_patterns = {
    "PRODUCT": r"\b([A-Z][a-zA-Z]+(?:\s+[A-Z][a-zA-Z]+)*)\b"
}
extractor = NERExtractor(
    method="regex",
    patterns=custom_patterns
)
entities = extractor.extract_entities("I love using OpenAI GPT-4.")

# Detects "OpenAI GPT-4" as PRODUCT

HuggingFace Integration

extractor = NERExtractor(
    method="huggingface",
    huggingface_model="dslim/bert-base-NER"
)
entities = extractor.extract_entities(text)

Key Source Files and Components

File Purpose Key Functions/Classes
semantica/semantic_extract/ner_extractor.py High-level orchestrator NERExtractor, extract_entities(), _vote_entities()
semantica/semantic_extract/methods.py Method implementations extract_entities_pattern(), extract_entities_ml(), get_entity_method(), load_spacy_model()
semantica/semantic_extract/types.py Core data structures Entity, Relation, Triplet dataclasses
semantica/semantic_extract/providers.py LLM provider factory create_provider()
semantica/semantic_extract/cache.py Process-level caching _spacy_model_cache, _embedder_cache

Summary

Semantica's Named Entity Recognition implementation offers production-ready flexibility through these key design decisions:

  • Modular architecture: Six pluggable extraction methods from regex to LLMs
  • Intelligent fallback: Automatic degradation when preferred methods fail
  • Performance optimization: Process-level caching for spaCy models and HuggingFace pipelines
  • Quality controls: Weighted confidence scoring and optional ensemble voting
  • Operational resilience: Regex fallback ensures entity extraction even in complete failure scenarios

The NERExtractor class abstracts complexity while exposing fine-grained control through method chains, custom patterns, and confidence thresholds.


Frequently Asked Questions

What NER methods does Semantica support?

Semantica supports six extraction methods: pattern (lightweight regexes), regex (customizable patterns), rules (linguistic heuristics), ml (spaCy models), huggingface (transformer token classification), and llm (structured prompting with any provider). These are configured via the method parameter in NERExtractor.

How does Semantica handle missing dependencies like spaCy?

The extract_entities() method filters unusable methods before execution. If spaCy fails to import or the model fails to load, that method is skipped. If all methods fail, the pipeline executes _extract_fallback()—a regex capturing capitalized words. This ensures entity extraction continues regardless of environment constraints.

Can Semantica combine multiple NER methods?

Yes. Set ensemble_voting=True to enable _vote_entities(), which aggregates results from all successful methods and applies majority voting with a confidence threshold (default 0.5). Without voting, methods are tried in order until one succeeds, creating a fallback chain behavior.

How do I use a custom HuggingFace model for NER?

Pass method="huggingface" with huggingface_model set to your model identifier:

extractor = NERExtractor(
    method="huggingface",
    huggingface_model="your-username/your-ner-model"
)

The implementation in extract_entities_huggingface() auto-detects IOB formats and handles manual aggregation when needed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →