Semantica EntityResolver Parameters: Complete Configuration Guide for Knowledge Graph Deduplication

Semantica's EntityResolver accepts three primary parameters—strategy (matching algorithm), similarity_threshold (minimum duplicate score), and **config (dictionary for deduplication and merger settings)—defined in the constructor at lines 51-53 of semantica/kg/entity_resolver.py.

The EntityResolver class in the semantica-agi/semantica repository serves as the core component for deduplicating and merging entities when building knowledge graphs. Understanding the Semantica EntityResolver parameters is essential for tuning entity matching behavior, whether you require fuzzy string matching, exact equality comparisons, or semantic similarity using embeddings.

Core EntityResolver Parameters

The constructor signature in semantica/kg/entity_resolver.py exposes three distinct configuration options that control how entities are identified as duplicates and subsequently merged.

strategy

The strategy parameter accepts a string that determines the matching algorithm used to identify duplicate entities.

  • Type: str
  • Default: "fuzzy"
  • Options:
    • "fuzzy" – Uses fuzzy string similarity algorithms (default behavior)
    • "exact" – Requires exact string equality between entity identifiers
    • "semantic" – Employs semantic similarity using vector embeddings

similarity_threshold

The similarity_threshold parameter sets the minimum similarity score required for two entities to be considered duplicates when using "fuzzy" or "semantic" strategies.

  • Type: float
  • Default: 0.7
  • Range: 0.0 to 1.0

Lower values make the resolver more permissive (catching more potential duplicates), while higher values enforce stricter matching criteria. This threshold is ignored when using the "exact" strategy, which implicitly requires a perfect match (1.0).

**config

The **config parameter accepts an optional dictionary that forwards configuration to internal components, enabling fine-grained control over the deduplication pipeline.

  • Type: dict
  • Default: {}

Valid configuration keys include:

  • deduplication – Passed to DuplicateDetector (e.g., custom similarity functions, blocking techniques, or ignore_case flags)
  • merger – Passed to EntityMerger (e.g., field-selection rules like prefer_field, conflict-resolution policies like keep_longest_description)

Source Code Implementation

According to the semantica-agi/semantica source code, the EntityResolver class is implemented in semantica/kg/entity_resolver.py, with the constructor handling parameter validation at lines 51-53. The class exposes public methods including resolve_entities() and merge_duplicates(), along with internal helpers that consume the supplied configuration.

The resolver delegates specific tasks to specialized modules:

Configuration Examples

Basic Usage with Default Fuzzy Strategy

from semantica.kg import EntityResolver

# strategy="fuzzy", threshold=0.7 (defaults)

resolver = EntityResolver()

entities = [
    {"id": "1", "name": "Apple Inc.", "type": "Company"},
    {"id": "2", "name": "Apple", "type": "Company"},
    {"id": "3", "name": "Microsoft", "type": "Company"},
]

resolved = resolver.resolve_entities(entities)
print(resolved)  # Returns two entities: merged Apple entry + Microsoft

Exact-Match Strategy with Custom Configuration

resolver = EntityResolver(
    strategy="exact",
    similarity_threshold=1.0,
    deduplication={"ignore_case": True},
    merger={"prefer_field": "name"}
)
resolved = resolver.resolve_entities(entities)

Semantic Similarity with High Precision

resolver = EntityResolver(
    strategy="semantic",
    similarity_threshold=0.85,
    merger={"keep_longest_description": True}
)
resolved = resolver.resolve_entities(entities)

Summary

  • Three primary parameters control the EntityResolver: strategy for algorithm selection, similarity_threshold for match strictness, and **config for component-specific settings.
  • Strategy options include fuzzy string matching ("fuzzy"), exact equality ("exact"), and embedding-based semantic similarity ("semantic").
  • Configuration dictionary accepts deduplication and merger keys to customize the DuplicateDetector and EntityMerger behavior.
  • Source location is semantica/kg/entity_resolver.py (lines 51-53), with dependencies in the deduplication/ and utils/ modules.

Frequently Asked Questions

What are the valid values for the strategy parameter in Semantica's EntityResolver?

The strategy parameter accepts three string values: "fuzzy" for Levenshtein-style string similarity (default), "exact" for literal string equality, and "semantic" for vector-based embedding similarity. Each strategy uses the DuplicateDetector in semantica/deduplication/duplicate_detector.py with different comparison logic.

How does the similarity_threshold affect entity deduplication?

The similarity_threshold is a float between 0.0 and 1.0 that sets the minimum score required to flag two entities as duplicates when using "fuzzy" or "semantic" strategies. The default value of 0.7 provides balanced recall and precision; increasing to 0.85 or higher reduces false positives but may miss variant spellings or abbreviations.

Can I configure custom deduplication logic through the config parameter?

Yes, the **config dictionary accepts a deduplication key that forwards settings to the DuplicateDetector class, allowing you to specify custom similarity functions, blocking techniques, or case-sensitivity rules (ignore_case). Similarly, the merger key configures the EntityMerger with field-selection policies like prefer_field or keep_longest_description.

Where is the EntityResolver class implemented in the Semantica repository?

The EntityResolver class is implemented in semantica/kg/entity_resolver.py within the semantica-agi/semantica repository. The constructor parameters are defined at lines 51-53, and the class utilizes semantica/deduplication/duplicate_detector.py for matching logic and semantica/deduplication/entity_merger.py for consolidating duplicate groups.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →