Semantica EntityResolver Parameters: Complete Configuration Guide for Knowledge Graph Deduplication
Semantica's EntityResolver accepts three primary parameters—strategy (matching algorithm), similarity_threshold (minimum duplicate score), and **config (dictionary for deduplication and merger settings)—defined in the constructor at lines 51-53 of semantica/kg/entity_resolver.py.
The EntityResolver class in the semantica-agi/semantica repository serves as the core component for deduplicating and merging entities when building knowledge graphs. Understanding the Semantica EntityResolver parameters is essential for tuning entity matching behavior, whether you require fuzzy string matching, exact equality comparisons, or semantic similarity using embeddings.
Core EntityResolver Parameters
The constructor signature in semantica/kg/entity_resolver.py exposes three distinct configuration options that control how entities are identified as duplicates and subsequently merged.
strategy
The strategy parameter accepts a string that determines the matching algorithm used to identify duplicate entities.
- Type:
str - Default:
"fuzzy" - Options:
"fuzzy"– Uses fuzzy string similarity algorithms (default behavior)"exact"– Requires exact string equality between entity identifiers"semantic"– Employs semantic similarity using vector embeddings
similarity_threshold
The similarity_threshold parameter sets the minimum similarity score required for two entities to be considered duplicates when using "fuzzy" or "semantic" strategies.
- Type:
float - Default:
0.7 - Range:
0.0to1.0
Lower values make the resolver more permissive (catching more potential duplicates), while higher values enforce stricter matching criteria. This threshold is ignored when using the "exact" strategy, which implicitly requires a perfect match (1.0).
**config
The **config parameter accepts an optional dictionary that forwards configuration to internal components, enabling fine-grained control over the deduplication pipeline.
- Type:
dict - Default:
{}
Valid configuration keys include:
deduplication– Passed toDuplicateDetector(e.g., custom similarity functions, blocking techniques, orignore_caseflags)merger– Passed toEntityMerger(e.g., field-selection rules likeprefer_field, conflict-resolution policies likekeep_longest_description)
Source Code Implementation
According to the semantica-agi/semantica source code, the EntityResolver class is implemented in semantica/kg/entity_resolver.py, with the constructor handling parameter validation at lines 51-53. The class exposes public methods including resolve_entities() and merge_duplicates(), along with internal helpers that consume the supplied configuration.
The resolver delegates specific tasks to specialized modules:
semantica/deduplication/duplicate_detector.py– Implements the duplicate-detection logic configurable via thededuplicationkeysemantica/deduplication/entity_merger.py– Handles the merging of duplicate groups, receiving settings through themergerkeysemantica/utils/entity_ids.py– Provides utilities for extracting stable entity identifierssemantica/utils/logging.py– Supplies the centralized logger used throughout the resolver
Configuration Examples
Basic Usage with Default Fuzzy Strategy
from semantica.kg import EntityResolver
# strategy="fuzzy", threshold=0.7 (defaults)
resolver = EntityResolver()
entities = [
{"id": "1", "name": "Apple Inc.", "type": "Company"},
{"id": "2", "name": "Apple", "type": "Company"},
{"id": "3", "name": "Microsoft", "type": "Company"},
]
resolved = resolver.resolve_entities(entities)
print(resolved) # Returns two entities: merged Apple entry + Microsoft
Exact-Match Strategy with Custom Configuration
resolver = EntityResolver(
strategy="exact",
similarity_threshold=1.0,
deduplication={"ignore_case": True},
merger={"prefer_field": "name"}
)
resolved = resolver.resolve_entities(entities)
Semantic Similarity with High Precision
resolver = EntityResolver(
strategy="semantic",
similarity_threshold=0.85,
merger={"keep_longest_description": True}
)
resolved = resolver.resolve_entities(entities)
Summary
- Three primary parameters control the
EntityResolver:strategyfor algorithm selection,similarity_thresholdfor match strictness, and**configfor component-specific settings. - Strategy options include fuzzy string matching (
"fuzzy"), exact equality ("exact"), and embedding-based semantic similarity ("semantic"). - Configuration dictionary accepts
deduplicationandmergerkeys to customize theDuplicateDetectorandEntityMergerbehavior. - Source location is
semantica/kg/entity_resolver.py(lines 51-53), with dependencies in thededuplication/andutils/modules.
Frequently Asked Questions
What are the valid values for the strategy parameter in Semantica's EntityResolver?
The strategy parameter accepts three string values: "fuzzy" for Levenshtein-style string similarity (default), "exact" for literal string equality, and "semantic" for vector-based embedding similarity. Each strategy uses the DuplicateDetector in semantica/deduplication/duplicate_detector.py with different comparison logic.
How does the similarity_threshold affect entity deduplication?
The similarity_threshold is a float between 0.0 and 1.0 that sets the minimum score required to flag two entities as duplicates when using "fuzzy" or "semantic" strategies. The default value of 0.7 provides balanced recall and precision; increasing to 0.85 or higher reduces false positives but may miss variant spellings or abbreviations.
Can I configure custom deduplication logic through the config parameter?
Yes, the **config dictionary accepts a deduplication key that forwards settings to the DuplicateDetector class, allowing you to specify custom similarity functions, blocking techniques, or case-sensitivity rules (ignore_case). Similarly, the merger key configures the EntityMerger with field-selection policies like prefer_field or keep_longest_description.
Where is the EntityResolver class implemented in the Semantica repository?
The EntityResolver class is implemented in semantica/kg/entity_resolver.py within the semantica-agi/semantica repository. The constructor parameters are defined at lines 51-53, and the class utilizes semantica/deduplication/duplicate_detector.py for matching logic and semantica/deduplication/entity_merger.py for consolidating duplicate groups.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →