How to Use Semantic Similarity for Entity Deduplication in Semantica: A Complete Developer Guide
Semantica resolves duplicate entities by computing multi-factor semantic similarity scores—combining string, property, relationship, and embedding metrics—then grouping or merging pairs that exceed configurable thresholds.
Entity deduplication is critical for maintaining clean knowledge graphs, and the Semantica framework provides a robust pipeline for identifying and merging duplicate entities using semantic similarity. The system evaluates entities across multiple dimensions to generate confidence-weighted similarity scores, enabling both pairwise matching and transitive clustering. This guide walks through the core components, workflow, and implementation patterns found in the semantica-agi/semantica repository.
Core Deduplication Components
The deduplication architecture consists of three primary classes that handle calculation, detection orchestration, and API convenience.
SimilarityCalculator
Located in semantica/deduplication/similarity_calculator.py, the SimilarityCalculator class computes normalized similarity scores across four distinct factors. It implements a fast pre-filter via _prefilter_pair to reject obvious non-matches before expensive calculations occur. The calculator evaluates string similarity (Levenshtein, Jaro-Winkler, cosine), property similarity (weighted value comparison), relationship similarity (Jaccard index of relationship sets), and embedding similarity (vector cosine similarity). Individual component scores are aggregated using configurable weights—embedding_weight, string_weight, property_weight, and relationship_weight—which are automatically normalized if they do not sum to 1.0 (lines 44-50).
DuplicateDetector
The DuplicateDetector class in semantica/deduplication/duplicate_detector.py wraps the calculator and manages the detection pipeline. It applies configurable similarity_threshold and confidence_threshold values, supports optional min_similarity floors, and implements ranking controls via max_results and top_k_per_entity. The sort_by parameter allows sorting by either confidence or similarity_score. When use_clustering=True, the detector employs a union-find algorithm to group transitive duplicates into DuplicateGroup objects, automatically selecting the most complete entity as the representative.
High-Level API Methods
For simplified access, semantica/deduplication/methods.py exposes convenience functions including detect_duplicates(), calculate_similarity(), build_clusters(), and merge_entities(). These wrappers interface with the detector and calculator while supporting custom strategies through the method registry defined in semantica/deduplication/registry.py.
The Semantic Similarity Workflow
The deduplication process follows a seven-stage pipeline that transforms raw entities into canonical, merged records.
1. Pre-processing and Normalization
Before similarity calculation, entity names undergo normalization via _lower_name and optional EntityNormalizer components. This canonicalization of punctuation and case improves accuracy during the string similarity stage.
2. Fast Pre-filtering
The _prefilter_pair method quickly rejects candidate pairs that differ in entity type or exhibit extreme name-length ratio disparities. This optimization prevents wasted computation on obvious non-matches.
3. Multi-Factor Score Calculation
For each surviving pair, the calculator evaluates the four component scores: string distance metrics, property value overlaps, relationship set Jaccard similarity, and embedding vector cosine similarity.
4. Weighted Aggregation
Component scores combine according to user-defined weights. The calculator automatically renormalizes weights that do not sum to unity, ensuring consistent scoring behavior.
5. Thresholding and Ranking
The detect_duplicates() method filters candidates against similarity_threshold and confidence_threshold, then applies result limits. Users can constrain output to top_k_per_entity matches or an absolute max_results ceiling.
6. Transitive Group Formation
When clustering is enabled, the union-find algorithm identifies transitive duplicate relationships (A matches B, B matches C, therefore A matches C) and collapses them into DuplicateGroup instances.
7. Entity Merging
Detected groups feed into EntityMerger (accessible via merge_entities()), which applies strategies like keep_most_complete to consolidate records while preserving provenance metadata.
Implementation Examples
Basic Duplicate Detection
from semantica.deduplication import DuplicateDetector
entities = [
{"id": "e1", "name": "Apple Inc.", "type": "Company"},
{"id": "e2", "name": "Apple", "type": "Company"},
{"id": "e3", "name": "Microsoft", "type": "Company"},
]
detector = DuplicateDetector(similarity_threshold=0.75, confidence_threshold=0.7)
candidates = detector.detect_duplicates(entities)
for c in candidates:
print(f"Duplicate: {c.entity1['id']} ↔ {c.entity2['id']} (score={c.similarity_score:.2f})")
Custom Weight Configuration
from semantica.deduplication import DuplicateDetector
detector = DuplicateDetector(
similarity_threshold=0.6,
confidence_threshold=0.5,
config={"similarity": {
"embedding_weight": 0.5,
"string_weight": 0.2,
"property_weight": 0.2,
"relationship_weight": 0.1
}},
)
candidates = detector.detect_duplicates(entities)
High-Level API Usage
from semantica.deduplication.methods import detect_duplicates, calculate_similarity
# Direct pairwise similarity
entity_a = {"name": "Google LLC", "type": "Company"}
entity_b = {"name": "Google", "type": "Company"}
sim_res = calculate_similarity(entity_a, entity_b, method="jaro_winkler")
print(f"Jaro-Winkler similarity: {sim_res.score:.3f}")
# Bulk detection
duplicates = detect_duplicates(entities, method="pairwise", similarity_threshold=0.8)
Merging Detected Duplicates
from semantica.deduplication import EntityMerger, MergeStrategyManager
merger = EntityMerger(strategy=MergeStrategyManager.get("keep_most_complete"))
merged = merger.merge_duplicates(candidates)
print(f"Merged {len(candidates)} pairs into {len(merged)} canonical entities.")
End-to-End Pipeline Integration
from semantica.kg import GraphBuilder
from semantica.deduplication import DuplicateDetector, EntityMerger
graph_builder = GraphBuilder(merge_entities=True)
graph = graph_builder.build(entities)
Summary
- Multi-factor scoring: Semantica combines string, property, relationship, and embedding similarities with configurable weights in
similarity_calculator.py. - Optimized pipeline: The
_prefilter_pairmechanism and union-find clustering minimize computational overhead when processing large entity sets. - Flexible thresholds: Control precision via
similarity_threshold,confidence_threshold,min_similarity, and result limits (top_k_per_entity,max_results). - Simple API: High-level functions in
methods.pyprovide one-line access to complex deduplication workflows, while the registry system supports custom strategies. - Graph integration: Setting
merge_entities=TrueinGraphBuilderautomates deduplication during knowledge graph construction.
Frequently Asked Questions
How does Semantica handle transitive duplicate relationships?
When use_clustering=True, the DuplicateDetector employs a union-find algorithm to identify transitive relationships where entity A matches B and B matches C, automatically grouping all three into a single DuplicateGroup. This prevents fragmented deduplication chains and ensures consistent canonical entity selection.
What similarity metrics does the SimilarityCalculator support?
The calculator supports Levenshtein distance, Jaro-Winkler similarity, and cosine similarity for string comparisons, Jaccard index for relationship sets, and cosine similarity for vector embeddings. These metrics operate independently and combine via weighted aggregation to produce the final similarity score.
How can I customize which factors influence the similarity score most?
Pass custom weights via the config parameter when initializing DuplicateDetector. Set embedding_weight, string_weight, property_weight, and relationship_weight to values between 0.0 and 1.0; the system automatically normalizes these if they do not sum to exactly 1.0. Higher weights increase that component's influence on the final match decision.
What is the difference between similarity_threshold and confidence_threshold?
The similarity_threshold filters based on the raw aggregated similarity score (0.0-1.0), while confidence_threshold applies additional heuristics or model-based confidence estimates that may incorporate data quality signals. Both must be satisfied for a pair to be considered duplicate, allowing fine-grained control over precision and recall trade-offs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →