How to Use Semantic Similarity for Entity Deduplication in Semantica: A Complete Developer Guide

Semantica resolves duplicate entities by computing multi-factor semantic similarity scores—combining string, property, relationship, and embedding metrics—then grouping or merging pairs that exceed configurable thresholds.

Entity deduplication is critical for maintaining clean knowledge graphs, and the Semantica framework provides a robust pipeline for identifying and merging duplicate entities using semantic similarity. The system evaluates entities across multiple dimensions to generate confidence-weighted similarity scores, enabling both pairwise matching and transitive clustering. This guide walks through the core components, workflow, and implementation patterns found in the semantica-agi/semantica repository.

Core Deduplication Components

The deduplication architecture consists of three primary classes that handle calculation, detection orchestration, and API convenience.

SimilarityCalculator

Located in semantica/deduplication/similarity_calculator.py, the SimilarityCalculator class computes normalized similarity scores across four distinct factors. It implements a fast pre-filter via _prefilter_pair to reject obvious non-matches before expensive calculations occur. The calculator evaluates string similarity (Levenshtein, Jaro-Winkler, cosine), property similarity (weighted value comparison), relationship similarity (Jaccard index of relationship sets), and embedding similarity (vector cosine similarity). Individual component scores are aggregated using configurable weights—embedding_weight, string_weight, property_weight, and relationship_weight—which are automatically normalized if they do not sum to 1.0 (lines 44-50).

DuplicateDetector

The DuplicateDetector class in semantica/deduplication/duplicate_detector.py wraps the calculator and manages the detection pipeline. It applies configurable similarity_threshold and confidence_threshold values, supports optional min_similarity floors, and implements ranking controls via max_results and top_k_per_entity. The sort_by parameter allows sorting by either confidence or similarity_score. When use_clustering=True, the detector employs a union-find algorithm to group transitive duplicates into DuplicateGroup objects, automatically selecting the most complete entity as the representative.

High-Level API Methods

For simplified access, semantica/deduplication/methods.py exposes convenience functions including detect_duplicates(), calculate_similarity(), build_clusters(), and merge_entities(). These wrappers interface with the detector and calculator while supporting custom strategies through the method registry defined in semantica/deduplication/registry.py.

The Semantic Similarity Workflow

The deduplication process follows a seven-stage pipeline that transforms raw entities into canonical, merged records.

1. Pre-processing and Normalization

Before similarity calculation, entity names undergo normalization via _lower_name and optional EntityNormalizer components. This canonicalization of punctuation and case improves accuracy during the string similarity stage.

2. Fast Pre-filtering

The _prefilter_pair method quickly rejects candidate pairs that differ in entity type or exhibit extreme name-length ratio disparities. This optimization prevents wasted computation on obvious non-matches.

3. Multi-Factor Score Calculation

For each surviving pair, the calculator evaluates the four component scores: string distance metrics, property value overlaps, relationship set Jaccard similarity, and embedding vector cosine similarity.

4. Weighted Aggregation

Component scores combine according to user-defined weights. The calculator automatically renormalizes weights that do not sum to unity, ensuring consistent scoring behavior.

5. Thresholding and Ranking

The detect_duplicates() method filters candidates against similarity_threshold and confidence_threshold, then applies result limits. Users can constrain output to top_k_per_entity matches or an absolute max_results ceiling.

6. Transitive Group Formation

When clustering is enabled, the union-find algorithm identifies transitive duplicate relationships (A matches B, B matches C, therefore A matches C) and collapses them into DuplicateGroup instances.

7. Entity Merging

Detected groups feed into EntityMerger (accessible via merge_entities()), which applies strategies like keep_most_complete to consolidate records while preserving provenance metadata.

Implementation Examples

Basic Duplicate Detection

from semantica.deduplication import DuplicateDetector

entities = [
    {"id": "e1", "name": "Apple Inc.", "type": "Company"},
    {"id": "e2", "name": "Apple", "type": "Company"},
    {"id": "e3", "name": "Microsoft", "type": "Company"},
]

detector = DuplicateDetector(similarity_threshold=0.75, confidence_threshold=0.7)
candidates = detector.detect_duplicates(entities)

for c in candidates:
    print(f"Duplicate: {c.entity1['id']} ↔ {c.entity2['id']} (score={c.similarity_score:.2f})")

Custom Weight Configuration

from semantica.deduplication import DuplicateDetector

detector = DuplicateDetector(
    similarity_threshold=0.6,
    confidence_threshold=0.5,
    config={"similarity": {
        "embedding_weight": 0.5, 
        "string_weight": 0.2, 
        "property_weight": 0.2, 
        "relationship_weight": 0.1
    }},
)

candidates = detector.detect_duplicates(entities)

High-Level API Usage

from semantica.deduplication.methods import detect_duplicates, calculate_similarity

# Direct pairwise similarity

entity_a = {"name": "Google LLC", "type": "Company"}
entity_b = {"name": "Google", "type": "Company"}
sim_res = calculate_similarity(entity_a, entity_b, method="jaro_winkler")
print(f"Jaro-Winkler similarity: {sim_res.score:.3f}")

# Bulk detection

duplicates = detect_duplicates(entities, method="pairwise", similarity_threshold=0.8)

Merging Detected Duplicates

from semantica.deduplication import EntityMerger, MergeStrategyManager

merger = EntityMerger(strategy=MergeStrategyManager.get("keep_most_complete"))
merged = merger.merge_duplicates(candidates)

print(f"Merged {len(candidates)} pairs into {len(merged)} canonical entities.")

End-to-End Pipeline Integration

from semantica.kg import GraphBuilder
from semantica.deduplication import DuplicateDetector, EntityMerger

graph_builder = GraphBuilder(merge_entities=True)
graph = graph_builder.build(entities)

Summary

  • Multi-factor scoring: Semantica combines string, property, relationship, and embedding similarities with configurable weights in similarity_calculator.py.
  • Optimized pipeline: The _prefilter_pair mechanism and union-find clustering minimize computational overhead when processing large entity sets.
  • Flexible thresholds: Control precision via similarity_threshold, confidence_threshold, min_similarity, and result limits (top_k_per_entity, max_results).
  • Simple API: High-level functions in methods.py provide one-line access to complex deduplication workflows, while the registry system supports custom strategies.
  • Graph integration: Setting merge_entities=True in GraphBuilder automates deduplication during knowledge graph construction.

Frequently Asked Questions

How does Semantica handle transitive duplicate relationships?

When use_clustering=True, the DuplicateDetector employs a union-find algorithm to identify transitive relationships where entity A matches B and B matches C, automatically grouping all three into a single DuplicateGroup. This prevents fragmented deduplication chains and ensures consistent canonical entity selection.

What similarity metrics does the SimilarityCalculator support?

The calculator supports Levenshtein distance, Jaro-Winkler similarity, and cosine similarity for string comparisons, Jaccard index for relationship sets, and cosine similarity for vector embeddings. These metrics operate independently and combine via weighted aggregation to produce the final similarity score.

How can I customize which factors influence the similarity score most?

Pass custom weights via the config parameter when initializing DuplicateDetector. Set embedding_weight, string_weight, property_weight, and relationship_weight to values between 0.0 and 1.0; the system automatically normalizes these if they do not sum to exactly 1.0. Higher weights increase that component's influence on the final match decision.

What is the difference between similarity_threshold and confidence_threshold?

The similarity_threshold filters based on the raw aggregated similarity score (0.0-1.0), while confidence_threshold applies additional heuristics or model-based confidence estimates that may incorporate data quality signals. Both must be satisfied for a pair to be considered duplicate, allowing fine-grained control over precision and recall trade-offs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →