# How to Use Semantic Similarity for Entity Deduplication in Semantica: A Complete Developer Guide

> Learn how to use semantic similarity for entity deduplication in Semantica. This developer guide shows you how to combine metrics for accurate duplicate resolution.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-10

---

**Semantica resolves duplicate entities by computing multi-factor semantic similarity scores—combining string, property, relationship, and embedding metrics—then grouping or merging pairs that exceed configurable thresholds.**

Entity deduplication is critical for maintaining clean knowledge graphs, and the Semantica framework provides a robust pipeline for identifying and merging duplicate entities using semantic similarity. The system evaluates entities across multiple dimensions to generate confidence-weighted similarity scores, enabling both pairwise matching and transitive clustering. This guide walks through the core components, workflow, and implementation patterns found in the `semantica-agi/semantica` repository.

## Core Deduplication Components

The deduplication architecture consists of three primary classes that handle calculation, detection orchestration, and API convenience.

### SimilarityCalculator

Located in [`semantica/deduplication/similarity_calculator.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/deduplication/similarity_calculator.py), the `SimilarityCalculator` class computes normalized similarity scores across four distinct factors. It implements a fast pre-filter via `_prefilter_pair` to reject obvious non-matches before expensive calculations occur. The calculator evaluates **string similarity** (Levenshtein, Jaro-Winkler, cosine), **property similarity** (weighted value comparison), **relationship similarity** (Jaccard index of relationship sets), and **embedding similarity** (vector cosine similarity). Individual component scores are aggregated using configurable weights—`embedding_weight`, `string_weight`, `property_weight`, and `relationship_weight`—which are automatically normalized if they do not sum to 1.0 (lines 44-50).

### DuplicateDetector

The `DuplicateDetector` class in [`semantica/deduplication/duplicate_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/deduplication/duplicate_detector.py) wraps the calculator and manages the detection pipeline. It applies configurable `similarity_threshold` and `confidence_threshold` values, supports optional `min_similarity` floors, and implements ranking controls via `max_results` and `top_k_per_entity`. The `sort_by` parameter allows sorting by either `confidence` or `similarity_score`. When `use_clustering=True`, the detector employs a union-find algorithm to group transitive duplicates into `DuplicateGroup` objects, automatically selecting the most complete entity as the representative.

### High-Level API Methods

For simplified access, [`semantica/deduplication/methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/deduplication/methods.py) exposes convenience functions including `detect_duplicates()`, `calculate_similarity()`, `build_clusters()`, and `merge_entities()`. These wrappers interface with the detector and calculator while supporting custom strategies through the method registry defined in [`semantica/deduplication/registry.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/deduplication/registry.py).

## The Semantic Similarity Workflow

The deduplication process follows a seven-stage pipeline that transforms raw entities into canonical, merged records.

### 1. Pre-processing and Normalization

Before similarity calculation, entity names undergo normalization via `_lower_name` and optional `EntityNormalizer` components. This canonicalization of punctuation and case improves accuracy during the string similarity stage.

### 2. Fast Pre-filtering

The `_prefilter_pair` method quickly rejects candidate pairs that differ in entity type or exhibit extreme name-length ratio disparities. This optimization prevents wasted computation on obvious non-matches.

### 3. Multi-Factor Score Calculation

For each surviving pair, the calculator evaluates the four component scores: string distance metrics, property value overlaps, relationship set Jaccard similarity, and embedding vector cosine similarity.

### 4. Weighted Aggregation

Component scores combine according to user-defined weights. The calculator automatically renormalizes weights that do not sum to unity, ensuring consistent scoring behavior.

### 5. Thresholding and Ranking

The `detect_duplicates()` method filters candidates against `similarity_threshold` and `confidence_threshold`, then applies result limits. Users can constrain output to `top_k_per_entity` matches or an absolute `max_results` ceiling.

### 6. Transitive Group Formation

When clustering is enabled, the union-find algorithm identifies transitive duplicate relationships (A matches B, B matches C, therefore A matches C) and collapses them into `DuplicateGroup` instances.

### 7. Entity Merging

Detected groups feed into `EntityMerger` (accessible via `merge_entities()`), which applies strategies like `keep_most_complete` to consolidate records while preserving provenance metadata.

## Implementation Examples

### Basic Duplicate Detection

```python
from semantica.deduplication import DuplicateDetector

entities = [
    {"id": "e1", "name": "Apple Inc.", "type": "Company"},
    {"id": "e2", "name": "Apple", "type": "Company"},
    {"id": "e3", "name": "Microsoft", "type": "Company"},
]

detector = DuplicateDetector(similarity_threshold=0.75, confidence_threshold=0.7)
candidates = detector.detect_duplicates(entities)

for c in candidates:
    print(f"Duplicate: {c.entity1['id']} ↔ {c.entity2['id']} (score={c.similarity_score:.2f})")

```

### Custom Weight Configuration

```python
from semantica.deduplication import DuplicateDetector

detector = DuplicateDetector(
    similarity_threshold=0.6,
    confidence_threshold=0.5,
    config={"similarity": {
        "embedding_weight": 0.5, 
        "string_weight": 0.2, 
        "property_weight": 0.2, 
        "relationship_weight": 0.1
    }},
)

candidates = detector.detect_duplicates(entities)

```

### High-Level API Usage

```python
from semantica.deduplication.methods import detect_duplicates, calculate_similarity

# Direct pairwise similarity

entity_a = {"name": "Google LLC", "type": "Company"}
entity_b = {"name": "Google", "type": "Company"}
sim_res = calculate_similarity(entity_a, entity_b, method="jaro_winkler")
print(f"Jaro-Winkler similarity: {sim_res.score:.3f}")

# Bulk detection

duplicates = detect_duplicates(entities, method="pairwise", similarity_threshold=0.8)

```

### Merging Detected Duplicates

```python
from semantica.deduplication import EntityMerger, MergeStrategyManager

merger = EntityMerger(strategy=MergeStrategyManager.get("keep_most_complete"))
merged = merger.merge_duplicates(candidates)

print(f"Merged {len(candidates)} pairs into {len(merged)} canonical entities.")

```

### End-to-End Pipeline Integration

```python
from semantica.kg import GraphBuilder
from semantica.deduplication import DuplicateDetector, EntityMerger

graph_builder = GraphBuilder(merge_entities=True)
graph = graph_builder.build(entities)

```

## Summary

- **Multi-factor scoring**: Semantica combines string, property, relationship, and embedding similarities with configurable weights in [`similarity_calculator.py`](https://github.com/semantica-agi/semantica/blob/main/similarity_calculator.py).
- **Optimized pipeline**: The `_prefilter_pair` mechanism and union-find clustering minimize computational overhead when processing large entity sets.
- **Flexible thresholds**: Control precision via `similarity_threshold`, `confidence_threshold`, `min_similarity`, and result limits (`top_k_per_entity`, `max_results`).
- **Simple API**: High-level functions in [`methods.py`](https://github.com/semantica-agi/semantica/blob/main/methods.py) provide one-line access to complex deduplication workflows, while the registry system supports custom strategies.
- **Graph integration**: Setting `merge_entities=True` in `GraphBuilder` automates deduplication during knowledge graph construction.

## Frequently Asked Questions

### How does Semantica handle transitive duplicate relationships?

When `use_clustering=True`, the `DuplicateDetector` employs a union-find algorithm to identify transitive relationships where entity A matches B and B matches C, automatically grouping all three into a single `DuplicateGroup`. This prevents fragmented deduplication chains and ensures consistent canonical entity selection.

### What similarity metrics does the SimilarityCalculator support?

The calculator supports **Levenshtein distance**, **Jaro-Winkler similarity**, and **cosine similarity** for string comparisons, **Jaccard index** for relationship sets, and **cosine similarity** for vector embeddings. These metrics operate independently and combine via weighted aggregation to produce the final similarity score.

### How can I customize which factors influence the similarity score most?

Pass custom weights via the `config` parameter when initializing `DuplicateDetector`. Set `embedding_weight`, `string_weight`, `property_weight`, and `relationship_weight` to values between 0.0 and 1.0; the system automatically normalizes these if they do not sum to exactly 1.0. Higher weights increase that component's influence on the final match decision.

### What is the difference between similarity_threshold and confidence_threshold?

The `similarity_threshold` filters based on the raw aggregated similarity score (0.0-1.0), while `confidence_threshold` applies additional heuristics or model-based confidence estimates that may incorporate data quality signals. Both must be satisfied for a pair to be considered duplicate, allowing fine-grained control over precision and recall trade-offs.