# How the MemPalace dedup Module Identifies and Handles Duplicate Content

> Learn how the MemPalace dedup module efficiently identifies and handles duplicate content using cosine distance similarity. Preserves the most complete versions.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: deep-dive
- Published: 2026-06-06

---

**The dedup module groups text chunks by source file, sorts them by length, and uses cosine distance similarity scoring to identify and remove near-identical drawers while preserving the longest, most complete version.**

The dedup module in the MemPalace open-source repository provides chunk-precise deduplication for verbatim text drawers. It operates entirely on-device using the configured storage backend to scan source files, detect duplicate content using vector similarity, and safely remove redundant drawers while keeping the richest version of each text segment.

## The Three-Phase Deduplication Pipeline

### Grouping Drawers by Source File

The process begins in [`mempalace/dedup.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dedup.py) with the `get_source_groups` function. This function scans the `mempalace_drawers` collection and organizes drawers into a dictionary keyed by `source_file` values.

Only source groups containing at least **MIN_DRAWERS_TO_CHECK** (default 5) are processed, filtering out trivially small sources. The function also supports optional `wing` and `source_pattern` parameters, allowing targeted deduplication across specific project wings or filename patterns.

### Greedy Duplicate Detection

Once grouped, the `dedup_source_group` function handles the actual similarity analysis. The algorithm follows a longest-first strategy: drawers are sorted by text length in descending order so that the most complete version is always preserved as the canonical copy.

For each candidate drawer, the module queries the backend similarity index using `col.query` to compare the drawer's text against already-kept items. The system uses **cosine distance** metrics, flagging duplicates when the distance falls below the configured `threshold` (default 0.15, representing approximately 85% similarity). Candidates exceeding this threshold are marked for deletion, while unique items join the kept set.

Empty or trivially short drawers (fewer than 20 characters) are automatically purged regardless of similarity scores, as they provide minimal value to the palace.

### Batch Deletion and Dry-Run Safety

When `dry_run=False`, the dedup module executes deletions in batches of 500 records using `col.delete` to minimize backend overhead. The `dedup_source_group` function returns two distinct lists: surviving drawer IDs and purged duplicate IDs.

The top-level `dedup_palace` orchestrator coordinates the entire workflow by obtaining the collection, invoking `get_source_groups`, and iterating through each source to aggregate results and print progress reports.

## Running Deduplication via the Python API

Direct API access allows fine-grained control over the deduplication process. Import the core functions from `mempalace.dedup` to run targeted operations:

```python

# Dry-run: Preview duplicates without deleting anything

from mempalace.dedup import dedup_palace

dedup_palace(dry_run=True)

```

For stricter deduplication targeting only near-identical matches, adjust the threshold and restrict processing to specific wings:

```python

# Apply strict deduplication to a specific wing

dedup_palace(
    wing="my_project",
    threshold=0.10,
    dry_run=False,
)

```

To inspect duplication metrics without modifying data, use the `show_stats` function:

```python
from mempalace.dedup import show_stats

show_stats()

```

## Configuration Parameters and Tuning

Several configuration constants govern deduplication behavior in [`mempalace/dedup.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dedup.py):

* **MIN_DRAWERS_TO_CHECK**: Minimum group size (default 5) required before similarity analysis begins.
* **threshold**: Cosine distance limit (default 0.15) determining duplicate classification; lower values require higher similarity.
* **Batch size**: Deletions execute in groups of 500 drawers to optimize storage backend performance.

The threshold parameter offers direct control over deduplication sensitivity. A value of 0.10 catches only nearly identical content, while 0.20 identifies loosely paraphrased duplicates.

## Summary

* The dedup module groups drawers by `source_file` using `get_source_groups`, filtering out groups smaller than 5 items.
* The `dedup_source_group` function sorts candidates longest-first and queries the vector similarity index to calculate cosine distances.
* Drawers with distances below the 0.15 threshold (approximately 85% similarity) are flagged as duplicates and scheduled for removal.
* The system preserves the longest version of each text chunk and automatically deletes trivially short content under 20 characters.
* Deletions occur in 500-record batches, with full support for dry-run mode via `dedup_palace(dry_run=True)`.

## Frequently Asked Questions

### What similarity algorithm does the dedup module use?

The module relies on **cosine distance** scoring provided by the backend similarity index. When `col.query` compares drawer texts, it returns distance values where lower numbers indicate higher similarity. The default threshold of 0.15 corresponds to roughly 85% textual similarity.

### Why does the dedup module sort drawers by length before checking for duplicates?

The algorithm sorts drawers longest-first to ensure the **most complete version** of any text segment becomes the canonical copy. When cosine distance identifies near-matches, the system keeps the already-processed (longer) drawer and marks the shorter candidate as a duplicate, preserving the richest content.

### How can I test deduplication without risking data loss?

Pass `dry_run=True` to the `dedup_palace` function. This mode executes all grouping and similarity calculations but skips the final `col.delete` operations, returning only the lists of what would be removed versus preserved.

### What happens to very short text chunks during deduplication?

Drawers containing fewer than 20 characters are automatically deleted during the `dedup_source_group` phase regardless of similarity scores. This prevents the accumulation of trivial fragments that provide minimal semantic value to the palace collection.