How the MemPalace dedup Module Identifies and Handles Duplicate Content
The dedup module groups text chunks by source file, sorts them by length, and uses cosine distance similarity scoring to identify and remove near-identical drawers while preserving the longest, most complete version.
The dedup module in the MemPalace open-source repository provides chunk-precise deduplication for verbatim text drawers. It operates entirely on-device using the configured storage backend to scan source files, detect duplicate content using vector similarity, and safely remove redundant drawers while keeping the richest version of each text segment.
The Three-Phase Deduplication Pipeline
Grouping Drawers by Source File
The process begins in mempalace/dedup.py with the get_source_groups function. This function scans the mempalace_drawers collection and organizes drawers into a dictionary keyed by source_file values.
Only source groups containing at least MIN_DRAWERS_TO_CHECK (default 5) are processed, filtering out trivially small sources. The function also supports optional wing and source_pattern parameters, allowing targeted deduplication across specific project wings or filename patterns.
Greedy Duplicate Detection
Once grouped, the dedup_source_group function handles the actual similarity analysis. The algorithm follows a longest-first strategy: drawers are sorted by text length in descending order so that the most complete version is always preserved as the canonical copy.
For each candidate drawer, the module queries the backend similarity index using col.query to compare the drawer's text against already-kept items. The system uses cosine distance metrics, flagging duplicates when the distance falls below the configured threshold (default 0.15, representing approximately 85% similarity). Candidates exceeding this threshold are marked for deletion, while unique items join the kept set.
Empty or trivially short drawers (fewer than 20 characters) are automatically purged regardless of similarity scores, as they provide minimal value to the palace.
Batch Deletion and Dry-Run Safety
When dry_run=False, the dedup module executes deletions in batches of 500 records using col.delete to minimize backend overhead. The dedup_source_group function returns two distinct lists: surviving drawer IDs and purged duplicate IDs.
The top-level dedup_palace orchestrator coordinates the entire workflow by obtaining the collection, invoking get_source_groups, and iterating through each source to aggregate results and print progress reports.
Running Deduplication via the Python API
Direct API access allows fine-grained control over the deduplication process. Import the core functions from mempalace.dedup to run targeted operations:
# Dry-run: Preview duplicates without deleting anything
from mempalace.dedup import dedup_palace
dedup_palace(dry_run=True)
For stricter deduplication targeting only near-identical matches, adjust the threshold and restrict processing to specific wings:
# Apply strict deduplication to a specific wing
dedup_palace(
wing="my_project",
threshold=0.10,
dry_run=False,
)
To inspect duplication metrics without modifying data, use the show_stats function:
from mempalace.dedup import show_stats
show_stats()
Configuration Parameters and Tuning
Several configuration constants govern deduplication behavior in mempalace/dedup.py:
- MIN_DRAWERS_TO_CHECK: Minimum group size (default 5) required before similarity analysis begins.
- threshold: Cosine distance limit (default 0.15) determining duplicate classification; lower values require higher similarity.
- Batch size: Deletions execute in groups of 500 drawers to optimize storage backend performance.
The threshold parameter offers direct control over deduplication sensitivity. A value of 0.10 catches only nearly identical content, while 0.20 identifies loosely paraphrased duplicates.
Summary
- The dedup module groups drawers by
source_fileusingget_source_groups, filtering out groups smaller than 5 items. - The
dedup_source_groupfunction sorts candidates longest-first and queries the vector similarity index to calculate cosine distances. - Drawers with distances below the 0.15 threshold (approximately 85% similarity) are flagged as duplicates and scheduled for removal.
- The system preserves the longest version of each text chunk and automatically deletes trivially short content under 20 characters.
- Deletions occur in 500-record batches, with full support for dry-run mode via
dedup_palace(dry_run=True).
Frequently Asked Questions
What similarity algorithm does the dedup module use?
The module relies on cosine distance scoring provided by the backend similarity index. When col.query compares drawer texts, it returns distance values where lower numbers indicate higher similarity. The default threshold of 0.15 corresponds to roughly 85% textual similarity.
Why does the dedup module sort drawers by length before checking for duplicates?
The algorithm sorts drawers longest-first to ensure the most complete version of any text segment becomes the canonical copy. When cosine distance identifies near-matches, the system keeps the already-processed (longer) drawer and marks the shorter candidate as a duplicate, preserving the richest content.
How can I test deduplication without risking data loss?
Pass dry_run=True to the dedup_palace function. This mode executes all grouping and similarity calculations but skips the final col.delete operations, returning only the lists of what would be removed versus preserved.
What happens to very short text chunks during deduplication?
Drawers containing fewer than 20 characters are automatically deleted during the dedup_source_group phase regardless of similarity scores. This prevents the accumulation of trivial fragments that provide minimal semantic value to the palace collection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →