# MemPalace Deduplication Strategies for Mined Content: A Technical Guide

> Learn MemPalace deduplication strategies for mined content. Discover how MemPalace uses a greedy longest match algorithm to remove redundant text and preserve complete documents.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: tutorial
- Published: 2026-06-07

---

**MemPalace eliminates near-duplicate text fragments using a greedy \"keep the longest\" algorithm that compares cosine distances against a configurable threshold, removing redundant drawers while preserving the most complete version of each mined document.**

MemPalace stores every mined text fragment as a *drawer* in a vector-store collection (`mempalace_drawers`). When the same source file is mined repeatedly—whether through project-wide `mempalace mine` commands or after file edits—many drawers become near-duplicates that bloat the palace. According to the MemPalace source code, the deduplication pipeline removes these redundant entries using source-file grouping and similarity matching to maintain an efficient, non-redundant knowledge base.

## Grouping by Source File Metadata

The deduplication process first groups drawers that originated from the **same `source_file`** metadata field. To optimize performance on large palaces, only groups containing at least `MIN_DRAWERS_TO_CHECK` (default **5**) are inspected, filtering out sparse sources that are unlikely to contain significant duplication.

In [`mempalace/dedup.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dedup.py), the `get_source_groups` function handles this aggregation by iterating through the collection and collecting drawer IDs per source file:

```python

# mempalace/dedup.py – grouping logic (lines 53-78)

def get_source_groups(col, min_count=MIN_DRAWERS_TO_CHECK, source_pattern=None, wing=None):
    …
    groups[src].append(did)               # ↳ collect drawer IDs per source file

    …
    return {src: ids for src, ids in groups.items() if len(ids) >= min_count}

```

This grouping ensures that deduplication comparisons only occur between fragments originating from identical source documents, preventing false positives between different files.

## The Greedy "Keep the Longest" Algorithm

For each source-file group, MemPalace applies a greedy deduplication strategy that prioritizes **document completeness**. Drawers are sorted by length (longest first) and processed sequentially against a running set of kept items.

The algorithm in `dedup_source_group` queries the storage backend's similarity index using cosine distance:

- **First drawer**: Always preserved as the anchor for that content region.
- **Subsequent drawers**: Compared against all kept drawers using the vector store's similarity search.
- **Duplicate detection**: If the cosine distance to any kept drawer is **below the configured `threshold`** (default **0.15**, corresponding to ~85% cosine similarity), the drawer is marked for deletion.

```python

# mempalace/dedup.py – core dedup loop (lines 81-122)

def dedup_source_group(col, drawer_ids, threshold=DEFAULT_THRESHOLD, dry_run=True):
    …
    items.sort(key=lambda x: len(x[1] or ""), reverse=True)   # longest first

    …
    results = col.query(query_texts=[doc], n_results=min(len(kept), 5),
                        include=["distances"])
    is_dup = any(dist < threshold for rid, dist in zip(results["ids"][0], distances))
    …

```

This approach guarantees that the richest (longest) version of any semantic cluster survives while shorter, redundant fragments are eliminated.

## Configuring Similarity Thresholds

The **threshold parameter** controls deduplication aggressiveness by setting the maximum cosine distance considered "duplicate":

- **Strict mode** (`--threshold 0.10`): Catches only highly similar, near-identical fragments. Use this when preserving slight variations matters.
- **Default balance** (`--threshold 0.15`): Removes obvious duplicates while distinguishing meaningfully different passages (~85% similarity cutoff).
- **Loose mode** (`--threshold 0.30-0.35`): Also removes chunks that have been rephrased but convey identical information, useful for cleaning up repetitive technical documentation.

Because the threshold represents **cosine distance** (lower values indicate higher similarity), decreasing the number increases strictness.

## CLI Usage for Deduplication

The `mempalace dedup` command provides both safe preview modes and destructive cleanup operations:

```bash

# Show statistics overview without making changes

mempalace dedup --stats

# Preview which drawers would be removed (safe mode)

mempalace dedup --dry-run

# Execute live deduplication with default settings

mempalace dedup

# Aggressive deduplication for paraphrased content

mempalace dedup --threshold 0.35

# Limit scope to a specific wing (project)

mempalace dedup --wing blog

# Target specific source files using pattern matching

mempalace dedup --source "README.md"

```

The `--dry-run` flag reports statistics and identifies which drawers *would* be deleted without calling `col.delete`. In live mode, deletions occur in batches of **500** IDs to optimize vector-store performance.

## Programmatic API Access

You can invoke the deduplication pipeline directly from Python using functions defined in [`mempalace/dedup.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dedup.py):

```python
from mempalace.dedup import dedup_palace, show_stats

# Run a dry-run against the default palace location

dedup_palace(dry_run=True)

# Live deduplication with custom threshold and wing restriction

dedup_palace(
    threshold=0.30,
    dry_run=False,
    wing="my_project",
    source_pattern="docs/",
)

# Inspect palace statistics before cleaning

show_stats()

```

 The `dedup_palace` function coordinates the full workflow: retrieving the collection from [`mempalace/palace.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/palace.py), grouping by source metadata, executing the greedy algorithm, and batch-deleting duplicates through the storage backend interface.

## Summary

- **Source-based grouping**: Drawers are clustered by `source_file` metadata, with a minimum group size of 5 to optimize performance.
- **Longest-first greedy selection**: The algorithm keeps the longest drawer in each similarity cluster, ensuring maximum information retention.
- **Cosine distance threshold**: Default 0.15 (~85% similarity) balances precision and recall; tune between 0.10 (strict) and 0.35 (loose) based on content redundancy.
- **Safe execution**: Use `--dry-run` to preview deletions; live mode deletes in 500-item batches via the ChromaDB or compatible backend.
- **Scope control**: Filter operations by **wing** (project namespace) or **source pattern** to limit processing to specific content subsets.

## Frequently Asked Questions

### How does MemPalace determine which duplicate to keep?

MemPalace uses a **greedy longest-first algorithm** implemented in [`mempalace/dedup.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dedup.py). For each group of drawers from the same source file, it sorts documents by character length (descending) and keeps the longest fragment as the canonical version. Subsequent drawers are compared against this kept set using cosine similarity; if they fall below the distance threshold, they are discarded in favor of the longer original.

### What similarity threshold should I use for technical documentation?

For technical documentation with frequent minor edits, use the **default threshold of 0.15**, which corresponds to approximately 85% cosine similarity. This catches verbatim duplicates and trivial changes without removing meaningfully rewritten passages. If your documentation contains heavy copy-paste repetition or template boilerplate, consider `--threshold 0.20` or `0.25` to capture near-duplicates while preserving distinct explanations.

### Is it safe to run deduplication on a production palace?

Yes, when using the **`--dry-run` flag** first. This mode queries the vector store and reports which drawers would be deleted without executing `col.delete` operations. Review the dry-run output to verify that the threshold is not overly aggressive for your content. Once satisfied, run without `--dry-run` to execute the deletion; the operation only affects drawers sharing `source_file` metadata, leaving unrelated content untouched.

### Which vector-store backends support MemPalace deduplication?

The deduplication module relies on the abstract storage interface defined in [`mempalace/backends/base.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/base.py), which requires `query()` support for similarity searches with distance metrics. The default **ChromaDB** backend ([`mempalace/backends/chroma.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/chroma.py)) implements this using cosine-distance queries. Any backend implementing the base class's query interface with distance returns can support the deduplication pipeline.