MemPalace Deduplication Strategies for Mined Content: A Technical Guide
MemPalace eliminates near-duplicate text fragments using a greedy "keep the longest" algorithm that compares cosine distances against a configurable threshold, removing redundant drawers while preserving the most complete version of each mined document.
MemPalace stores every mined text fragment as a drawer in a vector-store collection (mempalace_drawers). When the same source file is mined repeatedly—whether through project-wide mempalace mine commands or after file edits—many drawers become near-duplicates that bloat the palace. According to the MemPalace source code, the deduplication pipeline removes these redundant entries using source-file grouping and similarity matching to maintain an efficient, non-redundant knowledge base.
Grouping by Source File Metadata
The deduplication process first groups drawers that originated from the same source_file metadata field. To optimize performance on large palaces, only groups containing at least MIN_DRAWERS_TO_CHECK (default 5) are inspected, filtering out sparse sources that are unlikely to contain significant duplication.
In mempalace/dedup.py, the get_source_groups function handles this aggregation by iterating through the collection and collecting drawer IDs per source file:
# mempalace/dedup.py – grouping logic (lines 53-78)
def get_source_groups(col, min_count=MIN_DRAWERS_TO_CHECK, source_pattern=None, wing=None):
…
groups[src].append(did) # ↳ collect drawer IDs per source file
…
return {src: ids for src, ids in groups.items() if len(ids) >= min_count}
This grouping ensures that deduplication comparisons only occur between fragments originating from identical source documents, preventing false positives between different files.
The Greedy "Keep the Longest" Algorithm
For each source-file group, MemPalace applies a greedy deduplication strategy that prioritizes document completeness. Drawers are sorted by length (longest first) and processed sequentially against a running set of kept items.
The algorithm in dedup_source_group queries the storage backend's similarity index using cosine distance:
- First drawer: Always preserved as the anchor for that content region.
- Subsequent drawers: Compared against all kept drawers using the vector store's similarity search.
- Duplicate detection: If the cosine distance to any kept drawer is below the configured
threshold(default 0.15, corresponding to ~85% cosine similarity), the drawer is marked for deletion.
# mempalace/dedup.py – core dedup loop (lines 81-122)
def dedup_source_group(col, drawer_ids, threshold=DEFAULT_THRESHOLD, dry_run=True):
…
items.sort(key=lambda x: len(x[1] or ""), reverse=True) # longest first
…
results = col.query(query_texts=[doc], n_results=min(len(kept), 5),
include=["distances"])
is_dup = any(dist < threshold for rid, dist in zip(results["ids"][0], distances))
…
This approach guarantees that the richest (longest) version of any semantic cluster survives while shorter, redundant fragments are eliminated.
Configuring Similarity Thresholds
The threshold parameter controls deduplication aggressiveness by setting the maximum cosine distance considered "duplicate":
- Strict mode (
--threshold 0.10): Catches only highly similar, near-identical fragments. Use this when preserving slight variations matters. - Default balance (
--threshold 0.15): Removes obvious duplicates while distinguishing meaningfully different passages (~85% similarity cutoff). - Loose mode (
--threshold 0.30-0.35): Also removes chunks that have been rephrased but convey identical information, useful for cleaning up repetitive technical documentation.
Because the threshold represents cosine distance (lower values indicate higher similarity), decreasing the number increases strictness.
CLI Usage for Deduplication
The mempalace dedup command provides both safe preview modes and destructive cleanup operations:
# Show statistics overview without making changes
mempalace dedup --stats
# Preview which drawers would be removed (safe mode)
mempalace dedup --dry-run
# Execute live deduplication with default settings
mempalace dedup
# Aggressive deduplication for paraphrased content
mempalace dedup --threshold 0.35
# Limit scope to a specific wing (project)
mempalace dedup --wing blog
# Target specific source files using pattern matching
mempalace dedup --source "README.md"
The --dry-run flag reports statistics and identifies which drawers would be deleted without calling col.delete. In live mode, deletions occur in batches of 500 IDs to optimize vector-store performance.
Programmatic API Access
You can invoke the deduplication pipeline directly from Python using functions defined in mempalace/dedup.py:
from mempalace.dedup import dedup_palace, show_stats
# Run a dry-run against the default palace location
dedup_palace(dry_run=True)
# Live deduplication with custom threshold and wing restriction
dedup_palace(
threshold=0.30,
dry_run=False,
wing="my_project",
source_pattern="docs/",
)
# Inspect palace statistics before cleaning
show_stats()
The dedup_palace function coordinates the full workflow: retrieving the collection from mempalace/palace.py, grouping by source metadata, executing the greedy algorithm, and batch-deleting duplicates through the storage backend interface.
Summary
- Source-based grouping: Drawers are clustered by
source_filemetadata, with a minimum group size of 5 to optimize performance. - Longest-first greedy selection: The algorithm keeps the longest drawer in each similarity cluster, ensuring maximum information retention.
- Cosine distance threshold: Default 0.15 (~85% similarity) balances precision and recall; tune between 0.10 (strict) and 0.35 (loose) based on content redundancy.
- Safe execution: Use
--dry-runto preview deletions; live mode deletes in 500-item batches via the ChromaDB or compatible backend. - Scope control: Filter operations by wing (project namespace) or source pattern to limit processing to specific content subsets.
Frequently Asked Questions
How does MemPalace determine which duplicate to keep?
MemPalace uses a greedy longest-first algorithm implemented in mempalace/dedup.py. For each group of drawers from the same source file, it sorts documents by character length (descending) and keeps the longest fragment as the canonical version. Subsequent drawers are compared against this kept set using cosine similarity; if they fall below the distance threshold, they are discarded in favor of the longer original.
What similarity threshold should I use for technical documentation?
For technical documentation with frequent minor edits, use the default threshold of 0.15, which corresponds to approximately 85% cosine similarity. This catches verbatim duplicates and trivial changes without removing meaningfully rewritten passages. If your documentation contains heavy copy-paste repetition or template boilerplate, consider --threshold 0.20 or 0.25 to capture near-duplicates while preserving distinct explanations.
Is it safe to run deduplication on a production palace?
Yes, when using the --dry-run flag first. This mode queries the vector store and reports which drawers would be deleted without executing col.delete operations. Review the dry-run output to verify that the threshold is not overly aggressive for your content. Once satisfied, run without --dry-run to execute the deletion; the operation only affects drawers sharing source_file metadata, leaving unrelated content untouched.
Which vector-store backends support MemPalace deduplication?
The deduplication module relies on the abstract storage interface defined in mempalace/backends/base.py, which requires query() support for similarity searches with distance metrics. The default ChromaDB backend (mempalace/backends/chroma.py) implements this using cosine-distance queries. Any backend implementing the base class's query interface with distance returns can support the deduplication pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →