How Hyperresearch Detects Duplicate Content: Cross-Provider and Note-Level Deduplication
Hyperresearch implements two complementary duplicate content detection pipelines—metadata deduplication for scholarly papers across multiple providers and shingle-based Jaccard similarity with MinHash LSH for user-generated notes—to ensure clean, canonical datasets without redundant entries.
Hyperresearch is an open-source scholarly research platform that aggregates bibliographic data from sources like OpenAlex, Crossref, and CORE while managing research notes in local vaults. To prevent data redundancy across these heterogeneous sources, the project implements sophisticated duplicate content detection mechanisms in src/hyperresearch/scholar/dedup.py for academic papers and src/hyperresearch/cli/dedup.py for note collections. These systems normalize identifiers, cluster similar records, and merge fields deterministically to produce canonical outputs.
Cross-Provider Paper Deduplication
When aggregating scholarly works from multiple APIs, Hyperresearch must resolve the same paper appearing with slight variations across providers. The deduplication engine in src/hyperresearch/scholar/dedup.py implements a hierarchical clustering strategy that prioritizes DOI matching while falling back to title-based heuristics for incomplete records.
DOI Normalization and Strict Compatibility
The system normalizes each incoming Paper using normalize_doi to handle case and URI variations consistently. Within the _Cluster class, the doi_compatible method (lines 53-60) enforces strict rules: papers sharing identical DOIs always merge, while disparate DOIs prevent clustering regardless of title similarity. This eliminates the risk of conflating distinct works that happen to share similar titles.
Title-Key Matching with Temporal Tolerance
For records lacking DOIs—common in preprint servers or older publications—Hyperresearch generates a title_key fingerprint and evaluates temporal proximity through _Cluster.year_compatible (lines 61-68). This method permits a one-year publication window to accommodate online-first versus print-year discrepancies, ensuring preprints correctly cluster with their final published versions despite minor date differences.
Canonical Field Selection and Provider Precedence
Once clustered, the _merge_cluster function (lines 96-140) constructs a canonical record by selecting optimal field values: the longest available abstract, highest citation count, and earliest publication date. The precedence configuration list determines provider priority (e.g., ["openalex", "crossref", "core"]), while first-seen ordering within ties guarantees deterministic, reproducible output across executions (lines 84-86).
The following example demonstrates merging papers from multiple providers:
# Example: merging papers from multiple providers
from hyperresearch.scholar.dedup import merge_papers
from hyperresearch.scholar.base import Paper
papers = [
Paper(source="openalex", title="Deep Learning", doi="10.1000/xyz123", year=2022, ...),
Paper(source="crossref", title="Deep Learning", doi="10.1000/xyz123", year=2022, ...),
Paper(source="core", title="Deep Learning", doi=None, year=2022, ...),
]
# Provider precedence: OpenAlex > Crossref > CORE
merged = merge_papers(papers, precedence=["openalex", "crossref", "core"])
print(merged[0].doi) # → 10.1000/xyz123
print(merged[0].source) # → openalex
Note-Level Content Deduplication
For user-generated research notes stored in vaults, Hyperresearch detects near-duplicate content using configurable text shingling and similarity algorithms. The CLI command hyperresearch dedup in src/hyperresearch/cli/dedup.py implements a hybrid approach that balances accuracy and performance based on collection size, excluding index notes and entries under 20 words.
Shingling and Jaccard Similarity
All qualifying notes are processed using shingling utilities from src/hyperresearch/core/similarity.py. The system generates word-level k-grams according to the shingle_size parameter, then calculates exact Jaccard similarity between documents to quantify content overlap. The jaccard function verifies candidate pairs against the configured threshold to confirm duplicates (lines 24-30).
MinHash LSH for Large Vaults
When the vault size exceeds lsh_switchover, Hyperresearch switches to probabilistic detection using minhash_signature and lsh_candidates (lines 14-22). Locality Sensitive Hashing reduces the search space from O(n²) to near-linear complexity by indexing similar MinHash signatures, enabling efficient deduplication of large note collections without exhaustive pairwise comparisons.
Brute-Force Fallback for Small Collections
For smaller vaults below the threshold, the system executes exact O(n²) Jaccard comparisons across all pairs (lines 2-12). This avoids the approximation errors inherent in probabilistic methods while maintaining acceptable performance for limited datasets. Results can be output as formatted text or JSON when using the --json flag (lines 78-88).
The CLI interface supports both interactive and programmatic usage:
# Example: finding near‑duplicate notes in a vault
$ hyperresearch dedup --threshold 0.85 --limit 10
# Output (example)
Similar notes (>85%, minhash+lsh):
92% | 42 (120w) <-> 57 (118w)
87% | 13 (95w) <-> 19 (97w)
Summary
- Cross-provider deduplication in
src/hyperresearch/scholar/dedup.pyclusters papers by normalized DOI or title-key with year tolerance, then merges fields using configurable provider precedence. - Note-level detection in
src/hyperresearch/cli/dedup.pyemploys shingle-based Jaccard similarity, automatically switching between brute-force O(n²) comparisons and MinHash LSH based on vault size. - Both systems use utilities from
src/hyperresearch/core/similarity.pyfor text processing and similarity calculations. - Provider precedence and similarity thresholds are user-configurable, ensuring deterministic output tailored to specific research workflows.
Frequently Asked Questions
How does Hyperresearch handle papers with missing DOIs?
When DOIs are absent, Hyperresearch falls back to title_key fingerprinting combined with _Cluster.year_compatible checking in src/hyperresearch/scholar/dedup.py. This applies a one-year publication window to account for online-first versus print date variations, allowing preprints to cluster correctly with their final versions despite missing identifiers.
What similarity threshold should I use for note deduplication?
The default threshold is configurable via the --threshold CLI argument, with 0.85 representing a conservative balance that catches near-duplicates while avoiding false positives. Higher values (0.90+) detect only nearly identical content, while lower values (0.70-0.80) capture more paraphrased or templated notes at the risk of increased false matches.
Why does Hyperresearch use both MinHash LSH and brute-force approaches?
The system dynamically selects algorithms based on the lsh_switchover parameter to optimize the performance-accuracy tradeoff. Brute-force Jaccard calculations provide exact similarity scores for small vaults where O(n²) complexity remains feasible, while MinHash LSH reduces computational overhead to near-linear time for large collections, making vault-wide deduplication practical with minimal accuracy loss.
Can provider precedence be customized when merging papers?
Yes, the merge_papers function accepts a precedence list parameter that determines which provider's metadata takes priority when fields conflict. For example, setting precedence=["openalex", "crossref", "core"] ensures that OpenAlex abstracts and citations are preferred over Crossref or CORE equivalents when constructing the canonical record.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →