# How Hyperresearch Detects Duplicate Content: Cross-Provider and Note-Level Deduplication

> Discover how Hyperresearch detects duplicate content using cross-provider metadata deduplication and note-level shingling. Ensure clean datasets with our advanced approach.

- Repository: [Jordan Gibbs/hyperresearch](https://github.com/jordan-gibbs/hyperresearch)
- Tags: how-to-guide
- Published: 2026-09-13

---

**Hyperresearch implements two complementary duplicate content detection pipelines—metadata deduplication for scholarly papers across multiple providers and shingle-based Jaccard similarity with MinHash LSH for user-generated notes—to ensure clean, canonical datasets without redundant entries.**

Hyperresearch is an open-source scholarly research platform that aggregates bibliographic data from sources like OpenAlex, Crossref, and CORE while managing research notes in local vaults. To prevent data redundancy across these heterogeneous sources, the project implements sophisticated **duplicate content detection** mechanisms in [`src/hyperresearch/scholar/dedup.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/scholar/dedup.py) for academic papers and [`src/hyperresearch/cli/dedup.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/dedup.py) for note collections. These systems normalize identifiers, cluster similar records, and merge fields deterministically to produce canonical outputs.

## Cross-Provider Paper Deduplication

When aggregating scholarly works from multiple APIs, Hyperresearch must resolve the same paper appearing with slight variations across providers. The deduplication engine in [`src/hyperresearch/scholar/dedup.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/scholar/dedup.py) implements a hierarchical clustering strategy that prioritizes DOI matching while falling back to title-based heuristics for incomplete records.

### DOI Normalization and Strict Compatibility

The system normalizes each incoming `Paper` using `normalize_doi` to handle case and URI variations consistently. Within the `_Cluster` class, the `doi_compatible` method (lines 53-60) enforces strict rules: papers sharing identical DOIs always merge, while disparate DOIs prevent clustering regardless of title similarity. This eliminates the risk of conflating distinct works that happen to share similar titles.

### Title-Key Matching with Temporal Tolerance

For records lacking DOIs—common in preprint servers or older publications—Hyperresearch generates a `title_key` fingerprint and evaluates temporal proximity through `_Cluster.year_compatible` (lines 61-68). This method permits a one-year publication window to accommodate online-first versus print-year discrepancies, ensuring preprints correctly cluster with their final published versions despite minor date differences.

### Canonical Field Selection and Provider Precedence

Once clustered, the `_merge_cluster` function (lines 96-140) constructs a canonical record by selecting optimal field values: the longest available abstract, highest citation count, and earliest publication date. The `precedence` configuration list determines provider priority (e.g., `["openalex", "crossref", "core"]`), while first-seen ordering within ties guarantees deterministic, reproducible output across executions (lines 84-86).

The following example demonstrates merging papers from multiple providers:

```python

# Example: merging papers from multiple providers

from hyperresearch.scholar.dedup import merge_papers
from hyperresearch.scholar.base import Paper

papers = [
    Paper(source="openalex", title="Deep Learning", doi="10.1000/xyz123", year=2022, ...),
    Paper(source="crossref", title="Deep Learning", doi="10.1000/xyz123", year=2022, ...),
    Paper(source="core", title="Deep Learning", doi=None, year=2022, ...),
]

# Provider precedence: OpenAlex > Crossref > CORE

merged = merge_papers(papers, precedence=["openalex", "crossref", "core"])
print(merged[0].doi)          # → 10.1000/xyz123

print(merged[0].source)       # → openalex

```

## Note-Level Content Deduplication

For user-generated research notes stored in vaults, Hyperresearch detects near-duplicate content using configurable text shingling and similarity algorithms. The CLI command `hyperresearch dedup` in [`src/hyperresearch/cli/dedup.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/dedup.py) implements a hybrid approach that balances accuracy and performance based on collection size, excluding index notes and entries under 20 words.

### Shingling and Jaccard Similarity

All qualifying notes are processed using shingling utilities from [`src/hyperresearch/core/similarity.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/similarity.py). The system generates word-level k-grams according to the `shingle_size` parameter, then calculates exact Jaccard similarity between documents to quantify content overlap. The `jaccard` function verifies candidate pairs against the configured `threshold` to confirm duplicates (lines 24-30).

### MinHash LSH for Large Vaults

When the vault size exceeds `lsh_switchover`, Hyperresearch switches to probabilistic detection using `minhash_signature` and `lsh_candidates` (lines 14-22). Locality Sensitive Hashing reduces the search space from O(n²) to near-linear complexity by indexing similar MinHash signatures, enabling efficient deduplication of large note collections without exhaustive pairwise comparisons.

### Brute-Force Fallback for Small Collections

For smaller vaults below the threshold, the system executes exact O(n²) Jaccard comparisons across all pairs (lines 2-12). This avoids the approximation errors inherent in probabilistic methods while maintaining acceptable performance for limited datasets. Results can be output as formatted text or JSON when using the `--json` flag (lines 78-88).

The CLI interface supports both interactive and programmatic usage:

```bash

# Example: finding near‑duplicate notes in a vault

$ hyperresearch dedup --threshold 0.85 --limit 10

# Output (example)

Similar notes (>85%, minhash+lsh):
  92% | 42 (120w) <-> 57 (118w)
  87% | 13 (95w)  <-> 19 (97w)

```

## Summary

- **Cross-provider deduplication** in [`src/hyperresearch/scholar/dedup.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/scholar/dedup.py) clusters papers by normalized DOI or title-key with year tolerance, then merges fields using configurable provider precedence.
- **Note-level detection** in [`src/hyperresearch/cli/dedup.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/dedup.py) employs shingle-based Jaccard similarity, automatically switching between brute-force O(n²) comparisons and MinHash LSH based on vault size.
- Both systems use utilities from [`src/hyperresearch/core/similarity.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/similarity.py) for text processing and similarity calculations.
- Provider precedence and similarity thresholds are user-configurable, ensuring deterministic output tailored to specific research workflows.

## Frequently Asked Questions

### How does Hyperresearch handle papers with missing DOIs?

When DOIs are absent, Hyperresearch falls back to `title_key` fingerprinting combined with `_Cluster.year_compatible` checking in [`src/hyperresearch/scholar/dedup.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/scholar/dedup.py). This applies a one-year publication window to account for online-first versus print date variations, allowing preprints to cluster correctly with their final versions despite missing identifiers.

### What similarity threshold should I use for note deduplication?

The default threshold is configurable via the `--threshold` CLI argument, with 0.85 representing a conservative balance that catches near-duplicates while avoiding false positives. Higher values (0.90+) detect only nearly identical content, while lower values (0.70-0.80) capture more paraphrased or templated notes at the risk of increased false matches.

### Why does Hyperresearch use both MinHash LSH and brute-force approaches?

The system dynamically selects algorithms based on the `lsh_switchover` parameter to optimize the performance-accuracy tradeoff. Brute-force Jaccard calculations provide exact similarity scores for small vaults where O(n²) complexity remains feasible, while MinHash LSH reduces computational overhead to near-linear time for large collections, making vault-wide deduplication practical with minimal accuracy loss.

### Can provider precedence be customized when merging papers?

Yes, the `merge_papers` function accepts a `precedence` list parameter that determines which provider's metadata takes priority when fields conflict. For example, setting `precedence=["openalex", "crossref", "core"]` ensures that OpenAlex abstracts and citations are preferred over Crossref or CORE equivalents when constructing the canonical record.