# How Marin Handles Text Deduplication: Exact, Paragraph, and Fuzzy Matching Explained

> Discover how Marin handles text deduplication with exact paragraph, exact document, and fuzzy matching. Learn about its powerful Zephyr engine and MinHash algorithms.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: deep-dive
- Published: 2026-09-10

---

**Marin handles text deduplication through three configurable strategies—exact paragraph, exact document, and fuzzy document detection—powered by the Zephyr execution engine and locality-sensitive MinHash algorithms.**

Marin is an open-source framework designed for large-scale text corpus cleaning and preprocessing. Understanding how Marin handles text deduplication is critical for preparing training datasets where duplicate content can bias machine learning models. The system implements a unified pipeline defined in [`lib/marin/src/marin/processing/classification/deduplication/dedup_commons.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/dedup_commons.py) that supports byte-identical matching and near-duplicate detection at scale.

## The Three Deduplication Strategies in Marin

Marin’s deduplication logic centers around the `DedupMode` enum defined in [`dedup_commons.py`](https://github.com/marin-community/marin/blob/main/dedup_commons.py) (lines 26-41). Each variant targets a distinct duplication pattern, from identical paragraphs to semantically similar documents.

### Exact Paragraph Deduplication (EXACT_PARAGRAPH)

The **EXACT_PARAGRAPH** mode targets repeated text blocks within and across documents. Implemented in `dedup_exact_paragraph` within [`lib/marin/src/marin/processing/classification/deduplication/exact.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/exact.py) (starting at line 75), this strategy splits each record into individual paragraphs.

It then hashes these paragraphs using **dupekit** with the `Xxh3_128` algorithm, groups records by hash value, and flags all occurrences after the first as duplicates. This method efficiently removes boilerplate text, headers, and footers that repeat across a corpus.

```python
from marin.processing.classification.deduplication.exact import dedup_exact_paragraph

result = dedup_exact_paragraph(
    input_paths="/data/raw_documents/*.parquet",
    output_path="/data/deduped/paragraphs",
    text_field="text",           # column containing raw text

)
print(result)   # → {'dedup/exact/paragraph/total': ..., 'dups': ..., 'unique': ...}

```

### Exact Document Deduplication (EXACT_DOCUMENT)

For completely identical files, Marin provides **EXACT_DOCUMENT** mode via `dedup_exact_document` in [`exact.py`](https://github.com/marin-community/marin/blob/main/exact.py) (starting at line 87). This approach computes a single hash for the entire document content rather than individual paragraphs.

Records are grouped by their document-level hash, and all non-canonical copies are marked as duplicates. This method offers the highest throughput for removing exact file copies in massive datasets.

```python
from marin.processing.classification.deduplication.exact import dedup_exact_document

result = dedup_exact_document(
    input_paths="/data/raw_documents",
    output_path="/data/deduped/documents",
    text_field="content",
    max_parallelism=8,           # number of workers for the Zephyr map stage

)
print(result)   # → {'dedup/exact/document/total': ..., 'dups': ..., 'unique': ...}

```

### Fuzzy Document Deduplication (FUZZY_DOCUMENT)

The **FUZZY_DOCUMENT** mode detects near-duplicates that are not byte-identical but share substantial content. This three-stage pipeline spans three specialized modules:

1. **MinHash fingerprinting** – [`lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py) generates locality-sensitive hashes for each document, creating compact signatures that preserve similarity relationships.

2. **Candidate generation** – [`lib/marin/src/marin/processing/classification/deduplication/fuzzy_dups.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/fuzzy_dups.py) groups documents sharing sufficient MinHash buckets, producing candidate duplicate sets without expensive pairwise comparisons.

3. **Verification** – [`lib/marin/src/marin/processing/classification/deduplication/fuzzy_verification.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/fuzzy_verification.py) performs a precise token-sequence similarity check to confirm or reject candidates, ensuring high precision in the final duplicate set.

This architecture catches documents with minor edits, formatting variations, or inserted boilerplate that exact hashing would miss.

```python
from marin.processing.classification.deduplication.fuzzy_dups import compute_fuzzy_dups_attrs_step
from marin.processing.classification.deduplication.fuzzy_minhash import compute_minhash_attrs_step

# In a Zephyr pipeline you would chain the two steps, then run verification:

# (simplified example – the real pipeline is built via `ZephyrContext.execute`)

pipeline = (
    Dataset.from_list(input_files)
    .flat_map(compute_minhash_attrs_step)
    .group_by(... )                 # bucket by MinHash signatures

    .flat_map(compute_fuzzy_dups_attrs_step)
    .group_by(... )                 # verification reducer

)

```

## Shared Infrastructure and Pipeline Execution

All three deduplication modes rely on common scaffolding defined in [`dedup_commons.py`](https://github.com/marin-community/marin/blob/main/dedup_commons.py). This shared infrastructure handles file I/O, batch processing, and metrics aggregation.

### Core Utilities in dedup_commons.py

The module provides essential helper functions that standardize the deduplication workflow:

- **File collection** – `_collect_input_files` discovers and validates input parquet and jsonl files (lines 55-71).

- **Batch loading** – `_load_batches` streams each file as PyArrow `RecordBatch` objects for memory-efficient processing (lines 94-108).

- **Result aggregation** – `_aggregate_shard_counters` and the `_DupTally` class combine per-shard statistics and write the final deduplicated parquet output.

- **Experiment tracking** – `_init_wandb` and `finalize_dedup` automatically initialize Weights & Biases runs and log deduplication counters, including total documents, duplicates found, and unique records retained.

### Zephyr Execution Engine

The computational heavy lifting executes on the **Zephyr** engine. Both exact and fuzzy pipelines instantiate a `ZephyrContext` (as seen in `dedup_exact_paragraph` lines 103-104) and invoke `ctx.execute(...)` to run the distributed computation graph.

The typical execution flow follows this pattern:

```text
Dataset → flat_map (hash generation) → group_by (hash → canonical record) → 
group_by (file index) → write_parquet → aggregate counters

```

For fuzzy deduplication, the graph extends with additional stages for MinHash computation and verification reducers while maintaining the same data-flow paradigm. This design allows Marin to scale deduplication operations across distributed clusters using the same API for exact and fuzzy matching.

## Summary

- Marin supports **three deduplication strategies** defined by the `DedupMode` enum in [`dedup_commons.py`](https://github.com/marin-community/marin/blob/main/dedup_commons.py): exact paragraph, exact document, and fuzzy document detection.

- **Exact deduplication** relies on `dupekit` (Xxh3_128) hashing via functions in [`exact.py`](https://github.com/marin-community/marin/blob/main/exact.py), offering high-speed removal of byte-identical content.

- **Fuzzy deduplication** implements a three-stage pipeline using [`fuzzy_minhash.py`](https://github.com/marin-community/marin/blob/main/fuzzy_minhash.py), [`fuzzy_dups.py`](https://github.com/marin-community/marin/blob/main/fuzzy_dups.py), and [`fuzzy_verification.py`](https://github.com/marin-community/marin/blob/main/fuzzy_verification.py) to detect near-duplicates through locality-sensitive hashing and token-sequence verification.

- All modes leverage shared infrastructure in [`dedup_commons.py`](https://github.com/marin-community/marin/blob/main/dedup_commons.py) for file collection, PyArrow batch loading, and automatic Weights & Biases metrics logging.

- The **Zephyr execution engine** powers the distributed computation graph, enabling scalable processing of massive text corpora.

## Frequently Asked Questions

### What hashing algorithm does Marin use for exact deduplication?

Marin uses the **Xxh3_128** algorithm provided by **dupekit** to generate document and paragraph hashes. This implementation appears in [`exact.py`](https://github.com/marin-community/marin/blob/main/exact.py) within the `dedup_exact_paragraph` and `dedup_exact_document` functions, providing fast, low-collision hashing suitable for large-scale deduplication tasks.

### How does Marin handle fuzzy matching for near-duplicate documents?

Marin implements fuzzy matching through a three-stage pipeline: first generating **MinHash fingerprints** in [`fuzzy_minhash.py`](https://github.com/marin-community/marin/blob/main/fuzzy_minhash.py), then grouping candidates by shared buckets in [`fuzzy_dups.py`](https://github.com/marin-community/marin/blob/main/fuzzy_dups.py), and finally running **token-sequence verification** in [`fuzzy_verification.py`](https://github.com/marin-community/marin/blob/main/fuzzy_verification.py). This approach identifies documents with minor textual variations while avoiding expensive pairwise comparisons across the entire corpus.

### Can Marin deduplicate both paragraphs and full documents?

Yes. Marin provides distinct modes for both granularities. The **EXACT_PARAGRAPH** mode in `dedup_exact_paragraph` targets repeated text blocks within documents, while **EXACT_DOCUMENT** and **FUZZY_DOCUMENT` operate on complete files. Users select the appropriate strategy via the `DedupMode` enum depending on whether they need to remove boilerplate text or eliminate duplicate files.

### What execution engine powers Marin's deduplication pipeline?

Marin uses the **Zephyr** execution engine to distribute deduplication workloads. The pipelines create a `ZephyrContext` and build execution graphs that map hash generation, group-by operations, and parquet writes across compute clusters. This architecture appears consistently across exact and fuzzy implementations in the `lib/marin/src/marin/processing/classification/deduplication/` directory.