How Marin Handles Text Deduplication: Exact, Paragraph, and Fuzzy Matching Explained

Marin handles text deduplication through three configurable strategies—exact paragraph, exact document, and fuzzy document detection—powered by the Zephyr execution engine and locality-sensitive MinHash algorithms.

Marin is an open-source framework designed for large-scale text corpus cleaning and preprocessing. Understanding how Marin handles text deduplication is critical for preparing training datasets where duplicate content can bias machine learning models. The system implements a unified pipeline defined in lib/marin/src/marin/processing/classification/deduplication/dedup_commons.py that supports byte-identical matching and near-duplicate detection at scale.

The Three Deduplication Strategies in Marin

Marin’s deduplication logic centers around the DedupMode enum defined in dedup_commons.py (lines 26-41). Each variant targets a distinct duplication pattern, from identical paragraphs to semantically similar documents.

Exact Paragraph Deduplication (EXACT_PARAGRAPH)

The EXACT_PARAGRAPH mode targets repeated text blocks within and across documents. Implemented in dedup_exact_paragraph within lib/marin/src/marin/processing/classification/deduplication/exact.py (starting at line 75), this strategy splits each record into individual paragraphs.

It then hashes these paragraphs using dupekit with the Xxh3_128 algorithm, groups records by hash value, and flags all occurrences after the first as duplicates. This method efficiently removes boilerplate text, headers, and footers that repeat across a corpus.

from marin.processing.classification.deduplication.exact import dedup_exact_paragraph

result = dedup_exact_paragraph(
    input_paths="/data/raw_documents/*.parquet",
    output_path="/data/deduped/paragraphs",
    text_field="text",           # column containing raw text

)
print(result)   # → {'dedup/exact/paragraph/total': ..., 'dups': ..., 'unique': ...}

Exact Document Deduplication (EXACT_DOCUMENT)

For completely identical files, Marin provides EXACT_DOCUMENT mode via dedup_exact_document in exact.py (starting at line 87). This approach computes a single hash for the entire document content rather than individual paragraphs.

Records are grouped by their document-level hash, and all non-canonical copies are marked as duplicates. This method offers the highest throughput for removing exact file copies in massive datasets.

from marin.processing.classification.deduplication.exact import dedup_exact_document

result = dedup_exact_document(
    input_paths="/data/raw_documents",
    output_path="/data/deduped/documents",
    text_field="content",
    max_parallelism=8,           # number of workers for the Zephyr map stage

)
print(result)   # → {'dedup/exact/document/total': ..., 'dups': ..., 'unique': ...}

Fuzzy Document Deduplication (FUZZY_DOCUMENT)

The FUZZY_DOCUMENT mode detects near-duplicates that are not byte-identical but share substantial content. This three-stage pipeline spans three specialized modules:

  1. MinHash fingerprinting – lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py generates locality-sensitive hashes for each document, creating compact signatures that preserve similarity relationships.

  2. Candidate generation – lib/marin/src/marin/processing/classification/deduplication/fuzzy_dups.py groups documents sharing sufficient MinHash buckets, producing candidate duplicate sets without expensive pairwise comparisons.

  3. Verification – lib/marin/src/marin/processing/classification/deduplication/fuzzy_verification.py performs a precise token-sequence similarity check to confirm or reject candidates, ensuring high precision in the final duplicate set.

This architecture catches documents with minor edits, formatting variations, or inserted boilerplate that exact hashing would miss.

from marin.processing.classification.deduplication.fuzzy_dups import compute_fuzzy_dups_attrs_step
from marin.processing.classification.deduplication.fuzzy_minhash import compute_minhash_attrs_step

# In a Zephyr pipeline you would chain the two steps, then run verification:

# (simplified example – the real pipeline is built via `ZephyrContext.execute`)

pipeline = (
    Dataset.from_list(input_files)
    .flat_map(compute_minhash_attrs_step)
    .group_by(... )                 # bucket by MinHash signatures

    .flat_map(compute_fuzzy_dups_attrs_step)
    .group_by(... )                 # verification reducer

)

Shared Infrastructure and Pipeline Execution

All three deduplication modes rely on common scaffolding defined in dedup_commons.py. This shared infrastructure handles file I/O, batch processing, and metrics aggregation.

Core Utilities in dedup_commons.py

The module provides essential helper functions that standardize the deduplication workflow:

  • File collection – _collect_input_files discovers and validates input parquet and jsonl files (lines 55-71).

  • Batch loading – _load_batches streams each file as PyArrow RecordBatch objects for memory-efficient processing (lines 94-108).

  • Result aggregation – _aggregate_shard_counters and the _DupTally class combine per-shard statistics and write the final deduplicated parquet output.

  • Experiment tracking – _init_wandb and finalize_dedup automatically initialize Weights & Biases runs and log deduplication counters, including total documents, duplicates found, and unique records retained.

Zephyr Execution Engine

The computational heavy lifting executes on the Zephyr engine. Both exact and fuzzy pipelines instantiate a ZephyrContext (as seen in dedup_exact_paragraph lines 103-104) and invoke ctx.execute(...) to run the distributed computation graph.

The typical execution flow follows this pattern:

Dataset → flat_map (hash generation) → group_by (hash → canonical record) → 
group_by (file index) → write_parquet → aggregate counters

For fuzzy deduplication, the graph extends with additional stages for MinHash computation and verification reducers while maintaining the same data-flow paradigm. This design allows Marin to scale deduplication operations across distributed clusters using the same API for exact and fuzzy matching.

Summary

  • Marin supports three deduplication strategies defined by the DedupMode enum in dedup_commons.py: exact paragraph, exact document, and fuzzy document detection.

  • Exact deduplication relies on dupekit (Xxh3_128) hashing via functions in exact.py, offering high-speed removal of byte-identical content.

  • Fuzzy deduplication implements a three-stage pipeline using fuzzy_minhash.py, fuzzy_dups.py, and fuzzy_verification.py to detect near-duplicates through locality-sensitive hashing and token-sequence verification.

  • All modes leverage shared infrastructure in dedup_commons.py for file collection, PyArrow batch loading, and automatic Weights & Biases metrics logging.

  • The Zephyr execution engine powers the distributed computation graph, enabling scalable processing of massive text corpora.

Frequently Asked Questions

What hashing algorithm does Marin use for exact deduplication?

Marin uses the Xxh3_128 algorithm provided by dupekit to generate document and paragraph hashes. This implementation appears in exact.py within the dedup_exact_paragraph and dedup_exact_document functions, providing fast, low-collision hashing suitable for large-scale deduplication tasks.

How does Marin handle fuzzy matching for near-duplicate documents?

Marin implements fuzzy matching through a three-stage pipeline: first generating MinHash fingerprints in fuzzy_minhash.py, then grouping candidates by shared buckets in fuzzy_dups.py, and finally running token-sequence verification in fuzzy_verification.py. This approach identifies documents with minor textual variations while avoiding expensive pairwise comparisons across the entire corpus.

Can Marin deduplicate both paragraphs and full documents?

Yes. Marin provides distinct modes for both granularities. The EXACT_PARAGRAPH mode in dedup_exact_paragraph targets repeated text blocks within documents, while EXACT_DOCUMENT and **FUZZY_DOCUMENToperate on complete files. Users select the appropriate strategy via theDedupMode` enum depending on whether they need to remove boilerplate text or eliminate duplicate files.

What execution engine powers Marin's deduplication pipeline?

Marin uses the Zephyr execution engine to distribute deduplication workloads. The pipelines create a ZephyrContext and build execution graphs that map hash generation, group-by operations, and parquet writes across compute clusters. This architecture appears consistently across exact and fuzzy implementations in the lib/marin/src/marin/processing/classification/deduplication/ directory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →