How to Perform Text Deduplication with Marin-dupekit: 3 Methods Explained

Marin-dupekit provides three distributed deduplication modes—Exact Paragraph, Exact Document, and Fuzzy Document—that leverage the Rust-based dupekit library to remove duplicates from petabyte-scale text corpora using Zephyr's execution engine.

Marin-dupekit is the deduplication framework within the marin-community/marin repository designed for large-scale text processing pipelines. It offers deterministic exact matching and probabilistic fuzzy matching capabilities built on top of the high-performance dupekit Rust library. Whether cleaning web crawl data or deduplicating document uploads, Marin-dupekit integrates with Zephyr's distributed engine to process Parquet and JSONL shards at scale.

Understanding the Three Deduplication Modes

Marin-dupekit implements three distinct strategies for identifying redundant content:

  • Exact Paragraph Deduplication: Removes duplicated paragraphs within individual documents by hashing each paragraph and retaining the first occurrence. This strict, deterministic approach is implemented in dedup_exact_paragraph and is ideal when you need to eliminate repeated sections inside long texts.

  • Exact Document Deduplication: Detects completely identical documents by computing a 128-bit Xxh3 hash of the entire text field. The dedup_exact_document function identifies verbatim duplicates across your dataset, making it suitable for removing duplicate file uploads.

  • Fuzzy Document Deduplication: Finds near-duplicate documents using a MinHash and Locality-Sensitive Hashing (LSH) pipeline. This probabilistic method groups semantically similar texts through connected-components graph analysis, implemented across compute_minhash_attrs and compute_fuzzy_dups_attrs.

Core Architecture and Transformations

All deduplication pipelines in Marin-dupekit rely on low-level dupekit transformations from the Rust library:

  1. Text Normalization: CleanText lowercases input, strips punctuation, and collapses whitespace to ensure consistent hashing.

  2. Segmentation: SplitParagraphs breaks documents into discrete paragraph records for granular deduplication.

  3. Signature Generation:

    • Hash produces 128-bit Xxh3 hashes for exact matching scenarios.
    • MinHash creates signature vectors using configurable permutations, while MinHashLSH buckets these signatures into bands to enable efficient similarity search.

The pipelines execute on Zephyr's distributed engine, where Dataset objects load shards and apply flat_map or group_by operations. Throughout execution, Zephyr counters (counters.pipeline.update_counter) track metrics such as processed documents, empty signatures, and duplicate totals. Jobs also initialize Weights & Biases runs via _init_wandb for metric surfacing and provenance tracking.

Exact Deduplication Methods

For scenarios requiring deterministic de-duplication, Marin-dupekit provides two exact modes in exact.py.

Exact Paragraph Deduplication

The dedup_exact_paragraph function processes documents paragraph-by-paragraph, keeping only the first occurrence of each unique paragraph hash.

from marin.processing.classification.deduplication.exact import dedup_exact_paragraph

result = dedup_exact_paragraph(
    input_paths="gs://my-bucket/raw-data/",
    output_path="gs://my-bucket/deduped/paragraphs/",
    text_field="text",
)
print(result)   # dict with total/dups/unique counters

Implementation reference: [exact.py](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/exact.py#L75-L74).

Exact Document Deduplication

Use dedup_exact_document to identify completely identical files across your dataset by hashing the entire text content.

from marin.processing.classification.deduplication.exact import dedup_exact_document

result = dedup_exact_document(
    input_paths=["gs://my-bucket/raw-data/"],
    output_path="gs://my-bucket/deduped/documents/",
    text_field="text",
    max_parallelism=8,
)
print(result)   # dict with total/dups/unique counters

Implementation reference: [exact.py](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/exact.py#L77-L86).

Fuzzy Document Deduplication Pipeline

Fuzzy matching requires a two-stage process to handle the computational complexity of similarity detection across large corpora.

Step 1: Computing MinHash Signatures

First, generate MinHash attributes using compute_minhash_attrs in fuzzy_minhash.py. This step creates signature vectors and LSH buckets for each document.

from marin.processing.classification.deduplication.fuzzy_minhash import compute_minhash_attrs

minhash_attrs = compute_minhash_attrs(
    source=NormalizedData("gs://my-bucket/normalized/"),
    output_path="gs://my-bucket/minhash-attrs/",
    num_perms=286,
    num_bands=26,
    ngram_size=5,
    text_cap_chars=500_000,
    seed=42,
)

Key parameters include num_perms (hash permutations), which must be divisible by num_bands (LSH bands). The text_cap_chars parameter optionally truncates very long documents to improve performance.

Implementation reference: [fuzzy_minhash.py](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py#L41-L56).

Step 2: Detecting Fuzzy Duplicates

Consume the MinHash attributes to compute duplicate clusters using connected-components analysis.

from marin.processing.classification.deduplication.fuzzy_dups import compute_fuzzy_dups_attrs

fuzzy_attrs = compute_fuzzy_dups_attrs(
    inputs=[minhash_attrs],
    output_path="gs://my-bucket/fuzzy-dups/",
    max_parallelism=12,
)

This produces FuzzyDupsAttrData artifacts containing duplicate markers co-partitioned with your original data.

Implementation reference: [fuzzy_dups.py](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/fuzzy_dups.py#L24-L30).

Working with Deduplication Artifacts

Both exact and fuzzy modes produce co-partitioned Parquet artifacts that enable downstream processing without reshuffling data.

Exact modes output files containing id, is_dup, and optionally span columns. Fuzzy modes generate MinHashAttrData (from fuzzy_minhash.py) and FuzzyDupsAttrData (from fuzzy_dups.py) objects that expose the attr_dir path to per-shard Parquet files.

Load results using the artifact API:

from marin.execution.artifact import read_artifact

mh_attrs = read_artifact(step="minhash-step", artifact_type=MinHashAttrData)
fd_attrs = read_artifact(step="fuzzy-step", artifact_type=FuzzyDupsAttrData)

These Pydantic models provide type-safe access to deduplication metadata and file locations.

Key Source Files

The deduplication system spans four primary modules in lib/marin/src/marin/processing/classification/deduplication/:

  • fuzzy_minhash.py: Implements the MinHash signature and LSH bucketing pipeline.
  • fuzzy_dups.py: Executes global connected-components graph analysis on MinHash attributes.
  • exact.py: Contains dedup_exact_paragraph and dedup_exact_document hash-based pipelines.
  • dedup_commons.py: Provides shared utilities including WandB initialization, batch loading, and counter management.

Summary

  • Marin-dupekit offers three deduplication strategies: Exact Paragraph, Exact Document, and Fuzzy Document, each optimized for different redundancy patterns.
  • Exact modes use 128-bit Xxh3 hashing for deterministic duplicate detection at paragraph or document granularity.
  • Fuzzy mode employs a two-stage MinHash and LSH pipeline to identify near-duplicates through probabilistic similarity hashing.
  • All pipelines run on Zephyr's distributed engine, producing co-partitioned Parquet artifacts via MinHashAttrData and FuzzyDupsAttrData models.
  • Implementation files are located in lib/marin/src/marin/processing/classification/deduplication/, with exact.py handling hash-based deduplication and fuzzy_minhash.py/fuzzy_dups.py handling similarity clustering.

Frequently Asked Questions

What is the difference between exact and fuzzy deduplication in Marin-dupekit?

Exact deduplication uses cryptographic hashing (Xxh3) to identify byte-for-byte identical paragraphs or documents, making it deterministic and suitable for removing verbatim copies. Fuzzy deduplication uses MinHash signatures and LSH bucketing to detect near-duplicates where text may have slight variations, making it appropriate for identifying plagiarized or reformatted content that exact matching would miss.

How does the MinHash LSH pipeline work for fuzzy matching?

The pipeline first computes MinHash signatures using compute_minhash_attrs, which generates fixed-length vectors representing document content. These signatures are then divided into bands (num_bands) and hashed into buckets using MinHashLSH. Documents sharing buckets are candidate pairs, which compute_fuzzy_dups_attrs resolves into duplicate clusters via connected-components graph analysis, efficiently grouping similar documents without pairwise comparisons of the entire corpus.

Can I configure the sensitivity of fuzzy deduplication?

Yes, sensitivity is controlled through the num_perms (permutations) and num_bands parameters in compute_minhash_attrs. Increasing num_perms improves signature accuracy but requires more computation. The num_bands value must divide num_perms evenly and determines the LSH collision probability—fewer bands increase recall (find more duplicates) but may reduce precision, while more bands require higher similarity for matching.

Where are the deduplication results stored in Marin-dupekit?

Results are written to the output_path specified in each function call as co-partitioned Parquet files. Exact modes produce simple Parquet files with duplicate flags, while fuzzy modes generate MinHashAttrData and FuzzyDupsAttrData artifacts. These Pydantic models expose the attr_dir attribute pointing to the directory containing per-shard Parquet files, which can be loaded via read_artifact or accessed directly for downstream verification steps.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →