# How dupekit Performs Deduplication on Curated Datasets in Marin: A Technical Deep Dive

> Discover how dupekit performs deduplication on curated datasets in Marin using a hybrid approach of fuzzy MinHash-LSH and exact xxh3 hashing for efficient data cleaning.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: deep-dive
- Published: 2026-08-28

---

**dupekit performs deduplication on curated datasets in Marin through a hybrid approach combining fuzzy MinHash-LSH similarity detection and exact xxh3 hashing, implemented as a Rust-native extension wrapped in a thin Python proxy.**

Marin leverages the **dupekit** library as its core deduplication engine for processing curated text datasets at scale. dupekit operates as a Python proxy that re-exports a high-performance native Rust extension (`dupekit_native`), enabling both approximate fuzzy matching and deterministic exact duplicate detection across distributed data shards.

## Architectural Overview of dupekit Components

dupekit exposes several transformation primitives that Marin orchestrates into complete deduplication pipelines. According to the dupekit proxy implementation in [`lib/dupekit/src/dupekit/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/dupekit/src/dupekit/__init__.py), the library loads the compiled Rust implementation when the native wheel is installed, otherwise raising an informative `ImportError` on first use.

The core transformations employed by Marin include:

- **CleanText** – Normalizes raw text via lower-casing, punctuation stripping, and whitespace collapsing to ensure deterministic shingling
- **MinHash** – Generates MinHash signatures from shingled text using configurable permutations, n-gram size, and seed parameters
- **MinHashLSH** – Buckets signatures into Locality-Sensitive Hashing (LSH) bands for fast approximate similarity joins
- **Hash** – Computes fast 128-bit xxh3 hashes for exact deduplication using `dupekit.HashAlgorithm.Xxh3_128`
- **SplitParagraphs** – Decomposes documents into individual paragraphs with span tracking for granular duplicate detection
- **SelectColumns** – Prunes intermediate columns to maintain lightweight data frames throughout the pipeline

## Fuzzy Deduplication with MinHash-LSH

Marin implements fuzzy deduplication through a multi-stage pipeline defined in [`lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py). This approach identifies near-duplicate documents using probabilistic data structures.

### Batch Preparation and Truncation

Each Parquet shard is read as an Arrow `RecordBatch`. To prevent "mega-document" saturation that could skew MinHash signatures, the pipeline supports optional text truncation via the `text_cap_chars` parameter. Documents exceeding this length are sliced, with truncation events counted in the `minhash/text_truncated` counter.

### Transformation Pipeline Execution

The fuzzy pipeline executes the following transformation chain in native Rust:

```python

# Conceptual pipeline flow as implemented in fuzzy_minhash.py lines 15-25

CleanText → MinHash → MinHashLSH → SelectColumns

```

The `dupekit.transform` call processes the entire batch, yielding a table containing row `id` values and lists of LSH bucket strings. Key configuration parameters include:

- `num_perms` – Number of MinHash permutations (default typically 286)
- `ngram_size` – Size of character n-grams for shingling (default 5)
- `num_bands` – Number of LSH bands for bucketing signatures (default 26)
- `seed` – Random seed for reproducible hashing

### Filtering and Artifact Generation

Empty signatures are dropped and tallied in the `minhash/empty_signatures` counter. The remaining rows convert bucket lists to a list-of-strings column for downstream grouping. Marin persists these bucket attributes as Parquet files (one per source shard) wrapped in a `MinHashAttrData` artifact, which stores version, parameters, source key, and processing counters for compatibility validation.

## Exact Deduplication Strategies

For deterministic duplicate removal, Marin employs dupekit's `Hash` transformation with the xxh3_128 algorithm, operating at both paragraph and document granularity.

### Paragraph-Level Deduplication

The paragraph-level pipeline in [`lib/marin/src/marin/processing/classification/deduplication/exact.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/classification/deduplication/exact.py) uses `SplitParagraphs` to break documents into individual spans:

```python

# From exact.py lines 90-94

SplitParagraphs creates records with:
- paragraph_text
- paragraph_span (start/end offsets)

```

Each paragraph receives a deterministic 128-bit hash via `dupekit.Transformation.Hash` with `dupekit.HashAlgorithm.Xxh3_128`. Zephyr's distributed `group_by` operation clusters rows by hash value, selecting the first record (sorted by `id`) as the canonical keeper and flagging subsequent matches as duplicates (`is_dup=True`). When `include_span=True`, duplicate spans attach to output records for audit trails.

### Document-Level Deduplication

For whole-document deduplication, Marin hashes the complete text field directly via `dupekit.Transformation.Hash` (lines 98-100), bypassing paragraph splitting. The grouping logic remains identical—records sharing a hash cluster together, with the earliest `id` retained and others marked duplicate. This mode sets `include_span=False` since paragraph-level granularity is unnecessary.

Both exact pipelines leverage Zephyr's distributed execution (`ZephyrContext`) to scale hash computation across shards before writing results back to Parquet via `write_parquet_file`.

## Integration with Marin's Pipeline Orchestration

Marin orchestrates dupekit transformations through reusable `StepSpec` definitions and artifact management.

**Step Specifications** – The `compute_minhash_attrs_step` function creates a `StepSpec` that invokes the fuzzy pipeline while recording provenance in the artifact's hash attributes. This specification includes resource requirements and dependency management for Zephyr execution.

**Artifact Model** – The `MinHashAttrData` class encapsulates processing metadata including dupekit version, algorithmic parameters, source data keys, and performance counters. Downstream pipelines validate compatibility by inspecting these stored attributes before consuming fuzzy match results.

## Summary

- dupekit serves as Marin's deduplication engine through a Rust-native Python proxy architecture, falling back to informative errors if the compiled extension is unavailable.
- **Fuzzy deduplication** employs MinHash-LSH with configurable permutations, bands, and n-gram sizes, processing data through `CleanText → MinHash → MinHashLSH` pipelines.
- **Exact deduplication** uses 128-bit xxh3 hashing at both paragraph and document levels, leveraging `SplitParagraphs` for granular span detection when needed.
- The `MinHashAttrData` artifact system tracks processing parameters and counters, enabling reproducible and verifiable deduplication workflows across distributed Zephyr executions.

## Frequently Asked Questions

### How does dupekit handle missing native extensions?

If the compiled Rust wheel (`dupekit_native`) is not installed, the proxy package in [`lib/dupekit/src/dupekit/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/dupekit/src/dupekit/__init__.py) raises an informative `ImportError` on first use, directing users to install the appropriate platform-specific wheel.

### What hash algorithm does dupekit use for exact deduplication?

dupekit uses **xxh3_128**, a fast 128-bit non-cryptographic hash algorithm implemented in `dupekit.HashAlgorithm.Xxh3_128`. This provides collision-resistant hashing suitable for large-scale duplicate detection while maintaining high throughput.

### Can dupekit handle very long documents in fuzzy matching?

Yes. The fuzzy pipeline accepts a `text_cap_chars` parameter that truncates documents exceeding the specified character limit, preventing "mega-doc" saturation that could degrade MinHash signature quality. Truncation events are tracked in the `minhash/text_truncated` counter.

### How does Marin scale dupekit deduplication across large datasets?

Marin integrates dupekit with Zephyr's distributed execution framework (`ZephyrContext`), which parallelizes transformation pipelines across data shards. Both exact and fuzzy pipelines use this architecture to distribute computation while collecting results into unified Parquet outputs.