How dupekit Performs Deduplication on Curated Datasets in Marin: A Technical Deep Dive
dupekit performs deduplication on curated datasets in Marin through a hybrid approach combining fuzzy MinHash-LSH similarity detection and exact xxh3 hashing, implemented as a Rust-native extension wrapped in a thin Python proxy.
Marin leverages the dupekit library as its core deduplication engine for processing curated text datasets at scale. dupekit operates as a Python proxy that re-exports a high-performance native Rust extension (dupekit_native), enabling both approximate fuzzy matching and deterministic exact duplicate detection across distributed data shards.
Architectural Overview of dupekit Components
dupekit exposes several transformation primitives that Marin orchestrates into complete deduplication pipelines. According to the dupekit proxy implementation in lib/dupekit/src/dupekit/__init__.py, the library loads the compiled Rust implementation when the native wheel is installed, otherwise raising an informative ImportError on first use.
The core transformations employed by Marin include:
- CleanText – Normalizes raw text via lower-casing, punctuation stripping, and whitespace collapsing to ensure deterministic shingling
- MinHash – Generates MinHash signatures from shingled text using configurable permutations, n-gram size, and seed parameters
- MinHashLSH – Buckets signatures into Locality-Sensitive Hashing (LSH) bands for fast approximate similarity joins
- Hash – Computes fast 128-bit xxh3 hashes for exact deduplication using
dupekit.HashAlgorithm.Xxh3_128 - SplitParagraphs – Decomposes documents into individual paragraphs with span tracking for granular duplicate detection
- SelectColumns – Prunes intermediate columns to maintain lightweight data frames throughout the pipeline
Fuzzy Deduplication with MinHash-LSH
Marin implements fuzzy deduplication through a multi-stage pipeline defined in lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py. This approach identifies near-duplicate documents using probabilistic data structures.
Batch Preparation and Truncation
Each Parquet shard is read as an Arrow RecordBatch. To prevent "mega-document" saturation that could skew MinHash signatures, the pipeline supports optional text truncation via the text_cap_chars parameter. Documents exceeding this length are sliced, with truncation events counted in the minhash/text_truncated counter.
Transformation Pipeline Execution
The fuzzy pipeline executes the following transformation chain in native Rust:
# Conceptual pipeline flow as implemented in fuzzy_minhash.py lines 15-25
CleanText → MinHash → MinHashLSH → SelectColumns
The dupekit.transform call processes the entire batch, yielding a table containing row id values and lists of LSH bucket strings. Key configuration parameters include:
num_perms– Number of MinHash permutations (default typically 286)ngram_size– Size of character n-grams for shingling (default 5)num_bands– Number of LSH bands for bucketing signatures (default 26)seed– Random seed for reproducible hashing
Filtering and Artifact Generation
Empty signatures are dropped and tallied in the minhash/empty_signatures counter. The remaining rows convert bucket lists to a list-of-strings column for downstream grouping. Marin persists these bucket attributes as Parquet files (one per source shard) wrapped in a MinHashAttrData artifact, which stores version, parameters, source key, and processing counters for compatibility validation.
Exact Deduplication Strategies
For deterministic duplicate removal, Marin employs dupekit's Hash transformation with the xxh3_128 algorithm, operating at both paragraph and document granularity.
Paragraph-Level Deduplication
The paragraph-level pipeline in lib/marin/src/marin/processing/classification/deduplication/exact.py uses SplitParagraphs to break documents into individual spans:
# From exact.py lines 90-94
SplitParagraphs creates records with:
- paragraph_text
- paragraph_span (start/end offsets)
Each paragraph receives a deterministic 128-bit hash via dupekit.Transformation.Hash with dupekit.HashAlgorithm.Xxh3_128. Zephyr's distributed group_by operation clusters rows by hash value, selecting the first record (sorted by id) as the canonical keeper and flagging subsequent matches as duplicates (is_dup=True). When include_span=True, duplicate spans attach to output records for audit trails.
Document-Level Deduplication
For whole-document deduplication, Marin hashes the complete text field directly via dupekit.Transformation.Hash (lines 98-100), bypassing paragraph splitting. The grouping logic remains identical—records sharing a hash cluster together, with the earliest id retained and others marked duplicate. This mode sets include_span=False since paragraph-level granularity is unnecessary.
Both exact pipelines leverage Zephyr's distributed execution (ZephyrContext) to scale hash computation across shards before writing results back to Parquet via write_parquet_file.
Integration with Marin's Pipeline Orchestration
Marin orchestrates dupekit transformations through reusable StepSpec definitions and artifact management.
Step Specifications – The compute_minhash_attrs_step function creates a StepSpec that invokes the fuzzy pipeline while recording provenance in the artifact's hash attributes. This specification includes resource requirements and dependency management for Zephyr execution.
Artifact Model – The MinHashAttrData class encapsulates processing metadata including dupekit version, algorithmic parameters, source data keys, and performance counters. Downstream pipelines validate compatibility by inspecting these stored attributes before consuming fuzzy match results.
Summary
- dupekit serves as Marin's deduplication engine through a Rust-native Python proxy architecture, falling back to informative errors if the compiled extension is unavailable.
- Fuzzy deduplication employs MinHash-LSH with configurable permutations, bands, and n-gram sizes, processing data through
CleanText → MinHash → MinHashLSHpipelines. - Exact deduplication uses 128-bit xxh3 hashing at both paragraph and document levels, leveraging
SplitParagraphsfor granular span detection when needed. - The
MinHashAttrDataartifact system tracks processing parameters and counters, enabling reproducible and verifiable deduplication workflows across distributed Zephyr executions.
Frequently Asked Questions
How does dupekit handle missing native extensions?
If the compiled Rust wheel (dupekit_native) is not installed, the proxy package in lib/dupekit/src/dupekit/__init__.py raises an informative ImportError on first use, directing users to install the appropriate platform-specific wheel.
What hash algorithm does dupekit use for exact deduplication?
dupekit uses xxh3_128, a fast 128-bit non-cryptographic hash algorithm implemented in dupekit.HashAlgorithm.Xxh3_128. This provides collision-resistant hashing suitable for large-scale duplicate detection while maintaining high throughput.
Can dupekit handle very long documents in fuzzy matching?
Yes. The fuzzy pipeline accepts a text_cap_chars parameter that truncates documents exceeding the specified character limit, preventing "mega-doc" saturation that could degrade MinHash signature quality. Truncation events are tracked in the minhash/text_truncated counter.
How does Marin scale dupekit deduplication across large datasets?
Marin integrates dupekit with Zephyr's distributed execution framework (ZephyrContext), which parallelizes transformation pipelines across data shards. Both exact and fuzzy pipelines use this architecture to distribute computation while collecting results into unified Parquet outputs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →