Handling Near-Duplicate Detection in Instruction Datasets: A Practical Guide
The rasbt/LLMs-from-scratch repository provides a production-ready utility in ch07/02_dataset-utilities/find-near-duplicates.py that detects and removes near-duplicate entries in JSON instruction datasets using character-level TF-IDF vectorization and cosine similarity scoring.
Handling near-duplicate detection in instruction datasets is critical for preventing data leakage and ensuring model generalization during fine-tuning. The utility implements a robust three-stage pipeline that normalizes text variations, converts entries into vector representations, and calculates pairwise similarity scores to identify redundant training examples.
The Three-Stage Detection Pipeline
The core logic in find-near-duplicates.py processes instruction datasets through distinct phases designed to catch semantic duplicates while preserving legitimate variations.
Text Normalization via preprocess_text
Each text field undergoes standardization through the preprocess_text function. This lowercases all characters and strips punctuation, ensuring that superficial differences like "What's" versus "what is" do not prevent duplicate detection. This normalization step is crucial for instruction datasets where minor formatting variations often mask identical underlying queries.
Character-Level TF-IDF Vectorization
The cleaned strings are transformed into numerical vectors using TfidfVectorizer(analyzer="char", ngram_range=(1,3)). Character-level n-grams capture small lexical changes and typos while remaining robust to minor wording differences that word-level tokenizers might miss. This approach is particularly effective for instruction datasets where similar prompts may use synonymous terms or slight paraphrasing.
Cosine Similarity Thresholding
The pipeline computes a full pairwise cosine similarity matrix using cosine_similarity from scikit-learn. Any pair exceeding the user-defined threshold (default 0.75) is flagged as a near duplicate. The system distinguishes between detection and removal: duplicates found in the instruction field are reported but never auto-deleted, while matches in input or output fields can be marked for removal to prevent repetitive training examples.
Command-Line and Programmatic Usage
The utility supports both interactive command-line workflows and direct Python module integration.
CLI Workflow for Dataset Cleaning
The script accepts standard UNIX-style arguments for processing JSON files directly:
python find-near-duplicates.py \
--json_file data/instructions.json \
--threshold 0.80 \
--remove_duplicates \
--json_output_file data/instructions_clean.json
This invocation loads the dataset, prints near-duplicate pairs for each key (instruction, input, output), removes duplicates from the input and output fields only, and writes the cleaned list to instructions_clean.json.
Python API: find_near_duplicates Function
For integration into existing data pipelines, import the detection logic directly:
import json
from find_near_duplicates import find_near_duplicates
# Load dataset
with open("data/instructions.json") as f:
dataset = json.load(f)
# Detect near-duplicates in instruction field with strict threshold
filtered, dup_pairs = find_near_duplicates(
dataset,
threshold=0.90,
key="instruction"
)
print(f"Found {len(dup_pairs)} near-duplicate pairs")
for a, b, sim in dup_pairs:
print(f"Similarity: {sim:.2f}")
print(f" 1: {a['instruction']}")
print(f" 2: {b['instruction']}")
This returns a tuple containing the filtered dataset (identical to input when key="instruction" since instructions are never removed) and a list of duplicate triples (item_a, item_b, similarity_score).
Automated Removal with find_print_and_remove_near_duplicates
For comprehensive cleaning across all fields with automatic removal:
from find_near_duplicates import find_print_and_remove_near_duplicates
cleaned = find_print_and_remove_near_duplicates(
json_data=dataset,
remove_duplicates=True,
threshold=0.85
)
When remove_duplicates=True, the function iterates through all keys in the first JSON object, prints similarity reports for each field, and automatically removes duplicates from input and output fields while preserving all instruction prompts.
Key Implementation Details
The utility deliberately separates detection from removal to prevent over-aggressive pruning. This two-step workflow allows data scientists to inspect similarity scores before committing to deletions. The character-level n-gram approach (range 1-3) provides superior performance for short instruction texts compared to word-level embeddings, capturing typos and minor morphological variations without requiring GPU resources.
According to the source code in ch07/02_dataset-utilities/find-near-duplicates.py, the threshold parameter accepts float values between 0 and 1, where higher values increase strictness. The default of 0.75 strikes a balance between catching paraphrased duplicates and preserving legitimately distinct examples.
Summary
find-near-duplicates.pyimplements a lightweight, CPU-based near-duplicate detection system using TF-IDF vectorization and cosine similarity.- Character-level n-grams (1-3) capture small textual variations better than word-level approaches for short instruction texts.
- Selective removal logic preserves instruction prompts while allowing deduplication of inputs and outputs to prevent training data redundancy.
- Both CLI and Python API interfaces support integration into MLOps pipelines with configurable similarity thresholds.
Frequently Asked Questions
What similarity threshold should I use for instruction datasets?
The default 0.75 threshold works well for most instruction-following datasets. Increase to 0.85-0.90 for stricter deduplication when dealing with templated prompts, or decrease to 0.65-0.70 if you want to catch only near-identical matches while preserving paraphrased variations.
Why does the utility use character-level instead of word-level n-grams?
Character-level n-grams catch typos, punctuation differences, and minor spelling variations that word-level tokenizers miss. This is particularly important for user-generated instruction datasets where "color" and "colour" or "what's" and "whats" should be recognized as similar without requiring semantic embeddings.
Can I remove duplicates from the instruction field itself?
No. As implemented in find_near_duplicates(), the instruction key never triggers automatic deletion. This safety mechanism prevents accidental loss of distinct prompts that may happen to share similar phrasing. Only input and output fields support automatic removal when remove_duplicates=True.
How does this utility handle large datasets?
The implementation uses scikit-learn's optimized TF-IDF vectorizer and cosine similarity functions, which efficiently handle sparse matrices. For very large datasets (100k+ entries), consider processing in batches or increasing available RAM, as the algorithm computes full pairwise similarity matrices that scale O(n²) with dataset size.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →