# Handling Near-Duplicate Detection in Instruction Datasets: A Practical Guide

> Efficiently handle near-duplicate detection in instruction datasets with this practical guide. Learn to remove redundant entries using TF-IDF and cosine similarity.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-12

---

**The `rasbt/LLMs-from-scratch` repository provides a production-ready utility in [`ch07/02_dataset-utilities/find-near-duplicates.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/02_dataset-utilities/find-near-duplicates.py) that detects and removes near-duplicate entries in JSON instruction datasets using character-level TF-IDF vectorization and cosine similarity scoring.**

Handling near-duplicate detection in instruction datasets is critical for preventing data leakage and ensuring model generalization during fine-tuning. The utility implements a robust three-stage pipeline that normalizes text variations, converts entries into vector representations, and calculates pairwise similarity scores to identify redundant training examples.

## The Three-Stage Detection Pipeline

The core logic in [`find-near-duplicates.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/find-near-duplicates.py) processes instruction datasets through distinct phases designed to catch semantic duplicates while preserving legitimate variations.

### Text Normalization via preprocess_text

Each text field undergoes standardization through the `preprocess_text` function. This lowercases all characters and strips punctuation, ensuring that superficial differences like "What's" versus "what is" do not prevent duplicate detection. This normalization step is crucial for instruction datasets where minor formatting variations often mask identical underlying queries.

### Character-Level TF-IDF Vectorization

The cleaned strings are transformed into numerical vectors using `TfidfVectorizer(analyzer="char", ngram_range=(1,3))`. Character-level n-grams capture small lexical changes and typos while remaining robust to minor wording differences that word-level tokenizers might miss. This approach is particularly effective for instruction datasets where similar prompts may use synonymous terms or slight paraphrasing.

### Cosine Similarity Thresholding

The pipeline computes a full pairwise cosine similarity matrix using `cosine_similarity` from scikit-learn. Any pair exceeding the user-defined `threshold` (default **0.75**) is flagged as a near duplicate. The system distinguishes between **detection** and **removal**: duplicates found in the *instruction* field are reported but never auto-deleted, while matches in *input* or *output* fields can be marked for removal to prevent repetitive training examples.

## Command-Line and Programmatic Usage

The utility supports both interactive command-line workflows and direct Python module integration.

### CLI Workflow for Dataset Cleaning

The script accepts standard UNIX-style arguments for processing JSON files directly:

```bash
python find-near-duplicates.py \
    --json_file data/instructions.json \
    --threshold 0.80 \
    --remove_duplicates \
    --json_output_file data/instructions_clean.json

```

This invocation loads the dataset, prints near-duplicate pairs for each key (`instruction`, `input`, `output`), removes duplicates from the *input* and *output* fields only, and writes the cleaned list to [`instructions_clean.json`](https://github.com/rasbt/LLMs-from-scratch/blob/main/instructions_clean.json).

### Python API: find_near_duplicates Function

For integration into existing data pipelines, import the detection logic directly:

```python
import json
from find_near_duplicates import find_near_duplicates

# Load dataset

with open("data/instructions.json") as f:
    dataset = json.load(f)

# Detect near-duplicates in instruction field with strict threshold

filtered, dup_pairs = find_near_duplicates(
    dataset, 
    threshold=0.90, 
    key="instruction"
)

print(f"Found {len(dup_pairs)} near-duplicate pairs")
for a, b, sim in dup_pairs:
    print(f"Similarity: {sim:.2f}")
    print(f"  1: {a['instruction']}")
    print(f"  2: {b['instruction']}")

```

This returns a tuple containing the filtered dataset (identical to input when `key="instruction"` since instructions are never removed) and a list of duplicate triples `(item_a, item_b, similarity_score)`.

### Automated Removal with find_print_and_remove_near_duplicates

For comprehensive cleaning across all fields with automatic removal:

```python
from find_near_duplicates import find_print_and_remove_near_duplicates

cleaned = find_print_and_remove_near_duplicates(
    json_data=dataset,
    remove_duplicates=True,
    threshold=0.85
)

```

When `remove_duplicates=True`, the function iterates through all keys in the first JSON object, prints similarity reports for each field, and automatically removes duplicates from *input* and *output* fields while preserving all instruction prompts.

## Key Implementation Details

The utility deliberately separates **detection** from **removal** to prevent over-aggressive pruning. This two-step workflow allows data scientists to inspect similarity scores before committing to deletions. The character-level n-gram approach (range 1-3) provides superior performance for short instruction texts compared to word-level embeddings, capturing typos and minor morphological variations without requiring GPU resources.

According to the source code in [`ch07/02_dataset-utilities/find-near-duplicates.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/02_dataset-utilities/find-near-duplicates.py), the threshold parameter accepts float values between 0 and 1, where higher values increase strictness. The default of 0.75 strikes a balance between catching paraphrased duplicates and preserving legitimately distinct examples.

## Summary

- **[`find-near-duplicates.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/find-near-duplicates.py)** implements a lightweight, CPU-based near-duplicate detection system using **TF-IDF vectorization** and **cosine similarity**.
- **Character-level n-grams** (1-3) capture small textual variations better than word-level approaches for short instruction texts.
- **Selective removal** logic preserves instruction prompts while allowing deduplication of inputs and outputs to prevent training data redundancy.
- Both **CLI** and **Python API** interfaces support integration into MLOps pipelines with configurable similarity thresholds.

## Frequently Asked Questions

### What similarity threshold should I use for instruction datasets?

The default **0.75** threshold works well for most instruction-following datasets. Increase to **0.85-0.90** for stricter deduplication when dealing with templated prompts, or decrease to **0.65-0.70** if you want to catch only near-identical matches while preserving paraphrased variations.

### Why does the utility use character-level instead of word-level n-grams?

Character-level n-grams catch **typos, punctuation differences, and minor spelling variations** that word-level tokenizers miss. This is particularly important for user-generated instruction datasets where "color" and "colour" or "what's" and "whats" should be recognized as similar without requiring semantic embeddings.

### Can I remove duplicates from the instruction field itself?

No. As implemented in `find_near_duplicates()`, the **instruction key never triggers automatic deletion**. This safety mechanism prevents accidental loss of distinct prompts that may happen to share similar phrasing. Only *input* and *output* fields support automatic removal when `remove_duplicates=True`.

### How does this utility handle large datasets?

The implementation uses **scikit-learn's optimized TF-IDF vectorizer and cosine similarity functions**, which efficiently handle sparse matrices. For very large datasets (100k+ entries), consider processing in batches or increasing available RAM, as the algorithm computes full pairwise similarity matrices that scale O(n²) with dataset size.