# How CUPP Removes Duplicates from Generated Wordlists: The Ordered Dictionary Method

> Learn how CUPP efficiently removes duplicate wordlist entries using Python's ordered dictionary method. Discover the technique for streamlined password generation.

- Repository: [Mebus/cupp](https://github.com/Mebus/cupp)
- Tags: internals
- Published: 2026-07-03

---

**CUPP eliminates duplicate password candidates by converting intermediate and final lists to ordered dictionaries using Python's `dict.fromkeys()` method, then extracting the keys back into a list while preserving the original insertion order.**

CUPP (Common User Passwords Profiler), maintained in the Mebus/cupp repository, generates massive candidate password lists by combining names, dates, special characters, and leet-speak transformations. As combinatorial output expands exponentially, the tool removes duplicate entries efficiently to keep wordlists manageable and avoid redundant security testing entries.

## The Ordered Dictionary Pattern for Duplicate Removal

CUPP implements a deterministic, order-preserving deduplication technique using Python's built-in dictionary type. The pattern `list(dict.fromkeys(iterable).keys())` creates a dictionary where keys represent unique items from the iterable, maintaining their first appearance order, then converts the keys back to a list.

This approach appears throughout [`cupp.py`](https://github.com/Mebus/cupp/blob/main/cupp.py) whenever intermediate combination tables or final wordlists require cleanup.

## Duplicate Removal in Interactive Mode (`generate_wordlist_from_profile`)

When running CUPP with the `-i` flag for interactive profiling, the `generate_wordlist_from_profile` function creates multiple combination tables (named `kombi`, `bdss`, `wbdss`, etc.). Each table undergoes individual deduplication before final merging.

According to the source code at lines 441-445, CUPP processes 21 combination tables:

```python
for i in range(1, 22):
    komb_unique[i] = list(dict.fromkeys(kombi[i]).keys())

```

After merging all deduplicated sub-lists into a single `uniqlist`, the function performs a final pass to ensure no duplicates remain across the combined dataset (lines 580-585):

```python
unique_lista = list(dict.fromkeys(uniqlist).keys())

```

This two-stage process—first deduplicating individual pattern groups, then the final concatenated list—ensures memory efficiency while maintaining the intended password priority order.

## Duplicate Removal in Dictionary Improvement Mode (`improve_dictionary`)

For the `-w` option, which improves existing dictionaries, CUPP applies the same technique in the `improve_dictionary` function (lines 60-74). The function first builds several intermediate lists (`kombinacija`) and deduplicates each:

```python
for i in range(6):
    komb_unique[i] = list(dict.fromkeys(kombinacija[i]).keys())
komb_unique[6] = list(dict.fromkeys(listica).keys())
komb_unique[7] = list(dict.fromkeys(cont).keys())

```

After concatenating these into `uniqlist`, the final deduplication occurs at lines 70-74:

```python
unique_lista = list(dict.fromkeys(uniqlist).keys())

```

## Why CUPP Uses Ordered Dictionary Keys

**Order preservation** is critical for CUPP's "hyperspeed" printing and ranking features. The tool prioritizes passwords based on generation order, placing simpler patterns before complex mutations. Using `dict.fromkeys()` preserves this insertion sequence, unlike Python `set` operations which discard ordering.

**Zero dependencies** make this method ideal for CUPP's standalone design. The technique requires only built-in Python types, avoiding external libraries like `pandas` or `numpy` that would increase the attack surface and deployment complexity for penetration testers.

## Summary

- CUPP removes duplicates using `dict.fromkeys()` to create ordered dictionaries, then converts keys back to lists.
- In `generate_wordlist_from_profile` (lines 441-585), the tool deduplicates 21 combination tables individually before a final merge pass.
- In `improve_dictionary` (lines 60-74), intermediate lists are cleaned before final concatenation and deduplication.
- This method preserves the original insertion order of passwords, maintaining priority ranking for security testing.
- The technique requires no external dependencies, relying solely on Python's built-in dictionary implementation.

## Frequently Asked Questions

### Does CUPP remove duplicates case-insensitively?

No, CUPP's duplicate removal is case-sensitive. The `dict.fromkeys()` method treats "Password" and "password" as distinct keys because Python strings are case-sensitive by default. This preserves both variants in the final wordlist, which is beneficial for penetration testing since password authentication systems typically enforce case sensitivity.

### Why doesn't CUPP use Python sets for duplicate removal?

While Python sets provide O(1) lookup for deduplication, they do not preserve insertion order (in versions prior to Python 3.7). CUPP specifically requires order preservation to maintain the priority ranking of generated passwords—from simpler base terms to complex leet-speak variations. The `dict.fromkeys()` method achieves the same uniqueness guarantee while keeping the original sequence intact.

### At what stage does CUPP remove duplicates during wordlist generation?

CUPP performs duplicate removal at two stages: first, it deduplicates individual combination tables immediately after generation (such as `kombi` or `kombinacija` arrays), and second, it performs a final deduplication pass after merging all intermediate lists into the master `uniqlist`. This staged approach optimizes memory usage during the combinatorial explosion of password candidates.

### Can the duplicate removal method in CUPP handle very large wordlists?

Yes, the ordered dictionary method scales efficiently for large datasets because Python's dictionary implementation provides O(n) time complexity for creation while maintaining low memory overhead. However, since CUPP stores the entire wordlist in memory before writing to disk (as seen in the `unique_lista` construction), available RAM remains the limiting factor for massive wordlists rather than the deduplication algorithm itself.