How CUPP Removes Duplicates from Generated Wordlists: The Ordered Dictionary Method

CUPP eliminates duplicate password candidates by converting intermediate and final lists to ordered dictionaries using Python's dict.fromkeys() method, then extracting the keys back into a list while preserving the original insertion order.

CUPP (Common User Passwords Profiler), maintained in the Mebus/cupp repository, generates massive candidate password lists by combining names, dates, special characters, and leet-speak transformations. As combinatorial output expands exponentially, the tool removes duplicate entries efficiently to keep wordlists manageable and avoid redundant security testing entries.

The Ordered Dictionary Pattern for Duplicate Removal

CUPP implements a deterministic, order-preserving deduplication technique using Python's built-in dictionary type. The pattern list(dict.fromkeys(iterable).keys()) creates a dictionary where keys represent unique items from the iterable, maintaining their first appearance order, then converts the keys back to a list.

This approach appears throughout cupp.py whenever intermediate combination tables or final wordlists require cleanup.

Duplicate Removal in Interactive Mode (generate_wordlist_from_profile)

When running CUPP with the -i flag for interactive profiling, the generate_wordlist_from_profile function creates multiple combination tables (named kombi, bdss, wbdss, etc.). Each table undergoes individual deduplication before final merging.

According to the source code at lines 441-445, CUPP processes 21 combination tables:

for i in range(1, 22):
    komb_unique[i] = list(dict.fromkeys(kombi[i]).keys())

After merging all deduplicated sub-lists into a single uniqlist, the function performs a final pass to ensure no duplicates remain across the combined dataset (lines 580-585):

unique_lista = list(dict.fromkeys(uniqlist).keys())

This two-stage process—first deduplicating individual pattern groups, then the final concatenated list—ensures memory efficiency while maintaining the intended password priority order.

Duplicate Removal in Dictionary Improvement Mode (improve_dictionary)

For the -w option, which improves existing dictionaries, CUPP applies the same technique in the improve_dictionary function (lines 60-74). The function first builds several intermediate lists (kombinacija) and deduplicates each:

for i in range(6):
    komb_unique[i] = list(dict.fromkeys(kombinacija[i]).keys())
komb_unique[6] = list(dict.fromkeys(listica).keys())
komb_unique[7] = list(dict.fromkeys(cont).keys())

After concatenating these into uniqlist, the final deduplication occurs at lines 70-74:

unique_lista = list(dict.fromkeys(uniqlist).keys())

Why CUPP Uses Ordered Dictionary Keys

Order preservation is critical for CUPP's "hyperspeed" printing and ranking features. The tool prioritizes passwords based on generation order, placing simpler patterns before complex mutations. Using dict.fromkeys() preserves this insertion sequence, unlike Python set operations which discard ordering.

Zero dependencies make this method ideal for CUPP's standalone design. The technique requires only built-in Python types, avoiding external libraries like pandas or numpy that would increase the attack surface and deployment complexity for penetration testers.

Summary

  • CUPP removes duplicates using dict.fromkeys() to create ordered dictionaries, then converts keys back to lists.
  • In generate_wordlist_from_profile (lines 441-585), the tool deduplicates 21 combination tables individually before a final merge pass.
  • In improve_dictionary (lines 60-74), intermediate lists are cleaned before final concatenation and deduplication.
  • This method preserves the original insertion order of passwords, maintaining priority ranking for security testing.
  • The technique requires no external dependencies, relying solely on Python's built-in dictionary implementation.

Frequently Asked Questions

Does CUPP remove duplicates case-insensitively?

No, CUPP's duplicate removal is case-sensitive. The dict.fromkeys() method treats "Password" and "password" as distinct keys because Python strings are case-sensitive by default. This preserves both variants in the final wordlist, which is beneficial for penetration testing since password authentication systems typically enforce case sensitivity.

Why doesn't CUPP use Python sets for duplicate removal?

While Python sets provide O(1) lookup for deduplication, they do not preserve insertion order (in versions prior to Python 3.7). CUPP specifically requires order preservation to maintain the priority ranking of generated passwords—from simpler base terms to complex leet-speak variations. The dict.fromkeys() method achieves the same uniqueness guarantee while keeping the original sequence intact.

At what stage does CUPP remove duplicates during wordlist generation?

CUPP performs duplicate removal at two stages: first, it deduplicates individual combination tables immediately after generation (such as kombi or kombinacija arrays), and second, it performs a final deduplication pass after merging all intermediate lists into the master uniqlist. This staged approach optimizes memory usage during the combinatorial explosion of password candidates.

Can the duplicate removal method in CUPP handle very large wordlists?

Yes, the ordered dictionary method scales efficiently for large datasets because Python's dictionary implementation provides O(n) time complexity for creation while maintaining low memory overhead. However, since CUPP stores the entire wordlist in memory before writing to disk (as seen in the unique_lista construction), available RAM remains the limiting factor for massive wordlists rather than the deduplication algorithm itself.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →