How to Improve Existing Dictionaries with the CUPP -w Option

The -w option in CUPP activates dictionary-improvement mode that reads an existing wordlist, applies configurable transformations like concatenation, special-character suffixes, and leet-speak conversions, and outputs an expanded, de-duplicated password list as <filename>.cupp.txt.

The Common User Passwords Profiler (CUPP) provides a specialized workflow for enhancing existing wordlists through its -w command-line flag. When you use the CUPP -w option, the tool transforms static dictionaries into combinatorial engines capable of generating thousands of password variations based on interactive user choices. This feature leverages the improve_dictionary() function in cupp.py to apply pattern-based mutations commonly found in real-world passwords.

How the CUPP -w Option Works

Argument Parsing and Mode Activation

In cupp.py (lines 58-64), the argument parser maps -w <FILENAME> to the improve attribute. When invoked, the CLI immediately calls improve_dictionary() with the supplied filename, shifting the tool from profiling mode to dictionary-enhancement mode. This architectural separation ensures that the combinatorial logic only executes when explicitly requested, preventing accidental mutation of existing wordlists.

Loading and Processing the Source Wordlist

The improve_dictionary() function (lines 84-96) validates file existence and reads every word into a Python list named listica, splitting entries on whitespace. This initial load creates the foundational dataset that will undergo multiple transformation rounds. The function preserves the original entries throughout the process, ensuring the final output contains both the base words and their generated variants.

Interactive Enrichment Options

Between lines 100-142, CUPP presents four augmentation choices via interactive prompts:

  • Concatenation (cont): Combines every pair of words from the source list
  • Special-character suffixes (spechars): Appends 1-3 user-defined symbols to each entry
  • Random numbers (randnum): Adds numeric sequences ranging from numfrom to numto
  • Leet-mode (leetmode): Transforms characters into their leet-speak equivalents (e.g., a → 4, e → 3)

Each option toggles specific branches in the generation pipeline, allowing you to customize the expansion strategy based on your target's password complexity.

Combinatorial Generation and De-duplication

When you enable multiple options, CUPP utilizes helper generators komb (for string-to-string combinations with optional separators) and concats (for string-plus-numeric ranges) to build six distinct base pools stored in kombinacija[0] through kombinacija[5] (lines 142-166).

The tool applies rigorous de-duplication at two stages:

  1. Pool-level de-duplication (lines 190-199): Uses dict.fromkeys() to eliminate duplicates within each combination pool (komb_unique)
  2. Global de-duplication (lines 202-215): Merges all unique pools with the original listica into uniqlist, then applies another dict.fromkeys() pass to create unique_lista

If leet-mode is enabled, the code iterates through unique_lista (lines 224-233), passing each entry through make_leet() and appending results to unique_leet before the final merge.

Length Filtering and Final Output

Before writing the output, CUPP filters passwords by length using configuration values wcfrom and wcto (lines 236-242). This step removes entries that fall outside your target parameters, preventing bloated files with obviously invalid passwords. The surviving entries are written to <FILENAME>.cupp.txt via the print_to_file function (lines 244-247), with an optional "hyperspeed print" feature for terminal preview.

Configuration and Customization via cupp.cfg

The cupp.cfg file controls the behavior of the CUPP -w option without requiring code modifications. Key parameters include:

  • Numeric ranges: Define numfrom and numto to specify which years or sequences append to words
  • Character sets: Customize spechars to include symbols specific to your target's keyboard layout
  • Length thresholds: Adjust wcfrom and wcto to match minimum and maximum password policies

Editing these values before running python cupp.py -w allows you to tune the dictionary size and complexity for specific penetration-testing scenarios.

Practical Usage Examples

Run the dictionary improvement mode on an existing wordlist:

python cupp.py -w rockyou.txt

During execution, CUPP will prompt you for each augmentation type:


> Do you want to concatenate all words from wordlist? Y/[N]: y
> Do you want to add special chars at the end of words? Y/[N]: y
> Do you want to add some random numbers at the end of words? Y/[N]: y
> Leet mode? (i.e. leet = 1337) Y/[N]: n

This sequence produces rockyou.txt.cupp.txt containing entries like:

  • password123 (original + numbers)
  • admin!@# (original + special characters)
  • helloWorld2021 (concatenated pair + year)

Combine the CUPP -w option with custom configuration settings:


# Edit cupp.cfg to change numfrom=1990 and numto=2024 before running

python cupp.py -w mylist.txt

Pipe the output directly into other security tools for immediate processing:

python cupp.py -w small.txt | grep -i "admin"

Summary

  • The CUPP -w option triggers improve_dictionary() in cupp.py to transform static wordlists into combinatorial password databases.
  • The feature applies four distinct augmentations: word concatenation, special-character suffixes, numeric ranges, and leet-speak conversion.
  • De-duplication occurs twice—first at the pool level using dict.fromkeys(), then globally across all combinations—ensuring no duplicate entries in the final output.
  • Length filtering via wcfrom and wcto parameters keeps output files focused on viable password lengths.
  • Output is saved as <FILENAME>.cupp.txt in the working directory, preserving the original file while providing an expanded variant.

Frequently Asked Questions

What does the -w option do in CUPP?

The -w option activates dictionary-improvement mode, which reads an existing wordlist and applies interactive transformations to generate password variations. According to the Mebus/cupp source code, this mode calls improve_dictionary() to concatenate words, append special characters and numbers, and optionally convert text to leet-speak, producing an enhanced wordlist saved as <filename>.cupp.txt.

How does CUPP prevent duplicate passwords when using -w?

CUPP implements a two-stage de-duplication process using Python's dict.fromkeys() method. First, it removes duplicates within each combination pool (lines 190-199 in cupp.py), then it de-duplicates the final merged list containing all original words and generated variants (lines 202-215). This ensures that even when multiple transformations create identical strings, only unique entries appear in the output.

Can I customize the numeric ranges appended by the -w option?

Yes, you can customize the numeric ranges by editing the cupp.cfg configuration file before running the tool. The parameters numfrom and numto define the starting and ending values for number suffixes, allowing you to target specific years, PIN ranges, or sequential patterns relevant to your penetration test.

Where does CUPP save the improved dictionary?

CUPP saves the enhanced dictionary in the current working directory using the naming convention <original_filename>.cupp.txt. For example, running python cupp.py -w rockyou.txt creates rockyou.txt.cupp.txt containing the original entries plus all generated combinations that passed the length and de-duplication filters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →