How the CUPP -w Option Concatenates Words from Existing Wordlists
The -w option triggers CUPP's improve_dictionary workflow, which loads an existing wordlist and optionally generates every ordered pair of different words through a nested loop, concatenating them (cont1 + cont2) to expand the dictionary before applying standard mutations like years and special characters.
The Common User Passwords Profiler (CUPP) is a targeted wordlist generator that creates personalized password lists based on profiling data. When you need to enhance an existing dictionary rather than generate one from scratch, the -w option provides a powerful concatenation mechanism that doubles the semantic space by combining words from your source list.
How the -w Option Loads and Processes Wordlists
When you invoke python cupp.py -w mylist.txt, CUPP enters the improve-dictionary mode defined in cupp.py. The tool first ingests your wordlist file and stores all entries in memory for processing.
Loading the Wordlist into Memory
At lines 191-197 in cupp.py, CUPP reads the supplied file line-by-line and splits each line into individual words, populating the listica array:
# Inside improve_dictionary function
listica = []
for line in f:
for word in line.split():
if word not in listica:
listica.append(word)
This deduplication ensures that each unique word appears only once in the processing queue, preventing redundant combinations later.
The Concatenation Prompt
After loading completes, CUPP prompts the user at lines 204-210:
> Do you want to concatenate all words from wordlist? Y/[N]:
If you answer n, CUPP processes only the original words. If you answer y, the tool executes the concatenation algorithm before proceeding to standard mutations.
The Double-Loop Concatenation Algorithm
The core concatenation logic resides at lines 220-224 in cupp.py. When enabled, CUPP initializes a cont list with an empty string and performs a Cartesian product of distinct words:
cont = [""]
for cont1 in listica:
for cont2 in listica:
if listica.index(cont1) != listica.index(cont2):
cont.append(cont1 + cont2)
This implementation exhibits three key characteristics:
- Ordered pairs: The nested loops generate every possible combination where
cont1precedescont2 - Distinct indices: The
indexcheck ensurescont1andcont2are different words, preventing same-word concatenation likepasswordpassword - Empty string seed: The initial empty string allows the combination engine to later work with "no extra word" as a valid option
Integration with CUPP's Combination Engine
After populating cont, CUPP feeds this expanded list into its standard mutation pipeline. At lines 247-250, the concatenated words combine with years, special characters, and other modifiers:
# From the improve_dictionary function
kombination[1] = list(komb(cont, years))
kombination[2] = list(komb(cont, ['', '!', '@', '#', '$', '%']))
This means a wordlist containing ["admin", "root"] becomes ["adminroot", "rootadmin"] in the cont array, and each of these then receives suffixes like 2024, !, and @ through the normal combination machinery.
Configuration Thresholds and Limits
CUPP respects a safety threshold defined in cupp.cfg to prevent memory exhaustion with large wordlists. At lines 108-112 in cupp.py, the code checks CONFIG["global"]["threshold"]:
if len(listica) > int(CONFIG["global"]["threshold"]):
print("> Maximum number of words for concatenation is " + CONFIG["global"]["threshold"])
print("> Check configuration file for increasing this number.")
sys.exit()
If your wordlist exceeds this limit (default typically 5000-10000 words depending on version), CUPP warns you and exits to prevent the exponential explosion of combinations that would occur during the double-loop operation.
Practical Usage Examples
Basic Wordlist Enhancement
Enable concatenation to generate compound passwords from common base words:
python cupp.py -w common.txt
# > Do you want to concatenate all words from wordlist? Y/[N]: y
# Generates: adminroot, rootadmin, adminuser, useradmin, etc.
Processing Without Concatenation
Apply only standard mutations (years, special chars) to original words:
python cupp.py -w common.txt
# > Do you want to concatenate all words from wordlist? Y/[N]: n
# Processes only: admin, root, user (with mutations)
Handling Large Wordlists
When exceeding the configured threshold, CUPP enforces limits:
python cupp.py -w huge.txt
# > Do you want to concatenate all words from wordlist? Y/[N]: y
# [-] Maximum number of words for concatenation is 5000
# > Do you want to concatenate all words from wordlist? Y/[N]: n
Summary
- The
-woption activates theimprove_dictionaryfunction incupp.pyto enhance existing wordlists - Words load into the
listicaarray at lines 191-197, with deduplication applied during ingestion - A nested double-loop at lines 220-224 generates all ordered pairs of distinct words, storing results in the
contlist - The concatenated results feed into standard CUPP mutations (years, special characters) through the
kombfunction - The
thresholdsetting incupp.cfgprevents memory exhaustion by limiting the maximum wordlist size for concatenation operations
Frequently Asked Questions
Does CUPP concatenate a word with itself when using the -w option?
No. The code explicitly checks if listica.index(cont1) != listica.index(cont2) before concatenating, ensuring that passwordpassword or similar self-repeating patterns are excluded from the generated combinations.
How does the -w option affect memory usage with large wordlists?
The double-loop algorithm creates n × (n-1) new strings (where n is your wordlist length), which can exhaust RAM quickly. CUPP prevents this by checking against CONFIG["global"]["threshold"] at lines 108-112 in cupp.py and warning users if their list exceeds the safe limit defined in cupp.cfg.
Can I use the -w option with the interactive profiling mode?
No. The -w option triggers improve_dictionary as a standalone workflow in cupp.py. It operates exclusively on existing wordlist files and cannot combine with the interactive questionnaire mode that generates profiles from personal data like birthdates or pet names.
What happens if I answer "no" to the concatenation prompt?
If you decline concatenation at the prompt (lines 204-210), CUPP skips the double-loop entirely and proceeds directly to standard mutations. The original listica words still receive years, special characters, and leetspeak transformations, but no cross-word combinations are generated.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →