How CUPP Generates Password Combinations from User Profile Data
CUPP generates passwords by collecting personal data through an interactive questionnaire, normalizing that data into tokens (dates, names, keywords), and then executing a combinatorial engine that produces Cartesian products of these tokens with numeric ranges and special characters.
The Common User Passwords Profiler (CUPP) is a penetration testing tool authored in Python that transforms victim profiling data into targeted wordlists. According to the Mebus/cupp source code, the generation process follows a deterministic pipeline that expands a modest set of personal attributes into thousands of candidate passwords through systematic mutation and combination.
Phase 1: Interactive Profile Collection
When the analyst invokes CUPP with the -i or --interactive flag, the interactive() function in cupp.py (lines 308-340) prompts for specific personal identifiers. This function constructs a Python dictionary called profile that stores:
- Identity data: First name, surname, and nicknames
- Temporal data: Birthdates for the victim, partner, and children
- Associative data: Pet names, company names, and custom keywords
- Configuration flags: Boolean-like strings indicating whether to append special characters (
spechars1), random numbers (randnum), or apply leet-speak conversion (leetmode)
These raw strings serve as the foundational input for the combinatorial engine.
Phase 2: Token Preprocessing and Normalization
The generate_wordlist_from_profile(profile) function receives the profile dictionary and immediately expands each attribute into multiple token variants. This preprocessing occurs before any password combinations are generated.
Date Fragmentation
Birthdates undergo aggressive decomposition in cupp.py (lines 95-112). Each date string gets sliced into every plausible numeric fragment: yy, yyy, yyyy, d, m, dd, and mm. These fragments are then combined in 2- and 3-element permutations. For a birthdate of 15081990, this produces tokens like 15, 90, 1990, 08, 15, and compound sequences like 1590 or 081990.
Case Variations and Reversal
For every textual element—names, nicknames, partner names, and child names—CUPP creates three variants (lines 219-229 and 236-244):
- Original lower-case (e.g.,
john) - Title-cased (e.g.,
John) - Reversed (e.g.,
nhoj)
Additional user-supplied keywords from the words field are also added in both lower-case and title-case forms (lines 330-337).
Special Character Injection
If the analyst enabled special characters (profile["spechars1"] == "y"), CUPP loads the character set defined in CONFIG["global"]["chars"] from cupp.cfg. The engine then generates 1-, 2-, and 3-character permutations of these special characters (lines 81-88) to serve as suffixes in later combination stages.
Phase 3: The Combination Engine
With the token pool prepared, CUPP employs two helper generators to construct the actual password candidates: komb() and concats().
Cartesian Product Generation
The komb(seq, start, special) function (lines 110-114) creates Cartesian products between two sequences. It accepts an optional separator string (typically "_" or empty) and yields every possible combination of elements from seq and start. This allows the engine to combine base names with date fragments (e.g., john + 1990 producing john1990 and john_1990).
Numeric Suffix Appending
The concats(seq, start, stop) function (lines 103-107) appends numeric ranges to each element in a sequence. It iterates from numfrom to numto (defined in cupp.cfg) and concatenates these integers directly onto base tokens, generating candidates like john2022 or john1.
The main generation logic (lines 371-424) constructs a dictionary called kombi that systematically mixes:
- Base tokens with victim birthdate fragments
- Base tokens with partner and child birthdate fragments
- Base tokens with years from the configuration
- All variations with special-character suffixes
- Random numeric suffixes if
profile["randnum"] == "y"
After each subset is generated, duplicates are eliminated using dict.fromkeys() (lines 558-582) before being merged into a master uniqlist.
Leet-Speak Conversion and Output Filtering
If leet-mode is enabled (profile["leetmode"] == "y"), every candidate in the final list passes through make_leet() (lines 95-100). This function substitutes characters according to the mapping defined in cupp.cfg (e.g., a becomes 4, e becomes 3).
Finally, the print_to_file() function (lines 693-704) filters the candidates against configurable length boundaries (wcfrom and wcto) and writes the results to a text file named <victim-name>.txt.
Complete Code Example
The following example demonstrates the profile structure required to trigger the full generation pipeline:
# Profile data as constructed by interactive() in cupp.py
profile = {
"name": "john",
"surname": "doe",
"nick": "jdo",
"birthdate": "15081990",
"wife": "mary",
"wifeb": "20021985",
"kid": "lisa",
"kidb": "08052010",
"pet": "fluffy",
"company": "acme",
"words": ["hacker", "juice"],
"spechars1": "y", # Enable special character suffixes
"randnum": "y", # Enable random numeric suffixes
"leetmode": "y", # Enable 1337-speak conversion
}
# Generate wordlist (creates john.txt)
generate_wordlist_from_profile(profile)
Summary
- Profile ingestion: The
interactive()function incupp.py(lines 308-340) collects victim data into a structured dictionary. - Token expansion: Dates are fragmented into substrings, names are reversed and case-varied, and special characters are permuted (lines 81-112, 219-244).
- Combinatorial generation: The
komb()andconcats()functions (lines 103-114) produce Cartesian products of tokens with dates and numeric ranges. - Deduplication and filtering: The engine uses
dict.fromkeys()(lines 558-582) to remove duplicates, applies optional leet-speak viamake_leet(), and filters by length before outputting to file viaprint_to_file()(lines 693-704).
Frequently Asked Questions
What configuration file controls CUPP's generation parameters?
The cupp.cfg file defines the CONFIG["global"] dictionary that specifies special characters (chars), numeric ranges (numfrom, numto), year ranges (wcfrom, wcto), and the leet-speak character mapping. The main script cupp.py reads these values during the preprocessing and output phases.
How does CUPP handle duplicate passwords?
After generating each subset of combinations, CUPP removes duplicates using dict.fromkeys() in lines 558-582 of cupp.py. This technique preserves insertion order while eliminating identical strings before merging results into the final uniqlist.
Can CUPP generate passwords without the interactive mode?
While the primary workflow uses interactive() to build the profile dictionary, you can programmatically construct the profile dictionary (as shown in the code example above) and pass it directly to generate_wordlist_from_profile(). This allows for automated batch processing without manual input prompts.
What date formats does CUPP extract from birthdates?
The extraction logic in lines 95-112 of cupp.py splits dates into eight distinct fragments: single-digit day (d), single-digit month (m), two-digit day (dd), two-digit month (mm), two-digit year (yy), three-digit year (yyy), and four-digit year (yyyy). It then generates 2- and 3-element permutations of these fragments to create date-based password components.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →