How CUPP Handles Unicode and UTF-8 Characters in Names for Password Generation
CUPP relies entirely on Python 3's native Unicode support to process UTF-8 characters in names, applying only standard case-folding operations without normalization or encoding conversion.
CUPP (Common User Passwords Profiler) is a Python-based password list generation tool maintained in the Mebus/cupp repository. Unlike tools that strip accents or normalize international characters to ASCII, CUPP preserves UTF-8 input exactly as entered while generating password candidates. Understanding how CUPP handles Unicode is essential for security researchers working with global name datasets and international wordlists.
How CUPP Processes Unicode Input in Names
Lower-Casing at the Input Stage
When CUPP collects the first name via interactive prompt at line 310 of cupp.py, it immediately applies str.lower():
name = input("> First Name: ").lower().strip() # cupp.py:310
This method is Unicode-aware, correctly handling characters like "Á", "ß", or "π" according to the Unicode standard without requiring explicit UTF-8 decoding.
Title-Casing for Password Variations
Later in the generation pipeline, at line 221 of cupp.py, the same value is transformed using str.title():
nameup = profile["name"].title() # cupp.py:221
This creates capitalized variants (e.g., "josé" → "José") while preserving the original UTF-8 encoding and character composition.
What CUPP Does Not Do with UTF-8 Characters
No Explicit Encoding or Decoding
Because CUPP runs exclusively on Python 3, all string handling uses Unicode objects directly. The source code never calls .encode() or .decode(), nor does it attempt to transcode UTF-8 to ASCII or other legacy encodings.
No Unicode Normalization
The code does not invoke unicodedata.normalize() to convert characters to NFC or NFD canonical forms. This means composed characters (like "é") and decomposed sequences (like "e" + "´") remain as entered, potentially creating distinct password candidates for visually identical strings.
No Diacritic Stripping
Unlike tools that remove accents to create ASCII-only passwords, CUPP preserves diacritics, umlauts, and non-Latin characters throughout the generation process. There is no transliteration logic in cupp.py to convert "ñ" to "n" or "ß" to "ss".
Impact on Generated Password Candidates
When CUPP generates passwords, UTF-8 characters from names appear in the final wordlist exactly where the original name, surname, or nickname is inserted. They participate in:
- Direct concatenations (e.g., "José2024")
- Reversals (e.g., "ésóJ")
- Numeric appending (e.g., "Łukasz123")
However, leet substitutions defined in cupp.cfg typically target ASCII characters only. Characters like "ß" or "ñ" remain unchanged during leet-speak conversion because the configuration file's mapping does not include Unicode equivalents, limiting 1337-speak transformations to standard ASCII letters.
Summary
- CUPP uses Python 3's native Unicode
strtype, eliminating manual encoding conversion - Only
str.lower()andstr.title()transformations are applied, both Unicode-aware - No normalization removes diacritics or converts composed/decomposed characters
- UTF-8 characters propagate directly into generated password candidates
- Leet-mode substitutions in
cupp.cfgaffect ASCII characters only
Frequently Asked Questions
Does CUPP convert Unicode names to ASCII for password generation?
No. According to the source code in cupp.py, the tool never strips diacritics or transliterates characters. Names containing "İstanbul" or "Łukasz" remain in their original UTF-8 form throughout the generation pipeline, creating passwords with international characters intact.
What happens to special characters when using CUPP's leet mode?
Leet substitutions only apply to ASCII characters defined in cupp.cfg. Unicode characters like "á" or "ß" pass through unchanged because the configuration file does not map them to 1337-speak equivalents, leaving them as literal UTF-8 in the output.
Can CUPP handle mixed-encoding inputs or Python 2?
CUPP is written for Python 3, which handles all strings as Unicode by default. The code does not support Python 2, and there are no explicit encoding conversions to handle mixed-encoding inputs—UTF-8 is expected and preserved natively through Python's internal string representation.
Why might the same name generate different passwords in CUPP?
Since CUPP does not normalize Unicode to NFC or NFD forms at line 310 or 221, visually identical names using different Unicode compositions (e.g., precomposed "é" vs. decomposed "e" + "´") generate distinct password candidates because the byte sequences differ, even when rendered identically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →