How CUPP Handles Unicode and UTF-8 Characters in Names for Password Generation

CUPP relies entirely on Python 3's native Unicode support to process UTF-8 characters in names, applying only standard case-folding operations without normalization or encoding conversion.

CUPP (Common User Passwords Profiler) is a Python-based password list generation tool maintained in the Mebus/cupp repository. Unlike tools that strip accents or normalize international characters to ASCII, CUPP preserves UTF-8 input exactly as entered while generating password candidates. Understanding how CUPP handles Unicode is essential for security researchers working with global name datasets and international wordlists.

How CUPP Processes Unicode Input in Names

Lower-Casing at the Input Stage

When CUPP collects the first name via interactive prompt at line 310 of cupp.py, it immediately applies str.lower():

name = input("> First Name: ").lower().strip()  # cupp.py:310

This method is Unicode-aware, correctly handling characters like "Á", "ß", or "π" according to the Unicode standard without requiring explicit UTF-8 decoding.

Title-Casing for Password Variations

Later in the generation pipeline, at line 221 of cupp.py, the same value is transformed using str.title():

nameup = profile["name"].title()  # cupp.py:221

This creates capitalized variants (e.g., "josé" → "José") while preserving the original UTF-8 encoding and character composition.

What CUPP Does Not Do with UTF-8 Characters

No Explicit Encoding or Decoding

Because CUPP runs exclusively on Python 3, all string handling uses Unicode objects directly. The source code never calls .encode() or .decode(), nor does it attempt to transcode UTF-8 to ASCII or other legacy encodings.

No Unicode Normalization

The code does not invoke unicodedata.normalize() to convert characters to NFC or NFD canonical forms. This means composed characters (like "é") and decomposed sequences (like "e" + "´") remain as entered, potentially creating distinct password candidates for visually identical strings.

No Diacritic Stripping

Unlike tools that remove accents to create ASCII-only passwords, CUPP preserves diacritics, umlauts, and non-Latin characters throughout the generation process. There is no transliteration logic in cupp.py to convert "ñ" to "n" or "ß" to "ss".

Impact on Generated Password Candidates

When CUPP generates passwords, UTF-8 characters from names appear in the final wordlist exactly where the original name, surname, or nickname is inserted. They participate in:

  • Direct concatenations (e.g., "José2024")
  • Reversals (e.g., "ésóJ")
  • Numeric appending (e.g., "Łukasz123")

However, leet substitutions defined in cupp.cfg typically target ASCII characters only. Characters like "ß" or "ñ" remain unchanged during leet-speak conversion because the configuration file's mapping does not include Unicode equivalents, limiting 1337-speak transformations to standard ASCII letters.

Summary

  • CUPP uses Python 3's native Unicode str type, eliminating manual encoding conversion
  • Only str.lower() and str.title() transformations are applied, both Unicode-aware
  • No normalization removes diacritics or converts composed/decomposed characters
  • UTF-8 characters propagate directly into generated password candidates
  • Leet-mode substitutions in cupp.cfg affect ASCII characters only

Frequently Asked Questions

Does CUPP convert Unicode names to ASCII for password generation?

No. According to the source code in cupp.py, the tool never strips diacritics or transliterates characters. Names containing "İstanbul" or "Łukasz" remain in their original UTF-8 form throughout the generation pipeline, creating passwords with international characters intact.

What happens to special characters when using CUPP's leet mode?

Leet substitutions only apply to ASCII characters defined in cupp.cfg. Unicode characters like "á" or "ß" pass through unchanged because the configuration file does not map them to 1337-speak equivalents, leaving them as literal UTF-8 in the output.

Can CUPP handle mixed-encoding inputs or Python 2?

CUPP is written for Python 3, which handles all strings as Unicode by default. The code does not support Python 2, and there are no explicit encoding conversions to handle mixed-encoding inputs—UTF-8 is expected and preserved natively through Python's internal string representation.

Why might the same name generate different passwords in CUPP?

Since CUPP does not normalize Unicode to NFC or NFD forms at line 310 or 221, visually identical names using different Unicode compositions (e.g., precomposed "é" vs. decomposed "e" + "´") generate distinct password candidates because the byte sequences differ, even when rendered identically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →