CUPP's Combination Generation Functions: Architecture and Implementation

CUPP generates password candidates by chaining two core generator functions—concats for numeric suffixes and komb for string concatenations—within modular pipelines that deduplicate and optionally apply leet transformations.

CUPP (Common User Passwords Profiler) is an open-source password list generation tool that constructs wordlists through combinatorial expansion. The architecture of CUPP's combination generation functions centers on reusable generator-based combinators defined in cupp.py, which orchestrate word fragments, dates, and special characters into comprehensive password candidates.

Core Combinator Functions in cupp.py

CUPP's combination engine revolves around three primary functions located in the main source file. These generators yield strings lazily, keeping memory usage minimal during large-scale wordlist generation.

concats(): Numeric Suffix Generation

The concats(seq, start, stop) function generates every possible combination of a string from seq suffixed with numbers in the range start to stop-1. This powers classic password patterns like "john123" or "alice2024".

In cupp.py lines 102-107, the implementation iterates through the input sequence and yields each element concatenated with every number in the specified range:

def concats(seq, start, stop):
    for word in seq:
        for i in range(start, stop):
            yield word + str(i)

komb(): String Concatenation Engine

The komb(seq, start, special="") function serves as the primary "word-word" combinator. It produces every concatenation of elements from seq with elements from start, optionally inserting a separator character. According to the source at lines 110-114, this handles combinations like "john_smith" or "mary2023":

def komb(seq, start, special=""):
    for word in seq:
        for i in start:
            yield word + special + str(i)

make_leet(): Leet-Speak Transformation

After generating base combinations, CUPP optionally transforms text into leet-speak using make_leet(x). Defined in lines 95-100, this function maps characters through the CONFIG["LEET"] dictionary (loaded from cupp.cfg), converting "a" to "4", "e" to "3", etc.

Wordlist Generation Pipelines

CUPP implements two distinct pipelines that reuse the same core combinators, ensuring consistent generation rules across the tool.

The improve_dictionary Pipeline (-w mode)

When users supply an existing wordlist via the -w option, the improve_dictionary function orchestrates the combination flow through six distinct stages:

  1. Load the wordlist into a list called listica
  2. Optional concatenation of every pair (cont) using nested loops
  3. Optional special-character suffixes (spechars) building 1-3 character combinations
  4. Optional numeric suffixes (randnum) using concats

The pipeline populates a kombinacija array where each index represents a specific combination type:

kombinacija[0] = list(komb(listica, years))                     # word + year

kombinacija[1] = list(komb(cont, years))                       # concatenated words + year

kombinacija[2] = list(komb(listica, spechars))                # word + special chars

kombinacija[3] = list(komb(cont, spechars))                   # concatenated + specials

kombinacija[4] = list(concats(listica, numfrom, numto))       # word + numbers

kombinacija[5] = list(concats(cont, numfrom, numto))          # concatenated + numbers

All resulting sub-lists are deduplicated using dict.fromkeys, merged together, filtered by length constraints, and written to disk.

The generate_wordlist_from_profile Pipeline (-i mode)

The interactive mode (-i) expands user profile data into combinatorial groups through the generate_wordlist_from_profile function. This pipeline:

  • Extracts date fragments (multiple slices of birthdates: yy, yyy, yyyy, d, m)
  • Generates reverse strings for each name field
  • Stores fragments in specialized lists: bdss, wbdss, kbdss, and reverse

The core combinations are built into a kombi array using the same komb and concats functions:

kombi[1] = list(komb(kombinaa, bdss))               # name combos + birthdates

kombi[2] = list(komb(kombinaaw, wbdss))             # wife-based combos + dates

kombi[3] = list(komb(kombinaak, kbdss))             # kid-based combos + dates

kombi[4] = list(komb(kombinaa, years))              # name combos + years

kombi[5] = list(komb(kombinaac, years))             # pet/company combos + years

kombi[12] = list(concats(kombinaa, numfrom, numto)) # numeric suffixes

kombi[17] = list(komb(reverse, years))              # reverse names + years

Each kombi[n] entry undergoes deduplication via komb_unique[i] = list(dict.fromkeys(kombi[i]).keys()) before concatenation into a master uniqlist. If leet mode is enabled, every entry passes through make_leet as a final transformation.

Architectural Design Patterns

CUPP's combination generation architecture exhibits several sophisticated design patterns that maximize efficiency and maintainability:

  • Generator-style design: Both concats and komb are Python generators that yield strings only when iterated, maintaining low memory footprints even with massive combinatorial spaces.

  • Config-driven parameters: All ranges (years, numfrom, numto) and special-character sets originate from cupp.cfg, allowing users to tune the combinatorial space without modifying source code.

  • Modular combinators: By separating numeric (concats) and textual (komb) logic, the same functions power both the "improve" and "interactive" generation paths, guaranteeing consistent behavior across the tool.

  • Order-preserving deduplication: After each combinatorial step, CUPP removes duplicates using dict.fromkeys(), which preserves insertion order while guaranteeing uniqueness in the final output.

Working with CUPP's Combinators

Generating Numeric Suffixes

To create passwords with numeric suffixes like "alice0" through "alice9":

from cupp import concats

for pw in concats(['alice'], 0, 10):
    print(pw)

# Output: alice0, alice1, ..., alice9

Combining Words with Separators

To generate underscore-separated combinations like "john_smith":

from cupp import komb

first = ['john']
last = ['smith']

for pw in komb(first, last, special='_'):
    print(pw)

# Output: john_smith

Full Pipeline Example

Combining names, years, and numbers in a nested pattern:

from cupp import komb, concats

names = ['bob']
years = ['2024']

# Generate base combinations: bob2024

for base in komb(names, years):
    # Append numeric suffixes: bob20240, bob20241, etc.

    for pw in concats([base], 0, 5):
        print(pw)

Summary

  • CUPP's combination generation functions rely on two core generators: concats() for numeric suffixes and komb() for string concatenation, both defined in cupp.py.
  • The improve_dictionary pipeline (-w mode) enhances existing wordlists by appending years, special characters, and numbers to base words and their concatenations.
  • The generate_wordlist_from_profile pipeline (-i mode) builds comprehensive lists from personal data using date fragments, reversed names, and the same core combinators.
  • Deduplication occurs via dict.fromkeys() after each combinatorial stage, preserving order while ensuring uniqueness.
  • Configuration values from cupp.cfg drive all ranges and special characters, making the tool adjustable without code changes.
  • Optional leet-speak transformation via make_leet() applies as a final pass over the generated candidate list.

Frequently Asked Questions

How does CUPP handle memory efficiency with large wordlists?

CUPP uses Python generator functions for its core combinators (concats and komb), yielding one string at a time rather than building complete lists in memory. While the pipelines eventually materialize lists for deduplication purposes, the initial generation streams values lazily, preventing excessive memory consumption during the combinatorial expansion phase.

What is the difference between the concats and komb functions?

concats(seq, start, stop) specifically handles numeric suffixes, iterating through a range of integers and appending them to base strings. komb(seq, start, special="") is more generic, concatenating elements from two sequences with an optional separator string (like "_" or ""), making it suitable for combining words, dates, or any string elements.

How does CUPP prevent duplicate passwords in the output?

After each combinatorial stage, CUPP applies dict.fromkeys() to remove duplicates while preserving the original insertion order. This occurs in both pipelines: the improve_dictionary function deduplicates the kombinacija array, while generate_wordlist_from_profile creates komb_unique lists before merging them into the final uniqlist.

Can I customize the special characters used in password combinations?

Yes. CUPP loads all special character sets, year ranges, and leet-speak mappings from the cupp.cfg configuration file. By modifying this file, you can adjust the spechars variable (which defaults to combinations of special characters like "!", "@", "#") or change the years list without touching the Python source code in cupp.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →