# How the CUPP -w Option Concatenates Words from Existing Wordlists

> Learn how the CUPP -w option concatenates words from existing wordlists by generating ordered pairs to expand your dictionary. Enhance your password cracking capabilities efficiently.

- Repository: [Mebus/cupp](https://github.com/Mebus/cupp)
- Tags: how-to-guide
- Published: 2026-07-03

---

**The `-w` option triggers CUPP's `improve_dictionary` workflow, which loads an existing wordlist and optionally generates every ordered pair of different words through a nested loop, concatenating them (`cont1 + cont2`) to expand the dictionary before applying standard mutations like years and special characters.**

The Common User Passwords Profiler (CUPP) is a targeted wordlist generator that creates personalized password lists based on profiling data. When you need to enhance an existing dictionary rather than generate one from scratch, the `-w` option provides a powerful concatenation mechanism that doubles the semantic space by combining words from your source list.

## How the -w Option Loads and Processes Wordlists

When you invoke `python cupp.py -w mylist.txt`, CUPP enters the *improve-dictionary* mode defined in [`cupp.py`](https://github.com/Mebus/cupp/blob/main/cupp.py). The tool first ingests your wordlist file and stores all entries in memory for processing.

### Loading the Wordlist into Memory

At lines 191-197 in [`cupp.py`](https://github.com/Mebus/cupp/blob/main/cupp.py), CUPP reads the supplied file line-by-line and splits each line into individual words, populating the `listica` array:

```python

# Inside improve_dictionary function

listica = []
for line in f:
    for word in line.split():
        if word not in listica:
            listica.append(word)

```

This deduplication ensures that each unique word appears only once in the processing queue, preventing redundant combinations later.

### The Concatenation Prompt

After loading completes, CUPP prompts the user at lines 204-210:

```text
> Do you want to concatenate all words from wordlist? Y/[N]:

```

If you answer **n**, CUPP processes only the original words. If you answer **y**, the tool executes the concatenation algorithm before proceeding to standard mutations.

## The Double-Loop Concatenation Algorithm

The core concatenation logic resides at lines 220-224 in [`cupp.py`](https://github.com/Mebus/cupp/blob/main/cupp.py). When enabled, CUPP initializes a `cont` list with an empty string and performs a Cartesian product of distinct words:

```python
cont = [""]
for cont1 in listica:
    for cont2 in listica:
        if listica.index(cont1) != listica.index(cont2):
            cont.append(cont1 + cont2)

```

This implementation exhibits three key characteristics:

- **Ordered pairs**: The nested loops generate every possible combination where `cont1` precedes `cont2`
- **Distinct indices**: The `index` check ensures `cont1` and `cont2` are different words, preventing same-word concatenation like `passwordpassword`
- **Empty string seed**: The initial empty string allows the combination engine to later work with "no extra word" as a valid option

## Integration with CUPP's Combination Engine

After populating `cont`, CUPP feeds this expanded list into its standard mutation pipeline. At lines 247-250, the concatenated words combine with years, special characters, and other modifiers:

```python

# From the improve_dictionary function

kombination[1] = list(komb(cont, years))
kombination[2] = list(komb(cont, ['', '!', '@', '#', '$', '%']))

```

This means a wordlist containing `["admin", "root"]` becomes `["adminroot", "rootadmin"]` in the `cont` array, and each of these then receives suffixes like `2024`, `!`, and `@` through the normal combination machinery.

## Configuration Thresholds and Limits

CUPP respects a safety threshold defined in [`cupp.cfg`](https://github.com/Mebus/cupp/blob/main/cupp.cfg) to prevent memory exhaustion with large wordlists. At lines 108-112 in [`cupp.py`](https://github.com/Mebus/cupp/blob/main/cupp.py), the code checks `CONFIG["global"]["threshold"]`:

```python
if len(listica) > int(CONFIG["global"]["threshold"]):
    print("> Maximum number of words for concatenation is " + CONFIG["global"]["threshold"])
    print("> Check configuration file for increasing this number.")
    sys.exit()

```

If your wordlist exceeds this limit (default typically 5000-10000 words depending on version), CUPP warns you and exits to prevent the exponential explosion of combinations that would occur during the double-loop operation.

## Practical Usage Examples

### Basic Wordlist Enhancement

Enable concatenation to generate compound passwords from common base words:

```bash
python cupp.py -w common.txt

# > Do you want to concatenate all words from wordlist? Y/[N]: y

# Generates: adminroot, rootadmin, adminuser, useradmin, etc.

```

### Processing Without Concatenation

Apply only standard mutations (years, special chars) to original words:

```bash
python cupp.py -w common.txt

# > Do you want to concatenate all words from wordlist? Y/[N]: n

# Processes only: admin, root, user (with mutations)

```

### Handling Large Wordlists

When exceeding the configured threshold, CUPP enforces limits:

```bash
python cupp.py -w huge.txt

# > Do you want to concatenate all words from wordlist? Y/[N]: y

# [-] Maximum number of words for concatenation is 5000

# > Do you want to concatenate all words from wordlist? Y/[N]: n

```

## Summary

- The `-w` option activates the `improve_dictionary` function in [`cupp.py`](https://github.com/Mebus/cupp/blob/main/cupp.py) to enhance existing wordlists
- Words load into the `listica` array at lines 191-197, with deduplication applied during ingestion
- A nested double-loop at lines 220-224 generates all ordered pairs of distinct words, storing results in the `cont` list
- The concatenated results feed into standard CUPP mutations (years, special characters) through the `komb` function
- The `threshold` setting in [`cupp.cfg`](https://github.com/Mebus/cupp/blob/main/cupp.cfg) prevents memory exhaustion by limiting the maximum wordlist size for concatenation operations

## Frequently Asked Questions

### Does CUPP concatenate a word with itself when using the -w option?

No. The code explicitly checks `if listica.index(cont1) != listica.index(cont2)` before concatenating, ensuring that `passwordpassword` or similar self-repeating patterns are excluded from the generated combinations.

### How does the -w option affect memory usage with large wordlists?

The double-loop algorithm creates **n × (n-1)** new strings (where n is your wordlist length), which can exhaust RAM quickly. CUPP prevents this by checking against `CONFIG["global"]["threshold"]` at lines 108-112 in [`cupp.py`](https://github.com/Mebus/cupp/blob/main/cupp.py) and warning users if their list exceeds the safe limit defined in [`cupp.cfg`](https://github.com/Mebus/cupp/blob/main/cupp.cfg).

### Can I use the -w option with the interactive profiling mode?

No. The `-w` option triggers `improve_dictionary` as a standalone workflow in [`cupp.py`](https://github.com/Mebus/cupp/blob/main/cupp.py). It operates exclusively on existing wordlist files and cannot combine with the interactive questionnaire mode that generates profiles from personal data like birthdates or pet names.

### What happens if I answer "no" to the concatenation prompt?

If you decline concatenation at the prompt (lines 204-210), CUPP skips the double-loop entirely and proceeds directly to standard mutations. The original `listica` words still receive years, special characters, and leetspeak transformations, but no cross-word combinations are generated.