# How CUPP Parses and Extracts Data from the Alecto DB CSV Format: A Complete Guide

> Learn how CUPP parses Alecto DB CSV data Extract usernames and passwords from the 6th and 7th columns respectively. Includes deduplication and export.

- Repository: [Mebus/cupp](https://github.com/Mebus/cupp)
- Tags: how-to-guide
- Published: 2026-07-03

---

**CUPP downloads the Alecto database as a gzipped CSV file, extracts usernames from the 6th column and passwords from the 7th column using Python's `csv` module, deduplicates the results with `set()` operations, and exports them to separate text files.**

The Common User Passwords Profiler (CUPP) is a specialized tool for creating targeted wordlists used in security assessments. When processing the Alecto database—a compressed CSV containing default credentials—CUPP implements a specific parsing pipeline to extract and clean credential data according to the Mebus/cupp source code.

## How CUPP Accesses the Alecto Database

### Configuration and URL Lookup

The Alecto CSV URL is defined in [`cupp.cfg`](https://github.com/Mebus/cupp/blob/main/cupp.cfg) under the `alectourl` key. The `read_config()` function loads this configuration value into `CONFIG["global"]["alectourl"]` at runtime, allowing users to modify the download source without editing the core script.

### Download and Cache Management

The `alectodb_download()` function handles file acquisition. It first checks whether `alectodb.csv.gz` exists in the working directory. If the file is missing, the function invokes `download_http()` to fetch the compressed database from the configured URL. This caching mechanism ensures CUPP can operate offline after the initial download.

## Parsing the Alecto DB CSV Format

### Decompressing the Gzipped File

CUPP opens the compressed archive using `gzip.open(..., "rt")` in text mode. This approach provides a standard file handle compatible with Python's built-in `csv` module, allowing direct iteration over the decompressed content without extracting the file to disk.

### Extracting Credentials from CSV Columns

The `csv.reader` object iterates through each row of the Alecto database. According to the Alecto schema implemented in CUPP:

- **Username**: Extracted from `row[5]` (the 6th column)
- **Password**: Extracted from `row[6]` (the 7th column)

These specific column indices correspond to the credential fields in the Alecto database structure, as defined in the parsing logic.

### Deduplication and Sorting

After extraction, CUPP places usernames and passwords into separate lists. It removes duplicates using `list(set(...))` and sorts the results alphabetically. This processing ensures the final output contains only unique, organized entries suitable for password security testing.

## Output Generation

The processed credential sets are written to two distinct files in the working directory:

- [`alectodb-usernames.txt`](https://github.com/Mebus/cupp/blob/main/alectodb-usernames.txt) – Contains all unique usernames extracted from the CSV
- [`alectodb-passwords.txt`](https://github.com/Mebus/cupp/blob/main/alectodb-passwords.txt) – Contains all unique passwords extracted from the CSV

These files serve as ready-to-use wordlists for penetration testing workflows.

## Using the Alecto DB Feature in CUPP

To parse the Alecto database, execute CUPP with the `-a` flag:

```bash
python cupp.py -a

```

This command triggers the complete download-parse-export pipeline. If `alectodb.csv.gz` already exists locally, CUPP skips the network request and proceeds directly to parsing the cached file.

After execution, verify the generated wordlists:

```bash
ls -la alectodb-*.txt

```

## Summary

- **Configuration**: The Alecto URL is stored in [`cupp.cfg`](https://github.com/Mebus/cupp/blob/main/cupp.cfg) under `alectourl` and loaded by `read_config()` into `CONFIG["global"]["alectourl"]`.
- **Acquisition**: `alectodb_download()` fetches `alectodb.csv.gz` via `download_http()` only if the local cache is missing.
- **Decompression**: `gzip.open()` handles the compressed file in text mode (`"rt"`) for efficient streaming.
- **Extraction**: `csv.reader` extracts usernames from column 5 (`row[5]`) and passwords from column 6 (`row[6]`).
- **Processing**: Credentials are deduplicated with `list(set(...))` and sorted alphabetically before export.
- **Output**: Results are written to [`alectodb-usernames.txt`](https://github.com/Mebus/cupp/blob/main/alectodb-usernames.txt) and [`alectodb-passwords.txt`](https://github.com/Mebus/cupp/blob/main/alectodb-passwords.txt) in the working directory.

## Frequently Asked Questions

### What Python libraries does CUPP use to parse the Alecto CSV?

CUPP uses Python's standard `gzip` module to decompress the file and the `csv` module to read the data. The `gzip.open()` function is called with `"rt"` mode to provide a text-mode file handle that `csv.reader` can process directly, eliminating the need for temporary extraction.

### Which columns contain the username and password in the Alecto database?

The Alecto database CSV format stores usernames in the 6th column (index 5) and passwords in the 7th column (index 6). CUPP accesses these via `row[5]` and `row[6]` respectively during the parsing loop in `alectodb_download()`.

### Does CUPP download the Alecto database every time it runs?

No. The `alectodb_download()` function checks for the existence of `alectodb.csv.gz` locally before initiating a download. If the file exists, CUPP proceeds directly to parsing the cached version, saving bandwidth and improving execution speed.

### How does CUPP ensure the output files contain only unique entries?

CUPP converts the extracted credential lists to sets using `set()` to remove duplicates, then converts them back to lists with `list()`. Finally, it applies `sort()` to organize the entries alphabetically before writing to [`alectodb-usernames.txt`](https://github.com/Mebus/cupp/blob/main/alectodb-usernames.txt) and [`alectodb-passwords.txt`](https://github.com/Mebus/cupp/blob/main/alectodb-passwords.txt).