How CUPP Parses and Extracts Data from the Alecto DB CSV Format: A Complete Guide

CUPP downloads the Alecto database as a gzipped CSV file, extracts usernames from the 6th column and passwords from the 7th column using Python's csv module, deduplicates the results with set() operations, and exports them to separate text files.

The Common User Passwords Profiler (CUPP) is a specialized tool for creating targeted wordlists used in security assessments. When processing the Alecto database—a compressed CSV containing default credentials—CUPP implements a specific parsing pipeline to extract and clean credential data according to the Mebus/cupp source code.

How CUPP Accesses the Alecto Database

Configuration and URL Lookup

The Alecto CSV URL is defined in cupp.cfg under the alectourl key. The read_config() function loads this configuration value into CONFIG["global"]["alectourl"] at runtime, allowing users to modify the download source without editing the core script.

Download and Cache Management

The alectodb_download() function handles file acquisition. It first checks whether alectodb.csv.gz exists in the working directory. If the file is missing, the function invokes download_http() to fetch the compressed database from the configured URL. This caching mechanism ensures CUPP can operate offline after the initial download.

Parsing the Alecto DB CSV Format

Decompressing the Gzipped File

CUPP opens the compressed archive using gzip.open(..., "rt") in text mode. This approach provides a standard file handle compatible with Python's built-in csv module, allowing direct iteration over the decompressed content without extracting the file to disk.

Extracting Credentials from CSV Columns

The csv.reader object iterates through each row of the Alecto database. According to the Alecto schema implemented in CUPP:

  • Username: Extracted from row[5] (the 6th column)
  • Password: Extracted from row[6] (the 7th column)

These specific column indices correspond to the credential fields in the Alecto database structure, as defined in the parsing logic.

Deduplication and Sorting

After extraction, CUPP places usernames and passwords into separate lists. It removes duplicates using list(set(...)) and sorts the results alphabetically. This processing ensures the final output contains only unique, organized entries suitable for password security testing.

Output Generation

The processed credential sets are written to two distinct files in the working directory:

These files serve as ready-to-use wordlists for penetration testing workflows.

Using the Alecto DB Feature in CUPP

To parse the Alecto database, execute CUPP with the -a flag:

python cupp.py -a

This command triggers the complete download-parse-export pipeline. If alectodb.csv.gz already exists locally, CUPP skips the network request and proceeds directly to parsing the cached file.

After execution, verify the generated wordlists:

ls -la alectodb-*.txt

Summary

  • Configuration: The Alecto URL is stored in cupp.cfg under alectourl and loaded by read_config() into CONFIG["global"]["alectourl"].
  • Acquisition: alectodb_download() fetches alectodb.csv.gz via download_http() only if the local cache is missing.
  • Decompression: gzip.open() handles the compressed file in text mode ("rt") for efficient streaming.
  • Extraction: csv.reader extracts usernames from column 5 (row[5]) and passwords from column 6 (row[6]).
  • Processing: Credentials are deduplicated with list(set(...)) and sorted alphabetically before export.
  • Output: Results are written to alectodb-usernames.txt and alectodb-passwords.txt in the working directory.

Frequently Asked Questions

What Python libraries does CUPP use to parse the Alecto CSV?

CUPP uses Python's standard gzip module to decompress the file and the csv module to read the data. The gzip.open() function is called with "rt" mode to provide a text-mode file handle that csv.reader can process directly, eliminating the need for temporary extraction.

Which columns contain the username and password in the Alecto database?

The Alecto database CSV format stores usernames in the 6th column (index 5) and passwords in the 7th column (index 6). CUPP accesses these via row[5] and row[6] respectively during the parsing loop in alectodb_download().

Does CUPP download the Alecto database every time it runs?

No. The alectodb_download() function checks for the existence of alectodb.csv.gz locally before initiating a download. If the file exists, CUPP proceeds directly to parsing the cached version, saving bandwidth and improving execution speed.

How does CUPP ensure the output files contain only unique entries?

CUPP converts the extracted credential lists to sets using set() to remove duplicates, then converts them back to lists with list(). Finally, it applies sort() to organize the entries alphabetically before writing to alectodb-usernames.txt and alectodb-passwords.txt.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →