How CUPP Parses and Extracts Data from the Alecto DB CSV Format: A Complete Guide
CUPP downloads the Alecto database as a gzipped CSV file, extracts usernames from the 6th column and passwords from the 7th column using Python's csv module, deduplicates the results with set() operations, and exports them to separate text files.
The Common User Passwords Profiler (CUPP) is a specialized tool for creating targeted wordlists used in security assessments. When processing the Alecto database—a compressed CSV containing default credentials—CUPP implements a specific parsing pipeline to extract and clean credential data according to the Mebus/cupp source code.
How CUPP Accesses the Alecto Database
Configuration and URL Lookup
The Alecto CSV URL is defined in cupp.cfg under the alectourl key. The read_config() function loads this configuration value into CONFIG["global"]["alectourl"] at runtime, allowing users to modify the download source without editing the core script.
Download and Cache Management
The alectodb_download() function handles file acquisition. It first checks whether alectodb.csv.gz exists in the working directory. If the file is missing, the function invokes download_http() to fetch the compressed database from the configured URL. This caching mechanism ensures CUPP can operate offline after the initial download.
Parsing the Alecto DB CSV Format
Decompressing the Gzipped File
CUPP opens the compressed archive using gzip.open(..., "rt") in text mode. This approach provides a standard file handle compatible with Python's built-in csv module, allowing direct iteration over the decompressed content without extracting the file to disk.
Extracting Credentials from CSV Columns
The csv.reader object iterates through each row of the Alecto database. According to the Alecto schema implemented in CUPP:
- Username: Extracted from
row[5](the 6th column) - Password: Extracted from
row[6](the 7th column)
These specific column indices correspond to the credential fields in the Alecto database structure, as defined in the parsing logic.
Deduplication and Sorting
After extraction, CUPP places usernames and passwords into separate lists. It removes duplicates using list(set(...)) and sorts the results alphabetically. This processing ensures the final output contains only unique, organized entries suitable for password security testing.
Output Generation
The processed credential sets are written to two distinct files in the working directory:
alectodb-usernames.txt– Contains all unique usernames extracted from the CSValectodb-passwords.txt– Contains all unique passwords extracted from the CSV
These files serve as ready-to-use wordlists for penetration testing workflows.
Using the Alecto DB Feature in CUPP
To parse the Alecto database, execute CUPP with the -a flag:
python cupp.py -a
This command triggers the complete download-parse-export pipeline. If alectodb.csv.gz already exists locally, CUPP skips the network request and proceeds directly to parsing the cached file.
After execution, verify the generated wordlists:
ls -la alectodb-*.txt
Summary
- Configuration: The Alecto URL is stored in
cupp.cfgunderalectourland loaded byread_config()intoCONFIG["global"]["alectourl"]. - Acquisition:
alectodb_download()fetchesalectodb.csv.gzviadownload_http()only if the local cache is missing. - Decompression:
gzip.open()handles the compressed file in text mode ("rt") for efficient streaming. - Extraction:
csv.readerextracts usernames from column 5 (row[5]) and passwords from column 6 (row[6]). - Processing: Credentials are deduplicated with
list(set(...))and sorted alphabetically before export. - Output: Results are written to
alectodb-usernames.txtandalectodb-passwords.txtin the working directory.
Frequently Asked Questions
What Python libraries does CUPP use to parse the Alecto CSV?
CUPP uses Python's standard gzip module to decompress the file and the csv module to read the data. The gzip.open() function is called with "rt" mode to provide a text-mode file handle that csv.reader can process directly, eliminating the need for temporary extraction.
Which columns contain the username and password in the Alecto database?
The Alecto database CSV format stores usernames in the 6th column (index 5) and passwords in the 7th column (index 6). CUPP accesses these via row[5] and row[6] respectively during the parsing loop in alectodb_download().
Does CUPP download the Alecto database every time it runs?
No. The alectodb_download() function checks for the existence of alectodb.csv.gz locally before initiating a download. If the file exists, CUPP proceeds directly to parsing the cached version, saving bandwidth and improving execution speed.
How does CUPP ensure the output files contain only unique entries?
CUPP converts the extracted credential lists to sets using set() to remove duplicates, then converts them back to lists with list(). Finally, it applies sort() to organize the entries alphabetically before writing to alectodb-usernames.txt and alectodb-passwords.txt.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →