How Holehe Aggregates Module Results: A Complete Guide to the Core Pipeline
Holehe aggregates results from its individual modules by running concurrent async tasks that push uniform result dictionaries into a shared list, then sorting and formatting that list for display or CSV export.
The open-source OSINT tool megadose/holehe discovers email registrations across thousands of web services through a carefully orchestrated aggregation system. This article explains exactly how Holehe collects, combines, and presents results from its modular service checks.
How Holehe Module Aggregation Works
The aggregation pipeline lives entirely in holehe/core.py and follows six distinct stages from module discovery to final output.
1. Module Discovery via import_submodules()
Holehe begins by dynamically discovering every available service module. The import_submodules() function walks the holehe.modules package and imports each Python file containing a service check.
# From core.py lines 37-47
modules = import_submodules("holehe.modules")
This approach allows the tool to support new services automatically—simply adding a .py file to the modules directory makes it available without modifying core logic.
2. Callable Extraction via get_functions()
After discovery, get_functions() extracts the public async function from each imported module. These functions—named after their service like google, amazon, twitter—populate a list called websites.
# From core.py lines 50-63
website_funcs = get_functions(modules)
Each function follows the same signature: it accepts an email, an HTTP client, and appends its result to a shared collection.
3. Concurrent Execution with Trio
Holehe launches all service checks in parallel using a Trio nursery. For every module in websites, it starts a task calling launch_module() with:
- The target email address
- A shared
httpx.AsyncClientfor connection pooling - A shared list
outfor result collection
# From core.py lines 18-21 and 66-78
async with trio.open_nursery() as nursery:
for website in websites:
nursery.start_soon(launch_module, website, email, client, out)
This concurrent approach enables scanning hundreds of services in seconds rather than minutes.
4. Individual Result Collection
Each module appends a standardized dictionary to out with these fields:
| Field | Description |
|---|---|
name |
Service identifier (e.g., "google", "github") |
exists |
Boolean indicating if email is registered |
rateLimit |
Boolean if rate-limited during check |
emailrecovery |
Recovery email if exposed by service |
phoneNumber |
Phone number if exposed by service |
others |
Additional metadata from service response |
A successful check returns exists=True or exists=False with optional extra fields. If a module raises an exception, launch_module() catches it and records an error entry:
# Error handling in launch_module() lines 71-78
out.append({
"name": module.__name__,
"exists": False,
"rateLimit": False,
"error": True # Indicates exception occurred
})
5. Sorting and Terminal Rendering
After all tasks complete, the out list undergoes alphabetical sorting by module name (i['name']), then passes to print_result(). This function applies color-coded symbols for immediate visual parsing:
[+]— Account exists[-]— Account does not exist[!]— Rate limited or error occurred[x]— Other issues
# From core.py lines 22-24
out.sort(key=lambda i: i["name"])
print_result(out, args, start_time, email)
6. Optional CSV Export
When invoked with --csv, Holehe writes the raw out list to a CSV file via export_csv():
holehe target@email.com --csv results.csv
The CSV contains the complete dictionary data, enabling further analysis in spreadsheet tools or custom scripts.
Practical Usage Examples
Command-Line Aggregation
# Standard execution with colorized output
holehe alice@example.com
# Filter to show only existing accounts
holehe alice@example.com --only-used
# Generate CSV alongside terminal output
holehe alice@example.com -C ./results/
Programmatic Aggregation
For custom workflows, replicate Holehe's aggregation pipeline directly:
import asyncio
import httpx
from holehe.core import import_submodules, get_functions, launch_module
async def aggregate_results(email: str):
# Stage 1: Discover and extract modules
modules = import_submodules("holehe.modules")
website_funcs = get_functions(modules)
# Stage 2: Initialize shared resources
client = httpx.AsyncClient(timeout=10)
out = []
# Stage 3: Execute concurrently (asyncio version)
await asyncio.gather(*[
launch_module(func, email, client, out)
for func in website_funcs
])
await client.aclose()
# Stage 4: Sort and return
out.sort(key=lambda i: i["name"])
return out
# Execute
results = asyncio.run(aggregate_results("alice@example.com"))
print(f"Checked {len(results)} services")
This programmatic approach guarantees identical behavior to the CLI by using the same launch_module logic and sorting mechanism.
Key Technical Components
| File | Purpose | Critical Functions |
|---|---|---|
holehe/core.py |
Orchestrates entire aggregation pipeline | import_submodules(), get_functions(), launch_module(), print_result() |
holehe/modules/**/*.py |
Individual service implementations | One async function per service |
holehe/instruments.py |
Progress bar during execution | Trio instrumentation |
holehe/localuseragent.py |
Realistic User-Agent rotation | HTTP header management |
Summary
- Dynamic discovery in
core.pyautomatically loads all service modules without hardcoded lists - Concurrent execution via Trio nursery enables efficient parallel checking of hundreds of services
- Uniform result dictionaries pushed to a shared
outlist guarantee consistent data structure regardless of service - Automatic error capture ensures partial results remain usable even when individual modules fail
- Post-processing sorting and formatting produces human-readable output and machine-parseable CSV exports
Frequently Asked Questions
What happens if a Holehe module crashes during execution?
The launch_module() wrapper in core.py catches all exceptions and appends an error dictionary to the results. The aggregation continues with remaining modules, and the failed service appears in output with "error": True.
Can I control how many modules run simultaneously?
The official implementation uses Trio's default nursery behavior. For custom concurrency limits, modify the nursery configuration or replace with asyncio.Semaphore as shown in the programmatic example.
Why does Holehe sort results alphabetically rather than by outcome?
Alphabetical sorting provides deterministic, reproducible output across runs. This consistency aids debugging, diffing results, and programmatic parsing where service position matters more than success/failure grouping.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →