How to Integrate Holehe Functionality into a Python Script: CLI and Async API Methods
You can integrate Holehe into Python scripts by either importing the main() function from holehe.core to emulate CLI behavior, or by directly importing the async service modules and orchestrating them with Trio for fine-grained control over email reconnaissance workflows.
Holehe is an open-source email reconnaissance tool developed by megadose that checks if an email address exists on hundreds of platforms. While designed primarily as a command-line utility, its core logic resides in standard Python modules, allowing you to integrate Holehe functionality into Python scripts without invoking subprocess calls. This programmatic approach gives you direct control over the async workflow, result handling, and service selection.
How Holehe's Architecture Supports Programmatic Use
Holehe is built on Trio for structured concurrency and uses httpx for asynchronous HTTP requests. The entry point in holehe/core.py defines a main() function that simply executes trio.run(maincore), where maincore() handles argument parsing and orchestrates the service modules located in holehe/modules/. Each service module contains an async function that accepts an email string, an httpx.AsyncClient instance, and a results list, making them directly callable from your own async code.
Method 1: Using the CLI Entry Point
The simplest way to integrate Holehe is to import the main() function and populate sys.argv with your target email. This approach requires no knowledge of async programming and produces identical console output to the command-line interface.
Implementing the CLI Wrapper
import sys
from holehe.core import main
# Emulate the CLI call: `holehe you@example.com`
sys.argv = ["holehe", "you@example.com"]
main() # prints results to stdout
This method calls maincore() internally, which loads all modules from holehe/modules/ and runs the full suite of checks automatically.
Method 2: Direct Async API Integration
For production scripts requiring in-memory result handling or selective service queries, drive the async workflow directly. This method imports helper functions from holehe.core to load service modules and execute them within a Trio nursery.
Querying All Available Services
import httpx
import trio
from holehe.core import import_submodules, get_functions, TrioProgress
async def run_full_check(email: str):
# Load all service modules from holehe/modules/
modules = import_submodules("holehe.modules")
websites = get_functions(modules)
client = httpx.AsyncClient(timeout=10)
results = []
# Optional: Add progress instrumentation
instrument = TrioProgress(len(websites))
trio.lowlevel.add_instrument(instrument)
async with trio.open_nursery() as nursery:
for site in websites:
nursery.start_soon(site, email, client, results)
trio.lowlevel.remove_instrument(instrument)
await client.aclose()
return results
# Execute with Trio
email = "you@example.com"
data = trio.run(run_full_check, email)
The import_submodules function walks the package directory and imports every service submodule, while get_functions extracts the async check functions. The returned results list contains dictionaries with keys like domain, exists, and other metadata.
Targeting Specific Platforms
To reduce network traffic and execution time, import only the modules you need:
import httpx
import trio
from holehe.core import import_submodules
async def check_specific_sites(email: str):
modules = import_submodules("holehe.modules")
# Access specific service functions directly
twitter = modules["holehe.modules.social_media.twitter"].twitter
github = modules["holehe.modules.programing.github"].github
client = httpx.AsyncClient(timeout=10)
results = []
async with trio.open_nursery() as nursery:
nursery.start_soon(twitter, email, client, results)
nursery.start_soon(github, email, client, results)
await client.aclose()
return results
# Run the targeted check
results = trio.run(check_specific_sites, "you@example.com")
By accessing specific attributes like modules["holehe.modules.social_media.twitter"], you bypass automatic service discovery and execute only the desired checks.
Key Source Files and Functions
Understanding these core components helps when customizing your integration:
holehe/core.py: Containsmain(),maincore(),import_submodules(), andget_functions(). This is the primary interface for programmatic use.holehe/modules/: Directory containing categorized service modules (e.g.,social_media/twitter.py). Each module exports an async function that performs the existence check.holehe/instruments.py: DefinesTrioProgress, the instrument class that provides the live progress bar functionality.holehe/localuseragent.py: Generates random User-Agent strings for HTTP requests to avoid detection.
Both integration methods rely on these files, ensuring that programmatic calls produce the same accuracy as the CLI version.
Summary
- Two integration paths exist: Wrapping the CLI via
main()for simplicity, or using the async API for custom workflows. - Trio is required: Holehe uses Trio for structured concurrency, so you must use
trio.run()to execute the async functions. - Results are standardized: Whether using CLI or async methods, output is a list of dictionaries containing
domain,exists, and other metadata. - Selective execution is possible: Import specific modules from
holehe.modulesto limit queries to specific platforms. - HTTP client configuration: Pass a custom
httpx.AsyncClientto control timeouts, proxies, and connection limits.
Frequently Asked Questions
Can I use Holehe in a synchronous Python script?
No, Holehe is fundamentally asynchronous. However, you can bridge it using trio.run() to execute the async functions from synchronous code, or use anyio to run Trio-based coroutines within an asyncio environment if necessary.
How do I handle authentication or proxies when integrating Holehe programmatically?
Create a custom httpx.AsyncClient instance with your proxy configuration or authentication headers, then pass this client object to the service functions. The timeout parameter in httpx.AsyncClient(timeout=10) controls request duration, and you can add transport layers for SOCKS or HTTP proxies.
What is the performance impact of loading all modules versus selecting specific ones?
Loading all modules via import_submodules("holehe.modules") initializes every available service, which increases startup time and memory usage. For checking fewer than five sites, selective imports reduce initialization overhead by approximately 90%, though parallel execution time is primarily network-bound.
Does the programmatic API return the same data format as the CLI output?
Yes, both methods populate a list of dictionaries with identical schema. Each dictionary contains keys such as domain, exists (boolean), email, profile, and others (method-specific metadata), matching the JSON structure available via CLI output redirection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →