How to Use holehe Programmatically in Python: A Complete API Guide
You can use holehe programmatically by importing the core functions import_submodules(), get_functions(), and launch_module() from the holehe.core module to run async email checks without invoking the CLI.
The holehe tool by megadose/holehe is distributed primarily as a command-line utility for checking if an email address is registered on various websites. However, its internal architecture exposes a clean Python API that allows you to embed email reconnaissance directly into your applications, scripts, or web services. When you use holehe programmatically in Python, you bypass the argument parser and interact directly with the async discovery and execution engine.
Understanding holehe's Internal Architecture
The library is structured around three distinct layers that handle module discovery, function extraction, and concurrent execution. Understanding these layers helps you integrate the tool effectively into your codebase.
Module Discovery Layer
The function holehe.core.import_submodules() dynamically walks the holehe.modules package to load every site-specific checker as a Python module. Located in holehe/core.py at lines 37-47, this function returns a list of imported module objects representing services like Twitter, Instagram, and GitHub.
Function Extraction Layer
Once modules are loaded, holehe.core.get_functions() extracts the callable async functions from each module. As implemented in holehe/core.py at lines 50-63, this returns a list of coroutine objects where each follows the signature async def <site>(email, client, out).
Async Execution Engine
The holehe.core.launch_module() helper invokes individual site-checkers using an httpx.AsyncClient and appends standardized result dictionaries to a shared list. This logic resides in holehe/core.py at lines 66-78. The main entry point holehe.core.maincore() (lines 79-86) orchestrates these calls within a Trio nursery for concurrent execution.
Basic Programmatic Usage Example
To use holehe programmatically, you need to replicate the CLI's initialization logic while providing your own email string and handling the results list directly. The following example demonstrates the complete workflow:
import httpx
import trio
from holehe.core import import_submodules, get_functions, launch_module, TrioProgress
async def check_email_exists(email: str):
# Step 1: Discover all site modules
modules = import_submodules("holehe.modules")
# Step 2: Extract checker functions
websites = get_functions(modules)
# Step 3: Initialize HTTP client with custom timeout
client = httpx.AsyncClient(timeout=10)
# Step 4: Prepare output container and progress instrument
results = []
instrument = TrioProgress(len(websites))
trio.lowlevel.add_instrument(instrument)
# Step 5: Execute all checks concurrently using Trio
async with trio.open_nursery() as nursery:
for site_func in websites:
nursery.start_soon(launch_module, site_func, email, client, results)
trio.lowlevel.remove_instrument(instrument)
await client.aclose()
return results
# Run synchronously if needed
if __name__ == "__main__":
email = "target@example.com"
output = trio.run(check_email_exists, email)
for result in output:
print(f"{result['domain']}: {'Found' if result['exists'] else 'Not found'}")
Key API Functions and Their Roles
When integrating holehe into your projects, you will interact primarily with four functions from holehe/core.py.
import_submodules(package_name) accepts a string like "holehe.modules" and returns a list of imported module objects. This function uses pkgutil.walk_packages() to discover every Python file in the modules directory without requiring manual imports.
get_functions(modules) processes the list returned by import_submodules() and extracts the async checker functions. Each module in holehe contains exactly one public async function that implements the site-specific logic, making this extraction straightforward.
launch_module(func, email, client, out) executes a single site checker. It accepts the async function, the target email string, an httpx.AsyncClient instance for connection reuse, and a mutable list out where results are appended. This function handles exceptions internally to prevent one failed check from crashing the entire batch.
maincore(email, timeout) provides a higher-level interface that wraps the above steps, creates the nursery, and manages the full execution lifecycle. According to the source code at lines 79-86, this is the function that the CLI actually calls after parsing arguments.
Working with Result Dictionaries
Each site checker returns a standardized dictionary that is appended to your output list. Understanding this structure allows you to filter and process results effectively.
Result Dictionary Structure
The dictionary returned by each module contains consistent keys including "domain" (the service name), "exists" (boolean indicating registration status), and other metadata fields. When you use holehe programmatically, you can access these fields immediately without parsing console output.
Filtering and Formatting Results
After execution, you can process the results list using standard Python list operations:
# Filter only sites where email is registered
active_accounts = [r for r in results if r["exists"]]
# Format for reporting
for entry in results:
status = "[+] " if entry["exists"] else "[-] "
print(f"{status}{entry['domain']}: {'registered' if entry['exists'] else 'not found'}")
Synchronous vs Async Execution
While holehe is built on async/await patterns using Trio, you can invoke it from synchronous code using trio.run(). The CLI entry point in holehe/core.py at lines 32-34 uses exactly this pattern: trio.run(maincore) wraps the async execution in a blocking call suitable for scripts.
If your application already uses asyncio, you can bridge the two frameworks using anyio or run holehe in a separate thread. However, the native approach uses Trio instruments like TrioProgress for progress tracking, which is safe to include even when running programmatically as it only affects console output when attached.
Summary
- Import the core functions from
holehe.coreto bypass the CLI interface and access the internal API directly. - Use
import_submodules()andget_functions()to dynamically discover and load all available site checkers without hardcoding module names. - Execute checks concurrently by passing
launch_module()to a Trio nursery along with a sharedhttpx.AsyncClientfor efficient connection pooling. - Handle results programmatically by inspecting the standardized dictionary objects appended to your output list, checking the
"exists"and"domain"keys. - Choose your concurrency backend by using either
trio.run()for synchronous scripts orasyncio.run()with appropriate bridging for async applications.
Frequently Asked Questions
Can I use holehe without installing the command-line interface?
Yes, you can use holehe as a pure Python library by importing from holehe.core directly. While the package installs a CLI entry point, all functionality is exposed through importable async functions. Simply install the package via pip, import the core module discovery and execution functions, and invoke them with your target email to use holehe programmatically in Python without shelling out to subprocess.
What data structure does holehe return when used programmatically?
Each site checker appends a dictionary to your provided output list with standardized keys. The dictionary typically includes "domain" (the service name as a string), "exists" (a boolean indicating if the email is registered), and additional fields like "email_recovery" or "others" depending on the specific module. This structured format makes it easy to parse results programmatically compared to parsing CLI stdout.
How do I configure HTTP timeouts when using the holehe API?
When initializing the httpx.AsyncClient before passing it to launch_module(), you can set the timeout parameter to an integer representing seconds. For example, httpx.AsyncClient(timeout=10) creates a client that fails requests after 10 seconds. This client object is then passed to each site checker through the launch_module() function, ensuring consistent timeout behavior across all concurrent checks.
Is it safe to run holehe checks concurrently in a production application?
Yes, the library is designed for concurrent execution using Trio nurseries. The launch_module() function wraps each check in exception handling, so one failing site will not crash the entire batch. However, be mindful of rate limits when running many checks simultaneously, as the default behavior attempts to query every discovered module at once. You can implement custom throttling by modifying the nursery loop or using Trio semaphores.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →