How to Use Holehe Programmatically in a Python Script: Three Methods Explained

You can use Holehe programmatically by importing the main() function from holehe.core to emulate CLI behavior, or by importing the async helper functions import_submodules and get_functions to orchestrate service checks directly within your own async workflows.

Holehe, developed by megadose, is a command-line tool for email reconnaissance across hundreds of platforms. While designed as a CLI application, its core logic resides in standard Python modules located in holehe/core.py and holehe/modules/. Because the tool is built on regular Python coroutines using Trio, you can import and call its functions directly from any script to use Holehe programmatically without invoking subprocess commands.

Understanding Holehe's Core Architecture

The entry point for the command-line interface is the main() function in holehe/core.py. This function immediately delegates to maincore(), an asynchronous coroutine executed via trio.run(maincore). Inside maincore(), the code uses argparse to process arguments, dynamically imports every service module under holehe/modules/ using helper utilities, instantiates an httpx.AsyncClient, and launches each service's check function within a Trio nursery for parallel execution. The results are collected as a list of dictionaries containing keys like domain, exists, and emailrecovery.

Key files that enable programmatic use include:

Method 1: Emulating the CLI Entry Point

The simplest way to use Holehe programmatically is to import the existing CLI entry point and manipulate sys.argv before calling it. This requires no knowledge of the internal async implementation and produces identical console output to running holehe you@example.com from the shell.

import sys
from holehe.core import main

# Emulate the CLI call: `holehe you@example.com`

sys.argv = ["holehe", "you@example.com"]
main()                     # prints the result to stdout

By setting sys.argv before the call, the built-in argument parser inside maincore() receives the target email and runs the full suite of checks exactly as it would from the command line.

Method 2: Driving the Async Workflow Directly

For full control over execution and result handling, import the async infrastructure directly from holehe/core.py. This approach avoids console output entirely and allows you to process the results in-memory.

import asyncio
import httpx
import trio
from holehe.core import import_submodules, get_functions, TrioProgress

async def run_holehe(email: str):
    # Load all service modules

    modules = import_submodules("holehe.modules")
    # Retrieve the coroutine functions for each service

    websites = get_functions(modules)

    # Use the same timeout as the CLI (default 10 s)

    client = httpx.AsyncClient(timeout=10)

    # Prepare progress instrument (optional, just like the CLI)

    instrument = TrioProgress(len(websites))
    trio.lowlevel.add_instrument(instrument)

    results = []
    async with trio.open_nursery() as nursery:
        for site in websites:
            # launch each service coroutine

            nursery.start_soon(site, email, client, results)

    trio.lowlevel.remove_instrument(instrument)
    await client.aclose()
    return results

# Example usage

if __name__ == "__main__":
    email = "you@example.com"
    data = asyncio.run(run_holehe(email))
    for entry in data:
        print(entry)           # each entry is a dict with keys like 'domain', 'exists', …

The import_submodules utility walks the holehe.modules package and imports every submodule, while get_functions extracts the async check functions. The code creates an httpx.AsyncClient, opens a Trio nursery for structured concurrency, and runs each service coroutine in parallel. The results list captures the same dictionaries that the CLI would print, allowing you to handle the data programmatically.

Method 3: Querying Specific Services Selectively

When you only need to check a few platforms rather than the entire module catalog, access specific modules directly from the dictionary returned by import_submodules. This reduces network traffic and execution time significantly.

import asyncio
import httpx
import trio
from holehe.core import import_submodules

async def check_twitter_and_github(email: str):
    modules = import_submodules("holehe.modules")
    # Pull out just the two services we care about

    twitter = modules["holehe.modules.social_media.twitter"].twitter
    github  = modules["holehe.modules.programing.github"].github

    client = httpx.AsyncClient(timeout=10)
    results = []
    async with trio.open_nursery() as nursery:
        nursery.start_soon(twitter, email, client, results)
        nursery.start_soon(github,  email, client, results)
    await client.aclose()
    return results

if __name__ == "__main__":
    data = asyncio.run(check_twitter_and_github("you@example.com"))
    print(data)

By accessing the specific module objects (for example, modules["holehe.modules.social_media.twitter"]), you bypass the automatic discovery logic and invoke only the services you need, giving you precise control over the reconnaissance scope.

Summary

  • CLI Emulation: Import main() from holehe/core.py and set sys.argv to quickly run checks with console output exactly like the command-line tool.
  • Full Async Control: Use import_submodules() and get_functions() to load all services, then orchestrate them with trio.open_nursery() and an httpx.AsyncClient for complete in-memory result handling.
  • Selective Execution: Access specific modules by name from the imported dictionary to run targeted checks on individual platforms, optimizing for speed and reducing network footprint.

Frequently Asked Questions

Can I use Holehe in a standard asyncio application?

The source code in holehe/core.py is built on Trio for structured concurrency. While the examples above use asyncio.run() to launch the coroutines, Holehe internally uses trio.open_nursery() and Trio-specific instruments. For production use, consider running Holehe within a Trio-native context using trio.run() directly to avoid potential compatibility issues between the two async libraries.

What data structure does Holehe return when used programmatically?

When you drive the async workflow directly, Holehe appends results to the list you provide as the third argument to each service function. Each entry is a dictionary containing keys such as domain, exists, emailrecovery, and others, indicating whether the email was found on that platform and any additional metadata extracted during the check.

How do I customize the HTTP client configuration?

Both programmatic approaches allow you to instantiate your own httpx.AsyncClient. When using Method 2 or 3, create the client with your preferred settings (for example, timeout=10 for a 10-second timeout as used in the CLI, or custom proxies and headers) and pass it as the second argument to the service coroutines instead of relying on the default client created by maincore().

Is it possible to run Holehe without installing it as a package?

Yes, as long as the holehe directory is in your Python path, you can import from holehe.core regardless of installation method. However, the tool depends on external packages like httpx and trio, so ensure these are installed in your environment. The modular design in holehe/modules/__init__.py ensures all service sub-packages are importable once the parent package is accessible.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →