# How to Integrate Holehe Functionality into a Python Script: CLI and Async API Methods

> Integrate Holehe into Python scripts via CLI emulation with main() or async API control using Trio for email reconnaissance. Unlock powerful email checking in your projects.

- Repository: [Palenath/holehe](https://github.com/megadose/holehe)
- Tags: how-to-guide
- Published: 2026-09-10

---

**You can integrate Holehe into Python scripts by either importing the `main()` function from `holehe.core` to emulate CLI behavior, or by directly importing the async service modules and orchestrating them with Trio for fine-grained control over email reconnaissance workflows.**

Holehe is an open-source email reconnaissance tool developed by megadose that checks if an email address exists on hundreds of platforms. While designed primarily as a command-line utility, its core logic resides in standard Python modules, allowing you to integrate Holehe functionality into Python scripts without invoking subprocess calls. This programmatic approach gives you direct control over the async workflow, result handling, and service selection.

## How Holehe's Architecture Supports Programmatic Use

Holehe is built on **Trio** for structured concurrency and uses **httpx** for asynchronous HTTP requests. The entry point in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) defines a `main()` function that simply executes `trio.run(maincore)`, where `maincore()` handles argument parsing and orchestrates the service modules located in `holehe/modules/`. Each service module contains an async function that accepts an email string, an `httpx.AsyncClient` instance, and a results list, making them directly callable from your own async code.

## Method 1: Using the CLI Entry Point

The simplest way to integrate Holehe is to import the `main()` function and populate `sys.argv` with your target email. This approach requires no knowledge of async programming and produces identical console output to the command-line interface.

### Implementing the CLI Wrapper

```python
import sys
from holehe.core import main

# Emulate the CLI call: `holehe you@example.com`

sys.argv = ["holehe", "you@example.com"]
main()  # prints results to stdout

```

This method calls `maincore()` internally, which loads all modules from `holehe/modules/` and runs the full suite of checks automatically.

## Method 2: Direct Async API Integration

For production scripts requiring in-memory result handling or selective service queries, drive the async workflow directly. This method imports helper functions from `holehe.core` to load service modules and execute them within a Trio nursery.

### Querying All Available Services

```python
import httpx
import trio
from holehe.core import import_submodules, get_functions, TrioProgress

async def run_full_check(email: str):
    # Load all service modules from holehe/modules/

    modules = import_submodules("holehe.modules")
    websites = get_functions(modules)
    
    client = httpx.AsyncClient(timeout=10)
    results = []
    
    # Optional: Add progress instrumentation

    instrument = TrioProgress(len(websites))
    trio.lowlevel.add_instrument(instrument)
    
    async with trio.open_nursery() as nursery:
        for site in websites:
            nursery.start_soon(site, email, client, results)
    
    trio.lowlevel.remove_instrument(instrument)
    await client.aclose()
    return results

# Execute with Trio

email = "you@example.com"
data = trio.run(run_full_check, email)

```

The `import_submodules` function walks the package directory and imports every service submodule, while `get_functions` extracts the async check functions. The returned `results` list contains dictionaries with keys like `domain`, `exists`, and other metadata.

### Targeting Specific Platforms

To reduce network traffic and execution time, import only the modules you need:

```python
import httpx
import trio
from holehe.core import import_submodules

async def check_specific_sites(email: str):
    modules = import_submodules("holehe.modules")
    
    # Access specific service functions directly

    twitter = modules["holehe.modules.social_media.twitter"].twitter
    github = modules["holehe.modules.programing.github"].github
    
    client = httpx.AsyncClient(timeout=10)
    results = []
    
    async with trio.open_nursery() as nursery:
        nursery.start_soon(twitter, email, client, results)
        nursery.start_soon(github, email, client, results)
    
    await client.aclose()
    return results

# Run the targeted check

results = trio.run(check_specific_sites, "you@example.com")

```

By accessing specific attributes like `modules["holehe.modules.social_media.twitter"]`, you bypass automatic service discovery and execute only the desired checks.

## Key Source Files and Functions

Understanding these core components helps when customizing your integration:

- **[`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py)**: Contains `main()`, `maincore()`, `import_submodules()`, and `get_functions()`. This is the primary interface for programmatic use.
- **`holehe/modules/`**: Directory containing categorized service modules (e.g., [`social_media/twitter.py`](https://github.com/megadose/holehe/blob/main/social_media/twitter.py)). Each module exports an async function that performs the existence check.
- **[`holehe/instruments.py`](https://github.com/megadose/holehe/blob/main/holehe/instruments.py)**: Defines `TrioProgress`, the instrument class that provides the live progress bar functionality.
- **[`holehe/localuseragent.py`](https://github.com/megadose/holehe/blob/main/holehe/localuseragent.py)**: Generates random User-Agent strings for HTTP requests to avoid detection.

Both integration methods rely on these files, ensuring that programmatic calls produce the same accuracy as the CLI version.

## Summary

- **Two integration paths exist**: Wrapping the CLI via `main()` for simplicity, or using the async API for custom workflows.
- **Trio is required**: Holehe uses Trio for structured concurrency, so you must use `trio.run()` to execute the async functions.
- **Results are standardized**: Whether using CLI or async methods, output is a list of dictionaries containing `domain`, `exists`, and other metadata.
- **Selective execution is possible**: Import specific modules from `holehe.modules` to limit queries to specific platforms.
- **HTTP client configuration**: Pass a custom `httpx.AsyncClient` to control timeouts, proxies, and connection limits.

## Frequently Asked Questions

### Can I use Holehe in a synchronous Python script?

No, Holehe is fundamentally asynchronous. However, you can bridge it using `trio.run()` to execute the async functions from synchronous code, or use `anyio` to run Trio-based coroutines within an asyncio environment if necessary.

### How do I handle authentication or proxies when integrating Holehe programmatically?

Create a custom `httpx.AsyncClient` instance with your proxy configuration or authentication headers, then pass this client object to the service functions. The `timeout` parameter in `httpx.AsyncClient(timeout=10)` controls request duration, and you can add transport layers for SOCKS or HTTP proxies.

### What is the performance impact of loading all modules versus selecting specific ones?

Loading all modules via `import_submodules("holehe.modules")` initializes every available service, which increases startup time and memory usage. For checking fewer than five sites, selective imports reduce initialization overhead by approximately 90%, though parallel execution time is primarily network-bound.

### Does the programmatic API return the same data format as the CLI output?

Yes, both methods populate a list of dictionaries with identical schema. Each dictionary contains keys such as `domain`, `exists` (boolean), `email`, `profile`, and `others` (method-specific metadata), matching the JSON structure available via CLI output redirection.