# How Holehe Handles Concurrent Tasks for Scanning Websites

> Discover how Holehe uses Trio's structured concurrency to run hundreds of email checks simultaneously. Each website probe executes as an async function within a shared nursery scope for efficient scanning.

- Repository: [Palenath/holehe](https://github.com/megadose/holehe)
- Tags: internals
- Published: 2026-08-30

---

**Holehe uses Trio-based structured concurrency to run hundreds of email checks in parallel, with each site probe implemented as an async function executed through a shared nursery scope.**

Email reconnaissance tool **Holehe** by megadose needs to query dozens—sometimes hundreds—of websites to check if an email address exists on each platform. Doing this sequentially would take minutes; the project solves this through **asynchronous concurrent execution** powered by the Trio library. Every site check runs as a non-blocking coroutine, coordinated through Trio's nursery pattern for clean parallelism without callback hell.

## Async Module Functions for Non-Blocking Checks

Each supported service lives in `holehe/modules/` as a standalone file containing one `async def` function. These functions never block the event loop while waiting for HTTP responses.

Take the Twitter module as a representative example. In [`holehe/modules/social_media/twitter.py`](https://github.com/megadose/holehe/blob/main/holehe/modules/social_media/twitter.py), the function signature is:

```python
async def twitter(email, client, out):
    ...

```

The parameters are consistent across all modules:

- **`email`** – the target address to investigate
- **`client`** – a shared `httpx.AsyncClient` instance for HTTP requests
- **`out`** – a shared list where results are appended

This design means `twitter()` can `await` its HTTP requests without stalling other site checks. Instagram, Amazon, GitHub, and 100+ other services follow this identical pattern.

## Collecting and Scheduling Tasks

The orchestration logic lives in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py). Two utility functions gather all available probes:

```python
modules = import_submodules('holehe.modules')      # L50-L55

websites = get_functions(modules)                  # L56-L63

```

The `websites` list contains references to every async site-check function. This collection happens once at startup, then the same functions are reused across concurrent invocations.

## Progress Tracking with Trio Instruments

Before launching tasks, Holehe installs a custom **Trio instrument** that hooks into task lifecycle events. The `TrioProgress` class in [`holehe/instruments.py`](https://github.com/megadose/holehe/blob/main/holehe/instruments.py#L4-L10) updates a `tqdm` progress bar each time any site check completes:

```python
instrument = TrioProgress(total=len(websites))
trio.lowlevel.add_instrument(instrument)

```

This avoids cluttering the scanning logic with progress callbacks—the instrumentation stays orthogonal to the actual HTTP work.

## Concurrent Execution via Trio Nursery

The core concurrency mechanism appears in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) at lines 217-221:

```python
async with trio.open_nursery() as nursery:
    for module in websites:
        nursery.start_soon(
            launch_module, module, email, client, out
        )

```

Key components of this pattern:

- **`trio.open_nursery()`** – creates a scope that automatically waits for all child tasks
- **`nursery.start_soon()`** – schedules `launch_module` to run immediately without blocking the loop
- **`launch_module`** – a wrapper that awaits the site-specific function and normalizes exceptions into result dictionaries

When the `async with` block exits, every check has finished—guaranteed by Trio's **structured concurrency**. No zombie tasks, no orphaned connections.

## Shared HTTP Client for Efficiency

Connection pooling matters when opening hundreds of parallel requests. In [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py#L190-L197), a single `httpx.AsyncClient` is instantiated once:

```python
async with httpx.AsyncClient(timeout=10, headers=...) as client:
    # ... nursery execution happens here

```

Passing this shared client to every `launch_module` call allows HTTP/2 connection reuse and eliminates the overhead of per-request client creation. The client remains fully async-compatible with Trio's event loop.

## Result Aggregation

All async functions write to the same shared list `out` passed by reference. After the nursery closes, [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py#L166-L176) processes this list for display:

```python

# Inside the main flow

await Trio(...all checks...)  # nursery ensures everything finishes here

for result in out:
    # format and print or export to CSV

```

No locks are required because Python's `asyncio`-compatible structures and Trio's single-threaded concurrency model prevent race conditions on list appends.

## How to Use Holehe's Concurrency Model

### Command-Line Scanning

Run a standard concurrent scan from your terminal:

```bash

# Basic scan against all supported sites

holehe target@example.com

# Filter to only confirmed registrations

holehe target@example.com --only-used

# Export results for further analysis

holehe target@example.com --csv

```

### Programmatic Usage

Embed Holehe's concurrency in your own applications:

```python
import httpx
import trio
from holehe.core import import_submodules, get_functions, launch_module

async def scan_email(email: str):
    # Load all site-check functions

    modules = import_submodules('holehe.modules')
    websites = get_functions(modules)
    
    results = []
    
    async with httpx.AsyncClient(timeout=10) as client:
        async with trio.open_nursery() as nursery:
            for site_func in websites:
                nursery.start_soon(
                    launch_module, site_func, email, client, results
                )
    
    return results

# Execute with Trio's runner

trio.run(scan_email, 'someone@example.com')

```

### Custom Progress Tracking

Reuse Holehe's instrument for your own async workflows:

```python
from holehe.instruments import TrioProgress
import trio

async def monitored_scan(websites, email, client):
    progress = TrioProgress(total=len(websites))
    trio.lowlevel.add_instrument(progress)
    
    try:
        async with trio.open_nursery() as n:
            for func in websites:
                n.start_soon(launch_module, func, email, client, [])
    finally:
        trio.lowlevel.remove_instrument(progress)

```

## Why Trio Over Alternatives?

Holehe chose Trio specifically for three operational advantages:

- **Structured concurrency** – tasks cannot outlive their parent scope, eliminating cleanup bugs
- **Native instrumentation** – clean progress hooks without monkey-patching or decorators
- **Deterministic cancellation** – timeout handling that properly unwinds nested async operations

asyncio could achieve similar throughput, but Trio's semantics reduce edge-case bugs when managing hundreds of simultaneous network operations.

## Summary

- **Every site check is an `async def` function** in `holehe/modules/`, using non-blocking HTTP via `httpx.AsyncClient`
- **[`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py)** collects these functions and executes them through `trio.open_nursery()` for true parallelism
- **A single shared HTTP client** enables connection pooling across all concurrent requests
- **`TrioProgress`** instrument in [`holehe/instruments.py`](https://github.com/megadose/holehe/blob/main/holehe/instruments.py) provides clean progress reporting without invasive code
- **Structured concurrency guarantees** all tasks complete before final results are processed

## Frequently Asked Questions

### Why does Holehe use Trio instead of asyncio?

Trio provides **structured concurrency** with stricter guarantees about task lifetimes and cancellation. The instrument API also allows progress tracking without wrapping every function call. Both libraries achieve similar performance, but Trio's semantics reduce bugs in long-running concurrent network code.

### Can I adjust how many sites run simultaneously?

Trio's nursery doesn't use a fixed thread pool—coroutines yield cooperatively. However, you can partition the `websites` list and run multiple nurseries sequentially, or apply `anyio` limits if you fork the codebase. By default, all discovered modules run in parallel with whatever concurrency the event loop and OS can sustain.

### What happens if one site check crashes?

The `launch_module` wrapper in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) catches all exceptions and converts them to standardized result dictionaries. A failing Twitter check won't interrupt your Amazon or GitHub probes—the error is logged and the nursery continues.

### Is the shared `httpx.AsyncClient` thread-safe?

Yes, `httpx.AsyncClient` is designed for async contexts and safely handles concurrent requests from multiple coroutines. Holehe's single-threaded Trio approach means no actual thread contention occurs—cooperative multitasking handles the interleaving.