How Holehe Handles Concurrent Tasks for Scanning Websites

Holehe uses Trio-based structured concurrency to run hundreds of email checks in parallel, with each site probe implemented as an async function executed through a shared nursery scope.

Email reconnaissance tool Holehe by megadose needs to query dozens—sometimes hundreds—of websites to check if an email address exists on each platform. Doing this sequentially would take minutes; the project solves this through asynchronous concurrent execution powered by the Trio library. Every site check runs as a non-blocking coroutine, coordinated through Trio's nursery pattern for clean parallelism without callback hell.

Async Module Functions for Non-Blocking Checks

Each supported service lives in holehe/modules/ as a standalone file containing one async def function. These functions never block the event loop while waiting for HTTP responses.

Take the Twitter module as a representative example. In holehe/modules/social_media/twitter.py, the function signature is:

async def twitter(email, client, out):
    ...

The parameters are consistent across all modules:

  • email – the target address to investigate
  • client – a shared httpx.AsyncClient instance for HTTP requests
  • out – a shared list where results are appended

This design means twitter() can await its HTTP requests without stalling other site checks. Instagram, Amazon, GitHub, and 100+ other services follow this identical pattern.

Collecting and Scheduling Tasks

The orchestration logic lives in holehe/core.py. Two utility functions gather all available probes:

modules = import_submodules('holehe.modules')      # L50-L55

websites = get_functions(modules)                  # L56-L63

The websites list contains references to every async site-check function. This collection happens once at startup, then the same functions are reused across concurrent invocations.

Progress Tracking with Trio Instruments

Before launching tasks, Holehe installs a custom Trio instrument that hooks into task lifecycle events. The TrioProgress class in holehe/instruments.py updates a tqdm progress bar each time any site check completes:

instrument = TrioProgress(total=len(websites))
trio.lowlevel.add_instrument(instrument)

This avoids cluttering the scanning logic with progress callbacks—the instrumentation stays orthogonal to the actual HTTP work.

Concurrent Execution via Trio Nursery

The core concurrency mechanism appears in holehe/core.py at lines 217-221:

async with trio.open_nursery() as nursery:
    for module in websites:
        nursery.start_soon(
            launch_module, module, email, client, out
        )

Key components of this pattern:

  • trio.open_nursery() – creates a scope that automatically waits for all child tasks
  • nursery.start_soon() – schedules launch_module to run immediately without blocking the loop
  • launch_module – a wrapper that awaits the site-specific function and normalizes exceptions into result dictionaries

When the async with block exits, every check has finished—guaranteed by Trio's structured concurrency. No zombie tasks, no orphaned connections.

Shared HTTP Client for Efficiency

Connection pooling matters when opening hundreds of parallel requests. In holehe/core.py, a single httpx.AsyncClient is instantiated once:

async with httpx.AsyncClient(timeout=10, headers=...) as client:
    # ... nursery execution happens here

Passing this shared client to every launch_module call allows HTTP/2 connection reuse and eliminates the overhead of per-request client creation. The client remains fully async-compatible with Trio's event loop.

Result Aggregation

All async functions write to the same shared list out passed by reference. After the nursery closes, holehe/core.py processes this list for display:


# Inside the main flow

await Trio(...all checks...)  # nursery ensures everything finishes here

for result in out:
    # format and print or export to CSV

No locks are required because Python's asyncio-compatible structures and Trio's single-threaded concurrency model prevent race conditions on list appends.

How to Use Holehe's Concurrency Model

Command-Line Scanning

Run a standard concurrent scan from your terminal:


# Basic scan against all supported sites

holehe target@example.com

# Filter to only confirmed registrations

holehe target@example.com --only-used

# Export results for further analysis

holehe target@example.com --csv

Programmatic Usage

Embed Holehe's concurrency in your own applications:

import httpx
import trio
from holehe.core import import_submodules, get_functions, launch_module

async def scan_email(email: str):
    # Load all site-check functions

    modules = import_submodules('holehe.modules')
    websites = get_functions(modules)
    
    results = []
    
    async with httpx.AsyncClient(timeout=10) as client:
        async with trio.open_nursery() as nursery:
            for site_func in websites:
                nursery.start_soon(
                    launch_module, site_func, email, client, results
                )
    
    return results

# Execute with Trio's runner

trio.run(scan_email, 'someone@example.com')

Custom Progress Tracking

Reuse Holehe's instrument for your own async workflows:

from holehe.instruments import TrioProgress
import trio

async def monitored_scan(websites, email, client):
    progress = TrioProgress(total=len(websites))
    trio.lowlevel.add_instrument(progress)
    
    try:
        async with trio.open_nursery() as n:
            for func in websites:
                n.start_soon(launch_module, func, email, client, [])
    finally:
        trio.lowlevel.remove_instrument(progress)

Why Trio Over Alternatives?

Holehe chose Trio specifically for three operational advantages:

  • Structured concurrency – tasks cannot outlive their parent scope, eliminating cleanup bugs
  • Native instrumentation – clean progress hooks without monkey-patching or decorators
  • Deterministic cancellation – timeout handling that properly unwinds nested async operations

asyncio could achieve similar throughput, but Trio's semantics reduce edge-case bugs when managing hundreds of simultaneous network operations.

Summary

  • Every site check is an async def function in holehe/modules/, using non-blocking HTTP via httpx.AsyncClient
  • holehe/core.py collects these functions and executes them through trio.open_nursery() for true parallelism
  • A single shared HTTP client enables connection pooling across all concurrent requests
  • TrioProgress instrument in holehe/instruments.py provides clean progress reporting without invasive code
  • Structured concurrency guarantees all tasks complete before final results are processed

Frequently Asked Questions

Why does Holehe use Trio instead of asyncio?

Trio provides structured concurrency with stricter guarantees about task lifetimes and cancellation. The instrument API also allows progress tracking without wrapping every function call. Both libraries achieve similar performance, but Trio's semantics reduce bugs in long-running concurrent network code.

Can I adjust how many sites run simultaneously?

Trio's nursery doesn't use a fixed thread pool—coroutines yield cooperatively. However, you can partition the websites list and run multiple nurseries sequentially, or apply anyio limits if you fork the codebase. By default, all discovered modules run in parallel with whatever concurrency the event loop and OS can sustain.

What happens if one site check crashes?

The launch_module wrapper in holehe/core.py catches all exceptions and converts them to standardized result dictionaries. A failing Twitter check won't interrupt your Amazon or GitHub probes—the error is logged and the nursery continues.

Is the shared httpx.AsyncClient thread-safe?

Yes, httpx.AsyncClient is designed for async contexts and safely handles concurrent requests from multiple coroutines. Holehe's single-threaded Trio approach means no actual thread contention occurs—cooperative multitasking handles the interleaving.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →