How Holehe Handles Concurrent Tasks for Scanning Websites
Holehe uses Trio-based structured concurrency to run hundreds of email checks in parallel, with each site probe implemented as an async function executed through a shared nursery scope.
Email reconnaissance tool Holehe by megadose needs to query dozens—sometimes hundreds—of websites to check if an email address exists on each platform. Doing this sequentially would take minutes; the project solves this through asynchronous concurrent execution powered by the Trio library. Every site check runs as a non-blocking coroutine, coordinated through Trio's nursery pattern for clean parallelism without callback hell.
Async Module Functions for Non-Blocking Checks
Each supported service lives in holehe/modules/ as a standalone file containing one async def function. These functions never block the event loop while waiting for HTTP responses.
Take the Twitter module as a representative example. In holehe/modules/social_media/twitter.py, the function signature is:
async def twitter(email, client, out):
...
The parameters are consistent across all modules:
email– the target address to investigateclient– a sharedhttpx.AsyncClientinstance for HTTP requestsout– a shared list where results are appended
This design means twitter() can await its HTTP requests without stalling other site checks. Instagram, Amazon, GitHub, and 100+ other services follow this identical pattern.
Collecting and Scheduling Tasks
The orchestration logic lives in holehe/core.py. Two utility functions gather all available probes:
modules = import_submodules('holehe.modules') # L50-L55
websites = get_functions(modules) # L56-L63
The websites list contains references to every async site-check function. This collection happens once at startup, then the same functions are reused across concurrent invocations.
Progress Tracking with Trio Instruments
Before launching tasks, Holehe installs a custom Trio instrument that hooks into task lifecycle events. The TrioProgress class in holehe/instruments.py updates a tqdm progress bar each time any site check completes:
instrument = TrioProgress(total=len(websites))
trio.lowlevel.add_instrument(instrument)
This avoids cluttering the scanning logic with progress callbacks—the instrumentation stays orthogonal to the actual HTTP work.
Concurrent Execution via Trio Nursery
The core concurrency mechanism appears in holehe/core.py at lines 217-221:
async with trio.open_nursery() as nursery:
for module in websites:
nursery.start_soon(
launch_module, module, email, client, out
)
Key components of this pattern:
trio.open_nursery()– creates a scope that automatically waits for all child tasksnursery.start_soon()– scheduleslaunch_moduleto run immediately without blocking the looplaunch_module– a wrapper that awaits the site-specific function and normalizes exceptions into result dictionaries
When the async with block exits, every check has finished—guaranteed by Trio's structured concurrency. No zombie tasks, no orphaned connections.
Shared HTTP Client for Efficiency
Connection pooling matters when opening hundreds of parallel requests. In holehe/core.py, a single httpx.AsyncClient is instantiated once:
async with httpx.AsyncClient(timeout=10, headers=...) as client:
# ... nursery execution happens here
Passing this shared client to every launch_module call allows HTTP/2 connection reuse and eliminates the overhead of per-request client creation. The client remains fully async-compatible with Trio's event loop.
Result Aggregation
All async functions write to the same shared list out passed by reference. After the nursery closes, holehe/core.py processes this list for display:
# Inside the main flow
await Trio(...all checks...) # nursery ensures everything finishes here
for result in out:
# format and print or export to CSV
No locks are required because Python's asyncio-compatible structures and Trio's single-threaded concurrency model prevent race conditions on list appends.
How to Use Holehe's Concurrency Model
Command-Line Scanning
Run a standard concurrent scan from your terminal:
# Basic scan against all supported sites
holehe target@example.com
# Filter to only confirmed registrations
holehe target@example.com --only-used
# Export results for further analysis
holehe target@example.com --csv
Programmatic Usage
Embed Holehe's concurrency in your own applications:
import httpx
import trio
from holehe.core import import_submodules, get_functions, launch_module
async def scan_email(email: str):
# Load all site-check functions
modules = import_submodules('holehe.modules')
websites = get_functions(modules)
results = []
async with httpx.AsyncClient(timeout=10) as client:
async with trio.open_nursery() as nursery:
for site_func in websites:
nursery.start_soon(
launch_module, site_func, email, client, results
)
return results
# Execute with Trio's runner
trio.run(scan_email, 'someone@example.com')
Custom Progress Tracking
Reuse Holehe's instrument for your own async workflows:
from holehe.instruments import TrioProgress
import trio
async def monitored_scan(websites, email, client):
progress = TrioProgress(total=len(websites))
trio.lowlevel.add_instrument(progress)
try:
async with trio.open_nursery() as n:
for func in websites:
n.start_soon(launch_module, func, email, client, [])
finally:
trio.lowlevel.remove_instrument(progress)
Why Trio Over Alternatives?
Holehe chose Trio specifically for three operational advantages:
- Structured concurrency – tasks cannot outlive their parent scope, eliminating cleanup bugs
- Native instrumentation – clean progress hooks without monkey-patching or decorators
- Deterministic cancellation – timeout handling that properly unwinds nested async operations
asyncio could achieve similar throughput, but Trio's semantics reduce edge-case bugs when managing hundreds of simultaneous network operations.
Summary
- Every site check is an
async deffunction inholehe/modules/, using non-blocking HTTP viahttpx.AsyncClient holehe/core.pycollects these functions and executes them throughtrio.open_nursery()for true parallelism- A single shared HTTP client enables connection pooling across all concurrent requests
TrioProgressinstrument inholehe/instruments.pyprovides clean progress reporting without invasive code- Structured concurrency guarantees all tasks complete before final results are processed
Frequently Asked Questions
Why does Holehe use Trio instead of asyncio?
Trio provides structured concurrency with stricter guarantees about task lifetimes and cancellation. The instrument API also allows progress tracking without wrapping every function call. Both libraries achieve similar performance, but Trio's semantics reduce bugs in long-running concurrent network code.
Can I adjust how many sites run simultaneously?
Trio's nursery doesn't use a fixed thread pool—coroutines yield cooperatively. However, you can partition the websites list and run multiple nurseries sequentially, or apply anyio limits if you fork the codebase. By default, all discovered modules run in parallel with whatever concurrency the event loop and OS can sustain.
What happens if one site check crashes?
The launch_module wrapper in holehe/core.py catches all exceptions and converts them to standardized result dictionaries. A failing Twitter check won't interrupt your Amazon or GitHub probes—the error is logged and the nursery continues.
Is the shared httpx.AsyncClient thread-safe?
Yes, httpx.AsyncClient is designed for async contexts and safely handles concurrent requests from multiple coroutines. Holehe's single-threaded Trio approach means no actual thread contention occurs—cooperative multitasking handles the interleaving.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →