How Holehe Aggregates Results from Different Modules: Core Architecture Explained

Holehe aggregates results by dynamically importing every service module in holehe/modules, executing them concurrently through Trio nurseries, and collecting standardized dictionaries into a shared mutable list that is sorted alphabetically and formatted for terminal or CSV output.

Holehe is an open-source OSINT tool that verifies email registration across hundreds of platforms by querying individual services. Understanding how Holehe aggregates results from different modules reveals a modular architecture built on dynamic imports, structured concurrency, and strict data contracts. This design allows the tool to scale horizontally across services while maintaining fault isolation and uniform output.

Dynamic Module Discovery

The aggregation pipeline begins with automatic module enumeration. In holehe/core.py at line 37, the import_submodules("holehe.modules") function recursively walks the holehe/modules package directory to load every service implementation without hardcoded imports.

Once submodules are loaded, get_functions() extracts the callable exported by each module—specifically the function whose name matches the module name—and compiles the websites list (see holehe/core.py line 50). This dynamic approach allows contributors to add new services simply by dropping a Python file into the modules directory, with zero changes to the core aggregation logic.

Concurrent Execution via Trio Nurseries

After building the websites list, Holehe leverages structured concurrency through Python’s Trio library. An asynchronous nursery spawns launch_module for every entry in the sites list. Inside launch_module, the module is invoked with the signature:

await module(email, client, out)

This call occurs at line 66 of holehe/core.py, where email is the target address, client is a shared asynchronous HTTP session, and out is the mutable result accumulator. By running all modules concurrently within a single nursery scope, Holehe maximizes throughput while preserving the ability to cancel or manage the entire operation tree as a unit.

The Standardized Result Contract

Every module must return a standardized dictionary with a fixed schema to ensure uniform aggregation. As implemented in modules like holehe/modules/software/adobe.py, the required structure is:

{
    "name": "adobe",
    "domain": "adobe.com",
    "rateLimit": False,
    "error": False,
    "exists": True,
    "emailrecovery": "recovery@example.com",
    "phoneNumber": "+1234567890",
    "others": None
}

The aggregation engine relies on these eight keys. The exists boolean indicates whether the email is registered, rateLimit flags throttling, and emailrecovery or phoneNumber capture partial data leaks. This strict contract allows holehe/core.py to process results from disparate services—whether social media, shopping, or software platforms—using identical logic.

Fault Isolation and Error Aggregation

Holehe implements defensive error handling to prevent one failing service from crashing the entire scan. Within launch_module (lines 68–78 of holehe/core.py), a try-except block wraps the module execution. If a module raises an exception, the aggregator catches it and appends a failure entry to the shared list:

{
    "name": module_name,
    "error": True,
    "exists": None,
    # ... other fields

}

This pattern ensures that network timeouts, API changes, or parsing errors in individual modules are recorded as structured data rather than propagating as uncaught exceptions. The main loop continues processing remaining modules uninterrupted.

Result Normalization and Export

Once all coroutines complete, the shared out list contains one entry per module. The aggregation finalizes with two operations in holehe/core.py:

  1. Sorting: Results are sorted alphabetically by module name using out = sorted(out, key=lambda i: i['name']) (lines 23–24), ensuring deterministic output order regardless of completion timing.
  2. Formatting: print_result() renders the aggregated data to the terminal with visual indicators ([+], [-], [x], [!]), while export_csv() (lines 54–64) optionally serializes the list to a timestamped CSV file for downstream analysis.

The holehe/instruments.py file provides a Trio-aware progress bar that updates each time a module appends to out, giving real-time visibility into the aggregation state without blocking the concurrent flow.

Summary

  • Dynamic loading: import_submodules() discovers all service modules at runtime via filesystem introspection in holehe/core.py.
  • Structured concurrency: Trio nurseries execute all modules in parallel, passing a shared mutable list for result accumulation.
  • Uniform schema: Every module must return a dictionary with keys name, domain, rateLimit, error, exists, emailrecovery, phoneNumber, and others.
  • Fault tolerance: Exception handling in launch_module converts crashes into structured error entries rather than terminating the scan.
  • Deterministic output: The final list is sorted by module name and formatted through print_result() or export_csv().

Frequently Asked Questions

How does Holehe handle a module that crashes or times out?

If a module raises an exception during execution, the launch_module wrapper in holehe/core.py catches the error and appends a result dictionary with error=True to the shared out list. This ensures the aggregation completes with a full set of results—some indicating success, others indicating failure—rather than terminating prematurely.

What data structure must a custom Holehe module return?

Each module must return a Python dictionary containing eight specific keys: name (module identifier), domain (human-readable site), rateLimit (bool), error (bool), exists (bool or None), emailrecovery (str or None), phoneNumber (str or None), and others (dict or None). This standardized contract allows the aggregator to process results from any service uniformly.

Can I run Holehe modules sequentially instead of concurrently?

The current architecture in holehe/core.py is optimized for concurrent execution via Trio nurseries. While you could technically modify the source to await modules in a loop, the design intentionally uses structured concurrency to maximize throughput across hundreds of services. Sequential execution would require significant refactoring of the launch_module invocation logic.

How does Holehe sort the final aggregated results?

After all modules complete, the aggregator sorts the shared out list alphabetically by the name field using sorted(out, key=lambda i: i['name']). This occurs in holehe/core.py before the data reaches print_result() or export_csv(), ensuring consistent ordering regardless of which network requests finished first.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →