# How Holehe Aggregates Module Results: A Complete Guide to the Core Pipeline

> Discover how Holehe aggregates module results using concurrent async tasks. Learn about the core pipeline for efficient data processing and export.

- Repository: [Palenath/holehe](https://github.com/megadose/holehe)
- Tags: internals
- Published: 2026-08-31

---

**Holehe aggregates results from its individual modules by running concurrent async tasks that push uniform result dictionaries into a shared list, then sorting and formatting that list for display or CSV export.**

The open-source OSINT tool [megadose/holehe](https://github.com/megadose/holehe) discovers email registrations across thousands of web services through a carefully orchestrated aggregation system. This article explains exactly how Holehe collects, combines, and presents results from its modular service checks.

## How Holehe Module Aggregation Works

The aggregation pipeline lives entirely in **[`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py)** and follows six distinct stages from module discovery to final output.

### 1. Module Discovery via `import_submodules()`

Holehe begins by dynamically discovering every available service module. The `import_submodules()` function walks the **`holehe.modules`** package and imports each Python file containing a service check.

```python

# From core.py lines 37-47

modules = import_submodules("holehe.modules")

```

This approach allows the tool to support new services automatically—simply adding a `.py` file to the modules directory makes it available without modifying core logic.

### 2. Callable Extraction via `get_functions()`

After discovery, `get_functions()` extracts the public async function from each imported module. These functions—named after their service like `google`, `amazon`, `twitter`—populate a list called **`websites`**.

```python

# From core.py lines 50-63

website_funcs = get_functions(modules)

```

Each function follows the same signature: it accepts an email, an HTTP client, and appends its result to a shared collection.

### 3. Concurrent Execution with Trio

Holehe launches all service checks in parallel using a **Trio nursery**. For every module in `websites`, it starts a task calling **`launch_module()`** with:

- The target email address
- A shared **`httpx.AsyncClient`** for connection pooling
- A shared list **`out`** for result collection

```python

# From core.py lines 18-21 and 66-78

async with trio.open_nursery() as nursery:
    for website in websites:
        nursery.start_soon(launch_module, website, email, client, out)

```

This concurrent approach enables scanning hundreds of services in seconds rather than minutes.

### 4. Individual Result Collection

Each module appends a **standardized dictionary** to `out` with these fields:

| Field | Description |
|-------|-------------|
| `name` | Service identifier (e.g., "google", "github") |
| `exists` | Boolean indicating if email is registered |
| `rateLimit` | Boolean if rate-limited during check |
| `emailrecovery` | Recovery email if exposed by service |
| `phoneNumber` | Phone number if exposed by service |
| `others` | Additional metadata from service response |

A **successful check** returns `exists=True` or `exists=False` with optional extra fields. If a module raises an exception, `launch_module()` catches it and records an error entry:

```python

# Error handling in launch_module() lines 71-78

out.append({
    "name": module.__name__,
    "exists": False,
    "rateLimit": False,
    "error": True  # Indicates exception occurred

})

```

### 5. Sorting and Terminal Rendering

After all tasks complete, the `out` list undergoes **alphabetical sorting** by module name (`i['name']`), then passes to `print_result()`. This function applies color-coded symbols for immediate visual parsing:

- **`[+]`** — Account exists
- **`[-]`** — Account does not exist
- **`[!]`** — Rate limited or error occurred
- **`[x]`** — Other issues

```python

# From core.py lines 22-24

out.sort(key=lambda i: i["name"])
print_result(out, args, start_time, email)

```

### 6. Optional CSV Export

When invoked with `--csv`, Holehe writes the raw `out` list to a CSV file via `export_csv()`:

```bash
holehe target@email.com --csv results.csv

```

The CSV contains the complete dictionary data, enabling further analysis in spreadsheet tools or custom scripts.

## Practical Usage Examples

### Command-Line Aggregation

```bash

# Standard execution with colorized output

holehe alice@example.com

# Filter to show only existing accounts

holehe alice@example.com --only-used

# Generate CSV alongside terminal output

holehe alice@example.com -C ./results/

```

### Programmatic Aggregation

For custom workflows, replicate Holehe's aggregation pipeline directly:

```python
import asyncio
import httpx
from holehe.core import import_submodules, get_functions, launch_module

async def aggregate_results(email: str):
    # Stage 1: Discover and extract modules

    modules = import_submodules("holehe.modules")
    website_funcs = get_functions(modules)
    
    # Stage 2: Initialize shared resources

    client = httpx.AsyncClient(timeout=10)
    out = []
    
    # Stage 3: Execute concurrently (asyncio version)

    await asyncio.gather(*[
        launch_module(func, email, client, out)
        for func in website_funcs
    ])
    
    await client.aclose()
    
    # Stage 4: Sort and return

    out.sort(key=lambda i: i["name"])
    return out

# Execute

results = asyncio.run(aggregate_results("alice@example.com"))
print(f"Checked {len(results)} services")

```

This programmatic approach guarantees **identical behavior to the CLI** by using the same `launch_module` logic and sorting mechanism.

## Key Technical Components

| File | Purpose | Critical Functions |
|------|---------|-------------------|
| [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) | Orchestrates entire aggregation pipeline | `import_submodules()`, `get_functions()`, `launch_module()`, `print_result()` |
| `holehe/modules/**/*.py` | Individual service implementations | One async function per service |
| [`holehe/instruments.py`](https://github.com/megadose/holehe/blob/main/holehe/instruments.py) | Progress bar during execution | Trio instrumentation |
| [`holehe/localuseragent.py`](https://github.com/megadose/holehe/blob/main/holehe/localuseragent.py) | Realistic User-Agent rotation | HTTP header management |

## Summary

- **Dynamic discovery** in [`core.py`](https://github.com/megadose/holehe/blob/main/core.py) automatically loads all service modules without hardcoded lists
- **Concurrent execution** via Trio nursery enables efficient parallel checking of hundreds of services
- **Uniform result dictionaries** pushed to a shared `out` list guarantee consistent data structure regardless of service
- **Automatic error capture** ensures partial results remain usable even when individual modules fail
- **Post-processing sorting** and formatting produces human-readable output and machine-parseable CSV exports

## Frequently Asked Questions

### What happens if a Holehe module crashes during execution?

The `launch_module()` wrapper in [`core.py`](https://github.com/megadose/holehe/blob/main/core.py) catches all exceptions and appends an error dictionary to the results. The aggregation continues with remaining modules, and the failed service appears in output with `"error": True`.

### Can I control how many modules run simultaneously?

The official implementation uses Trio's default nursery behavior. For custom concurrency limits, modify the nursery configuration or replace with `asyncio.Semaphore` as shown in the programmatic example.

### Why does Holehe sort results alphabetically rather than by outcome?

Alphabetical sorting provides **deterministic, reproducible output** across runs. This consistency aids debugging, diffing results, and programmatic parsing where service position matters more than success/failure grouping.