# How Website Modules Are Dynamically Imported in Holehe: A Complete Technical Guide

> Learn how Holehe dynamically imports website modules using pkgutil and importlib for automatic discovery and loading at runtime. Explore the technical guide now.

- Repository: [Palenath/holehe](https://github.com/megadose/holehe)
- Tags: internals
- Published: 2026-08-30

---

**Holehe uses Python's `pkgutil.walk_packages` and `importlib` to automatically discover and load every Python file under `holehe.modules` at runtime, eliminating the need for hard-coded service lists.**

Holehe is an email footprinting tool that checks hundreds of websites for account existence. Rather than maintaining a static registry of services, it employs a **dynamic module import system** that discovers website modules automatically. This article examines the exact mechanism used in the megadose/holehe repository, walking through the three-stage pipeline defined in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py).

## Stage 1: Automatic Package Discovery with `pkgutil`

The foundation of Holehe's extensibility lies in `import_submodules`, defined at line 37 of [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py). This function traverses the entire `holehe.modules` package tree using **`pkgutil.walk_packages`**.

```python

# From holehe/core.py

import importlib
import pkgutil

def import_submodules(package, recursive=True):
    if isinstance(package, str):
        package = importlib.import_module(package)   # "holehe.modules"

    results = {}
    for loader, name, is_pkg in pkgutil.walk_packages(package.__path__):
        full_name = f"{package.__name__}.{name}"
        results[full_name] = importlib.import_module(full_name)
        if recursive and is_pkg:
            results.update(import_submodules(full_name))
    return results

```

The function returns a dictionary mapping **fully qualified module names** (e.g., `holehe.modules.social_media.twitter`) to imported module objects. This approach means new services are automatically discoverable immediately after dropping a `.py` file into the appropriate subdirectory—no configuration changes required.

## Stage 2: Extracting Callable Entry Points with `get_functions`

Raw module objects aren't directly executable. The **`get_functions`** utility at line 50 of [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) transforms the discovery results into callable coroutines.

```python

# From holehe/core.py

def get_functions(modules, args=None):
    websites = []
    for module in modules:
        if len(module.split(".")) > 3:                # Keep only leaf modules

            mod_obj = modules[module]
            site = module.split(".")[-1]              # Extract "twitter"

            
            # Optional: filter password-recovery-heavy sites

            if args and args.nopasswordrecovery:
                if "adobe" in str(mod_obj.__dict__.get(site, "")):
                    continue
            
            websites.append(mod_obj.__dict__[site])   # The coroutine itself

    return websites

```

The filtering logic `len(module.split(".")) > 3` ensures only actual **website implementation modules** are retained, excluding intermediate package directories. Each module exports a coroutine matching its filename—[`twitter.py`](https://github.com/megadose/holehe/blob/main/twitter.py) contains `async def twitter(email, client, out)`.

## Stage 3: Asynchronous Execution via `trio`

The **`maincore`** orchestration function (lines 24–31 in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py)) schedules all discovered modules for concurrent execution:

```python

# From holehe/core.py

async def maincore(email, args):
    modules = import_submodules("holehe.modules")
    websites = get_functions(modules, args)
    
    async with httpx.AsyncClient() as client:
        async with trio.open_nursery() as nursery:
            for website in websites:
                nursery.start_soon(
                    launch_module, 
                    website, 
                    email, 
                    client,
                    out  # Results list

                )

```

Each website module receives three arguments:
- `email` — the target email address to check
- `client` — a shared `httpx.AsyncClient` for connection pooling
- `out` — a list to append result dictionaries

The `launch_module` wrapper (also in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py)) provides standardized error handling, ensuring failed checks don't crash the entire scan:

```python
async def launch_module(module, email, client, out):
    try:
        await module(email, client, out)
    except Exception:
        out.append({
            "name": module.__name__,
            "domain": "unknown",
            "rateLimit": False,
            "error": True,
            "exists": False,
            "emailrecovery": None,
            "phoneNumber": None,
            "others": None,
        })

```

## Module Signature Requirements

For automatic discovery to work, every website module must follow a strict contract:

| Requirement | Description |
|-------------|-------------|
| **File location** | `holehe/modules/<category>/<service>.py` |
| **Function name** | Must match the filename (e.g., `twitter` in [`twitter.py`](https://github.com/megadose/holehe/blob/main/twitter.py)) |
| **Coroutine** | `async def service_name(email, client, out)` |
| **Result format** | Append `dict` to `out` with standardized keys |

Example implementation from [`holehe/modules/social_media/twitter.py`](https://github.com/megadose/holehe/blob/main/holehe/modules/social_media/twitter.py):

```python
async def twitter(email, client, out):
    data = {
        "email": email,
        "username": "",
        "password": ""
    }
    headers = {...}
    try:
        r = await client.post(
            "https://api.twitter.com/1.1/users/email_available.json",
            data=data,
            headers=headers
        )
        # ... parsing logic ...

        out.append({...result_dict...})
    except Exception as e:
        # Let launch_module handle via exception propagation

        raise

```

## Why This Architecture Matters

**Zero-configuration extensibility** — Contributors add services by creating a single file with a named coroutine. No registry edits, no imports to update.

**Category organization** — The `holehe.modules` package uses subpackages like `social_media`, `transport`, and `shopping` for logical grouping. `pkgutil.walk_packages` handles nested discovery recursively.

**Hot-reload friendly** — Since imports happen at runtime, restarting the application picks up new modules immediately without reinstallation.

## Summary

- **`import_submodules`** in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) uses `pkgutil.walk_packages` to discover every Python file under `holehe.modules`
- **`get_functions`** filters leaf modules and extracts coroutines matching filenames
- **`maincore`** schedules all discovered modules concurrently via `trio`
- **Module contract** requires: coroutine named after file, three parameters (`email`, `client`, `out`), append-only result handling
- **Adding services** requires only dropping a `.py` file in the correct subdirectory

## Frequently Asked Questions

### What Python standard library modules enable Holehe's dynamic imports?

Holehe relies on **`pkgutil`** for package tree walking and **`importlib`** for programmatic module importing. Specifically, `pkgutil.walk_packages` yields `(loader, name, is_pkg)` tuples for every module found, and `importlib.import_module` converts string names into live module objects.

### Can I exclude specific website modules from a scan?

Yes. The `get_functions` function accepts an `args` parameter that supports the `--no-password-recovery` flag. When enabled, it filters out modules containing "adobe" in their string representation, as Adobe's password recovery mechanism is particularly rate-limit-prone.

### How does Holehe handle errors in individual website modules?

The **`launch_module`** wrapper in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) catches all exceptions and appends a standardized error entry to the results list. This ensures one failing module doesn't interrupt checks against other services, and the failure is recorded with `error: True` in the output.

### Why use `trio` instead of `asyncio` for concurrency?

The Holehe codebase uses **`trio.open_nursery()`** for structured concurrency. This pattern creates a scope where all spawned tasks complete before the nursery exits, with automatic error propagation and cancellation. The `nursery.start_soon()` calls schedule each website module as a separate task without blocking.