# How Holehe Discovers and Loads Site Modules at Runtime

> Learn how Holehe dynamically discovers and loads site modules at runtime using pkgutil and importlib. Discover check functions by convention for efficient module management.

- Repository: [Palenath/holehe](https://github.com/megadose/holehe)
- Tags: internals
- Published: 2026-09-09

---

**Holehe dynamically discovers and loads site modules at runtime by scanning the `holehe/modules` directory with `pkgutil.walk_packages`, importing each submodule via `importlib`, and extracting check functions by naming convention rather than maintaining a static registry.**

Holehe (megadose/holehe) is an OSINT tool that checks whether an email address is registered on hundreds of online services. Instead of hard-coding every service into a central list, it uses a **runtime discovery mechanism** in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) to automatically find and execute site-specific check modules. This architecture allows contributors to add new services simply by dropping a properly named Python file into the modules hierarchy.

## The Discovery Mechanism in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py)

The core orchestration logic resides in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py), which implements a two-phase discovery process: first scanning the package structure, then extracting callable functions.

### Scanning the Package Tree with `import_submodules`

The `import_submodules` function implements recursive package scanning using Python’s standard library utilities. It accepts the root package name `"holehe.modules"` and traverses the entire directory tree:

```python
def import_submodules(package, recursive=True):
    if isinstance(package, str):
        package = importlib.import_module(package)
    results = {}
    for loader, name, is_pkg in pkgutil.walk_packages(package.__path__):
        full_name = package.__name__ + '.' + name  # e.g., holehe.modules.social_media.twitter

        results[full_name] = importlib.import_module(full_name)
        if recursive and is_pkg:
            results.update(import_submodules(full_name))
    return results

```

- **`pkgutil.walk_packages`** enumerates every submodule inside `holehe/modules`, including nested sub-packages like `holehe/modules/social_media/`.
- **`importlib.import_module`** loads each discovered module on-the-fly, populating a dictionary that maps full dotted names (e.g., `holehe.modules.social_media.twitter`) to module objects.
- The `recursive=True` flag ensures deep traversal of nested directories, enabling categorical organization of services.

### Extracting Site Check Functions with `get_functions`

After loading the modules, `get_functions` filters the dictionary to extract only the callable check functions. It assumes each leaf module exposes a single public function matching the filename:

```python
def get_functions(modules, args=None):
    websites = []
    for module in modules:
        if len(module.split(".")) > 3:  # ignore top-level package entries

            modu = modules[module]
            site = module.split(".")[-1]  # last component is the function name

            if args is not None and args.nopasswordrecovery:
                # filter out password-recovery-heavy sites when flag is set

                if "adobe" not in str(modu.__dict__[site]) and ...
                    websites.append(modu.__dict__[site])
            else:
                websites.append(modu.__dict__[site])
    return websites

```

The function identifies valid site modules by checking if the dotted path has more than three components (filtering out `holehe.modules` itself). It then retrieves the function from the module’s `__dict__` using the final component of the module path—meaning [`twitter.py`](https://github.com/megadose/holehe/blob/main/twitter.py) must define a `twitter()` function. The optional `--no-password-recovery` flag removes resource-intensive checks like Adobe from the execution list.

## Runtime Execution Flow

In the `maincore` function, Holehe orchestrates the discovery and execution phases sequentially before launching concurrent checks:

```python
modules = import_submodules("holehe.modules")
websites = get_functions(modules, args)
…
for website in websites:
    nursery.start_soon(launch_module, website, email, client, out)

```

The `modules` dictionary contains every file found under `holehe/modules/`, while `websites` contains only the callable functions ready for execution. Holehe uses **Trio** to run these checks concurrently, passing each function the target email, an HTTP client, and an output collector.

## Adding New Sites Without Code Changes

Because Holehe relies on filesystem discovery rather than a static registry, adding a new service requires no modifications to [`core.py`](https://github.com/megadose/holehe/blob/main/core.py). To add a check for `example.com`:

1. Create [`holehe/modules/social_media/example.py`](https://github.com/megadose/holehe/blob/main/holehe/modules/social_media/example.py) (or any appropriate subdirectory).
2. Define a function matching the filename:

```python
def example(email, client, out):
    # implement the check logic here …

    out.append({
        "name": "example",
        "domain": "example.com",
        "rateLimit": False,
        "error": False,
        "exists": True,
        "emailrecovery": None,
        "phoneNumber": None,
        "others": None,
    })

```

When Holehe runs, the new module is automatically imported as `holehe.modules.social_media.example` and added to the execution queue via `get_functions`.

## Summary

- **No static registry**: Holehe does not maintain a hard-coded list of supported sites in [`core.py`](https://github.com/megadose/holehe/blob/main/core.py).
- **Filesystem-driven discovery**: The `import_submodules` function in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py) uses `pkgutil.walk_packages` to scan `holehe/modules` at startup.
- **Convention over configuration**: Each module must define a function matching its filename (e.g., [`twitter.py`](https://github.com/megadose/holehe/blob/main/twitter.py) contains `def twitter(...)`).
- **Dynamic import chain**: `importlib.import_module` loads each discovered file into memory during the initialization phase.
- **Concurrent execution**: After discovery, `get_functions` compiles a list of callables that run simultaneously via Trio.

## Frequently Asked Questions

### How does Holehe know which function to call inside each module?

Holehe expects each site module to follow a strict naming convention according to the source code in [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py). The `get_functions` method extracts the last component of the module’s dotted path (e.g., `twitter` from `holehe.modules.social_media.twitter`) and looks for a function with that exact name inside the module’s `__dict__`. If [`twitter.py`](https://github.com/megadose/holehe/blob/main/twitter.py) defines a function named `twitter()`, it is added to the execution list.

### Can I organize site modules into subdirectories?

Yes. The `import_submodules` function uses `pkgutil.walk_packages` with `recursive=True`, which traverses nested packages automatically. You can create subdirectories like `holehe/modules/social_media/` or `holehe/modules/entertainment/`, and Holehe will discover modules at every depth level. The dotted module path will reflect the directory structure (e.g., `holehe.modules.social_media.instagram`).

### What happens if I add a malformed module to the directory?

If a file exists under `holehe/modules/` but does not define the expected function, `get_functions` will still attempt to access `modu.__dict__[site]` based on the filename. This will raise a `KeyError` at runtime when Holehe tries to extract the callable. The discovery mechanism imports all Python files indiscriminately, so every module must implement the standard function signature or be excluded by the length check (`len(module.split(".")) > 3`).

### Does Holehe support excluding specific sites from the scan?

Yes. The `get_functions` method accepts an `args` parameter that checks for the `nopasswordrecovery` flag. When this flag is present, the function filters out specific resource-intensive modules (such as Adobe) by checking if their string representation contains excluded keywords. You can extend this logic to add custom filtering criteria without modifying the discovery mechanism itself.