How Holehe Discovers and Loads Site Modules at Runtime

Holehe dynamically discovers and loads site modules at runtime by scanning the holehe/modules directory with pkgutil.walk_packages, importing each submodule via importlib, and extracting check functions by naming convention rather than maintaining a static registry.

Holehe (megadose/holehe) is an OSINT tool that checks whether an email address is registered on hundreds of online services. Instead of hard-coding every service into a central list, it uses a runtime discovery mechanism in holehe/core.py to automatically find and execute site-specific check modules. This architecture allows contributors to add new services simply by dropping a properly named Python file into the modules hierarchy.

The Discovery Mechanism in holehe/core.py

The core orchestration logic resides in holehe/core.py, which implements a two-phase discovery process: first scanning the package structure, then extracting callable functions.

Scanning the Package Tree with import_submodules

The import_submodules function implements recursive package scanning using Python’s standard library utilities. It accepts the root package name "holehe.modules" and traverses the entire directory tree:

def import_submodules(package, recursive=True):
    if isinstance(package, str):
        package = importlib.import_module(package)
    results = {}
    for loader, name, is_pkg in pkgutil.walk_packages(package.__path__):
        full_name = package.__name__ + '.' + name  # e.g., holehe.modules.social_media.twitter

        results[full_name] = importlib.import_module(full_name)
        if recursive and is_pkg:
            results.update(import_submodules(full_name))
    return results
  • pkgutil.walk_packages enumerates every submodule inside holehe/modules, including nested sub-packages like holehe/modules/social_media/.
  • importlib.import_module loads each discovered module on-the-fly, populating a dictionary that maps full dotted names (e.g., holehe.modules.social_media.twitter) to module objects.
  • The recursive=True flag ensures deep traversal of nested directories, enabling categorical organization of services.

Extracting Site Check Functions with get_functions

After loading the modules, get_functions filters the dictionary to extract only the callable check functions. It assumes each leaf module exposes a single public function matching the filename:

def get_functions(modules, args=None):
    websites = []
    for module in modules:
        if len(module.split(".")) > 3:  # ignore top-level package entries

            modu = modules[module]
            site = module.split(".")[-1]  # last component is the function name

            if args is not None and args.nopasswordrecovery:
                # filter out password-recovery-heavy sites when flag is set

                if "adobe" not in str(modu.__dict__[site]) and ...
                    websites.append(modu.__dict__[site])
            else:
                websites.append(modu.__dict__[site])
    return websites

The function identifies valid site modules by checking if the dotted path has more than three components (filtering out holehe.modules itself). It then retrieves the function from the module’s __dict__ using the final component of the module path—meaning twitter.py must define a twitter() function. The optional --no-password-recovery flag removes resource-intensive checks like Adobe from the execution list.

Runtime Execution Flow

In the maincore function, Holehe orchestrates the discovery and execution phases sequentially before launching concurrent checks:

modules = import_submodules("holehe.modules")
websites = get_functions(modules, args)
…
for website in websites:
    nursery.start_soon(launch_module, website, email, client, out)

The modules dictionary contains every file found under holehe/modules/, while websites contains only the callable functions ready for execution. Holehe uses Trio to run these checks concurrently, passing each function the target email, an HTTP client, and an output collector.

Adding New Sites Without Code Changes

Because Holehe relies on filesystem discovery rather than a static registry, adding a new service requires no modifications to core.py. To add a check for example.com:

  1. Create holehe/modules/social_media/example.py (or any appropriate subdirectory).
  2. Define a function matching the filename:
def example(email, client, out):
    # implement the check logic here …

    out.append({
        "name": "example",
        "domain": "example.com",
        "rateLimit": False,
        "error": False,
        "exists": True,
        "emailrecovery": None,
        "phoneNumber": None,
        "others": None,
    })

When Holehe runs, the new module is automatically imported as holehe.modules.social_media.example and added to the execution queue via get_functions.

Summary

  • No static registry: Holehe does not maintain a hard-coded list of supported sites in core.py.
  • Filesystem-driven discovery: The import_submodules function in holehe/core.py uses pkgutil.walk_packages to scan holehe/modules at startup.
  • Convention over configuration: Each module must define a function matching its filename (e.g., twitter.py contains def twitter(...)).
  • Dynamic import chain: importlib.import_module loads each discovered file into memory during the initialization phase.
  • Concurrent execution: After discovery, get_functions compiles a list of callables that run simultaneously via Trio.

Frequently Asked Questions

How does Holehe know which function to call inside each module?

Holehe expects each site module to follow a strict naming convention according to the source code in holehe/core.py. The get_functions method extracts the last component of the module’s dotted path (e.g., twitter from holehe.modules.social_media.twitter) and looks for a function with that exact name inside the module’s __dict__. If twitter.py defines a function named twitter(), it is added to the execution list.

Can I organize site modules into subdirectories?

Yes. The import_submodules function uses pkgutil.walk_packages with recursive=True, which traverses nested packages automatically. You can create subdirectories like holehe/modules/social_media/ or holehe/modules/entertainment/, and Holehe will discover modules at every depth level. The dotted module path will reflect the directory structure (e.g., holehe.modules.social_media.instagram).

What happens if I add a malformed module to the directory?

If a file exists under holehe/modules/ but does not define the expected function, get_functions will still attempt to access modu.__dict__[site] based on the filename. This will raise a KeyError at runtime when Holehe tries to extract the callable. The discovery mechanism imports all Python files indiscriminately, so every module must implement the standard function signature or be excluded by the length check (len(module.split(".")) > 3).

Does Holehe support excluding specific sites from the scan?

Yes. The get_functions method accepts an args parameter that checks for the nopasswordrecovery flag. When this flag is present, the function filters out specific resource-intensive modules (such as Adobe) by checking if their string representation contains excluded keywords. You can extend this logic to add custom filtering criteria without modifying the discovery mechanism itself.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →