How Holehe Discovers and Dynamically Loads Its Modules: Automated Plugin Architecture Explained

Holehe leverages Python's pkgutil.walk_packages and importlib to recursively scan, import, and execute async service-checking modules at runtime, enabling a plug-and-play architecture where new services require zero configuration changes.

The megadose/holehe repository implements a sophisticated extensibility system that automatically discovers and loads email-checking capabilities across dozens of web services. By utilizing Python's built-in import machinery rather than hardcoded registries, Holehe allows contributors to add new platforms simply by placing a properly structured Python file into the holehe/modules/ directory tree.

The Three-Phase Discovery Pipeline

Holehe's dynamic loading mechanism operates through three coordinated phases implemented in holehe/core.py, transforming the static file system into a living registry of asynchronous checkers.

Phase 1: Recursive Package Traversal with import_submodules

The discovery process begins in holehe/core.py at lines 37-47 with the import_submodules function. This routine uses pkgutil.walk_packages to iterate over the holehe.modules package path, yielding every sub-package and module discovered.

For each item found, the function constructs the full dotted name (e.g., holehe.modules.social_media.twitter) and imports it using importlib.import_module. When the walker encounters a sub-package (indicated by is_pkg=True), the function recurses into that namespace to capture nested modules. This recursive walk ensures that categories like social_media/ and software/ are fully explored without manual path enumeration.

Phase 2: Callable Extraction via get_functions

Once all modules reside in memory, get_functions (defined at lines 50-63 in holehe/core.py) filters the raw import dictionary to identify executable service checkers. The function applies a strict structural filter: only module paths containing more than three dotted components qualify (e.g., holehe.modules.<category>.<module>).

For each qualifying leaf module, the logic extracts the final component of the dotted path—which corresponds to the filename without extension—and retrieves that named attribute from the module's namespace. This convention enforces that the async function inside twitter.py must be named twitter, creating a predictable mapping that get_functions uses to populate the websites list with callable objects.

Phase 3: Asynchronous Execution in maincore

The orchestration layer resides in maincore at lines 105-108 of holehe/core.py. Here, the populated websites list enters a Trio nursery, where launch_module concurrently executes each discovered async function. The launcher awaits the module's main function—passing the target email, HTTP client, and output collector—and traps any exceptions, transforming them into uniform result dictionaries that indicate rate limits, errors, or account existence.

Anatomy of a Service Module

To participate in the automatic discovery system, a module must adhere to a strict but simple contract. Consider holehe/modules/social_media/twitter.py (lines 5-8):

async def twitter(email, client, out):
    # Implementation checks if email exists on Twitter

    # Appends result dict to 'out' list

    pass

The file must contain an async function matching its basename (twitter), accepting three parameters: the target email string, an HTTP client object, and a list to collect results. When get_functions processes this file, it extracts twitter from the module's __dict__ and adds it to the execution queue. Supporting infrastructure like holehe/localuseragent.py (custom user-agent strings) and holehe/instruments.py (Trio progress meters) remain available to these dynamically loaded modules.

Implementation Details and Code Examples

The following implementation mirrors the actual logic found in holehe/core.py, demonstrating how to replicate Holehe's discovery pattern:

import importlib
import pkgutil

def import_submodules(package):
    """Recursively import every submodule in the given package."""
    if isinstance(package, str):
        package = importlib.import_module(package)
    
    modules = {}
    for loader, name, is_pkg in pkgutil.walk_packages(package.__path__):
        full_name = f"{package.__name__}.{name}"
        modules[full_name] = importlib.import_module(full_name)
        if is_pkg:  # Recurse into sub-packages

            modules.update(import_submodules(full_name))
    return modules

Extracting the callable functions requires filtering for leaf modules and matching function names to filenames:

def get_functions(modules):
    """Extract service-checking functions from imported modules."""
    functions = []
    for dotted_name, mod in modules.items():
        # Only keep leaf modules like holehe.modules.social_media.twitter

        if len(dotted_name.split(".")) > 3:
            fn_name = dotted_name.split(".")[-1]  # e.g., "twitter"

            if hasattr(mod, fn_name):
                functions.append(getattr(mod, fn_name))
    return functions

Execution within an async context uses Trio for concurrency:

async def launch_module(module_func, email, client, out):
    try:
        await module_func(email, client, out)
    except Exception:
        # Uniform error handling

        out.append({
            "name": module_func.__name__,
            "domain": "...",  # Mapped elsewhere

            "rateLimit": False,
            "error": True,
            "exists": False,
        })

Summary

  • Recursive discovery: import_submodules in holehe/core.py uses pkgutil.walk_packages to map the entire holehe.modules package tree without hardcoded paths.
  • Structural filtering: get_functions enforces the holehe.modules.<category>.<module> convention by checking for more than three dotted name components.
  • Naming contract: Each service module must define an async function matching its filename (e.g., twitter in twitter.py) to be recognized.
  • Zero-configuration extensibility: Adding new services requires only creating a file in holehe/modules/<category>/; the discovery system handles registration automatically.
  • Concurrent execution: The maincore function passes discovered callables to a Trio nursery for parallel checking with standardized error handling.

Frequently Asked Questions

How does Holehe find new modules without explicit configuration?

Holehe treats the file system as the source of truth. The import_submodules function programmatically walks the holehe/modules/ directory using Python's pkgutil module, importing every Python file it encounters. Because this happens at runtime, any new file added to the tree is automatically imported and made available without edits to a master registry or configuration file.

What naming convention must service modules follow to be discovered?

Each module must reside at least three levels deep in the package hierarchy (holehe.modules.category.module_name) and contain an async function whose name exactly matches the filename. For example, holehe/modules/social_media/twitter.py must contain async def twitter(...). The get_functions logic specifically looks for this correlation using the final component of the dotted module path.

Why does Holehe require modules to have more than three components in their dotted name?

The check len(dotted_name.split(".")) > 3 serves as a filter to distinguish leaf service modules from category packages and the root holehe.modules namespace itself. This ensures that only actual service implementations (like holehe.modules.social_media.twitter) are collected, while excluding intermediate packages (like holehe.modules.social_media) that contain no executable checker functions.

How are errors handled when a dynamically loaded module fails during execution?

The launch_module function wraps each module call in a try-except block. If a module raises an exception, the error is caught and transformed into a standardized dictionary containing the function name, domain, and boolean flags indicating error: True and exists: False. This uniform error shape allows the main execution loop to continue processing other services even if one module crashes or encounters a network timeout.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →