How Website Modules Are Dynamically Imported in Holehe: A Complete Technical Guide

Holehe uses Python's pkgutil.walk_packages and importlib to automatically discover and load every Python file under holehe.modules at runtime, eliminating the need for hard-coded service lists.

Holehe is an email footprinting tool that checks hundreds of websites for account existence. Rather than maintaining a static registry of services, it employs a dynamic module import system that discovers website modules automatically. This article examines the exact mechanism used in the megadose/holehe repository, walking through the three-stage pipeline defined in holehe/core.py.

Stage 1: Automatic Package Discovery with pkgutil

The foundation of Holehe's extensibility lies in import_submodules, defined at line 37 of holehe/core.py. This function traverses the entire holehe.modules package tree using pkgutil.walk_packages.


# From holehe/core.py

import importlib
import pkgutil

def import_submodules(package, recursive=True):
    if isinstance(package, str):
        package = importlib.import_module(package)   # "holehe.modules"

    results = {}
    for loader, name, is_pkg in pkgutil.walk_packages(package.__path__):
        full_name = f"{package.__name__}.{name}"
        results[full_name] = importlib.import_module(full_name)
        if recursive and is_pkg:
            results.update(import_submodules(full_name))
    return results

The function returns a dictionary mapping fully qualified module names (e.g., holehe.modules.social_media.twitter) to imported module objects. This approach means new services are automatically discoverable immediately after dropping a .py file into the appropriate subdirectory—no configuration changes required.

Stage 2: Extracting Callable Entry Points with get_functions

Raw module objects aren't directly executable. The get_functions utility at line 50 of holehe/core.py transforms the discovery results into callable coroutines.


# From holehe/core.py

def get_functions(modules, args=None):
    websites = []
    for module in modules:
        if len(module.split(".")) > 3:                # Keep only leaf modules

            mod_obj = modules[module]
            site = module.split(".")[-1]              # Extract "twitter"

            
            # Optional: filter password-recovery-heavy sites

            if args and args.nopasswordrecovery:
                if "adobe" in str(mod_obj.__dict__.get(site, "")):
                    continue
            
            websites.append(mod_obj.__dict__[site])   # The coroutine itself

    return websites

The filtering logic len(module.split(".")) > 3 ensures only actual website implementation modules are retained, excluding intermediate package directories. Each module exports a coroutine matching its filename—twitter.py contains async def twitter(email, client, out).

Stage 3: Asynchronous Execution via trio

The maincore orchestration function (lines 24–31 in holehe/core.py) schedules all discovered modules for concurrent execution:


# From holehe/core.py

async def maincore(email, args):
    modules = import_submodules("holehe.modules")
    websites = get_functions(modules, args)
    
    async with httpx.AsyncClient() as client:
        async with trio.open_nursery() as nursery:
            for website in websites:
                nursery.start_soon(
                    launch_module, 
                    website, 
                    email, 
                    client,
                    out  # Results list

                )

Each website module receives three arguments:

  • email — the target email address to check
  • client — a shared httpx.AsyncClient for connection pooling
  • out — a list to append result dictionaries

The launch_module wrapper (also in holehe/core.py) provides standardized error handling, ensuring failed checks don't crash the entire scan:

async def launch_module(module, email, client, out):
    try:
        await module(email, client, out)
    except Exception:
        out.append({
            "name": module.__name__,
            "domain": "unknown",
            "rateLimit": False,
            "error": True,
            "exists": False,
            "emailrecovery": None,
            "phoneNumber": None,
            "others": None,
        })

Module Signature Requirements

For automatic discovery to work, every website module must follow a strict contract:

Requirement Description
File location holehe/modules/<category>/<service>.py
Function name Must match the filename (e.g., twitter in twitter.py)
Coroutine async def service_name(email, client, out)
Result format Append dict to out with standardized keys

Example implementation from holehe/modules/social_media/twitter.py:

async def twitter(email, client, out):
    data = {
        "email": email,
        "username": "",
        "password": ""
    }
    headers = {...}
    try:
        r = await client.post(
            "https://api.twitter.com/1.1/users/email_available.json",
            data=data,
            headers=headers
        )
        # ... parsing logic ...

        out.append({...result_dict...})
    except Exception as e:
        # Let launch_module handle via exception propagation

        raise

Why This Architecture Matters

Zero-configuration extensibility — Contributors add services by creating a single file with a named coroutine. No registry edits, no imports to update.

Category organization — The holehe.modules package uses subpackages like social_media, transport, and shopping for logical grouping. pkgutil.walk_packages handles nested discovery recursively.

Hot-reload friendly — Since imports happen at runtime, restarting the application picks up new modules immediately without reinstallation.

Summary

  • import_submodules in holehe/core.py uses pkgutil.walk_packages to discover every Python file under holehe.modules
  • get_functions filters leaf modules and extracts coroutines matching filenames
  • maincore schedules all discovered modules concurrently via trio
  • Module contract requires: coroutine named after file, three parameters (email, client, out), append-only result handling
  • Adding services requires only dropping a .py file in the correct subdirectory

Frequently Asked Questions

What Python standard library modules enable Holehe's dynamic imports?

Holehe relies on pkgutil for package tree walking and importlib for programmatic module importing. Specifically, pkgutil.walk_packages yields (loader, name, is_pkg) tuples for every module found, and importlib.import_module converts string names into live module objects.

Can I exclude specific website modules from a scan?

Yes. The get_functions function accepts an args parameter that supports the --no-password-recovery flag. When enabled, it filters out modules containing "adobe" in their string representation, as Adobe's password recovery mechanism is particularly rate-limit-prone.

How does Holehe handle errors in individual website modules?

The launch_module wrapper in holehe/core.py catches all exceptions and appends a standardized error entry to the results list. This ensures one failing module doesn't interrupt checks against other services, and the failure is recorded with error: True in the output.

Why use trio instead of asyncio for concurrency?

The Holehe codebase uses trio.open_nursery() for structured concurrency. This pattern creates a scope where all spawned tasks complete before the nursery exits, with automatic error propagation and cancellation. The nursery.start_soon() calls schedule each website module as a separate task without blocking.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →