How Holehe Dynamically Discovers and Loads Modules: A Deep Dive into Its Extensible Architecture

Holehe uses Python's pkgutil.walk_packages and importlib.import_module to dynamically discover and load service modules at runtime, eliminating hard-coded imports and enabling automatic detection of new services.

Holehe's modular design allows contributors to add new service checks simply by dropping Python files into the holehe/modules/ directory. This article explains exactly how the megadose/holehe repository implements dynamic module discovery, walking through the source code in holehe/core.py that makes this extensibility possible.

Understanding the Dynamic Module Discovery Pipeline

Holehe's discovery mechanism operates in three coordinated phases, all implemented in holehe/core.py. The system uses standard Python import utilities rather than custom directory scanning, ensuring cross-platform compatibility and proper package resolution.

Phase 1: Package Walk-through with import_submodules

The foundation of Holehe's discovery is the import_submodules function (lines 37-47 in holehe/core.py). This function leverages pkgutil.walk_packages to recursively traverse the holehe.modules package tree.

def import_submodules(package, built_in=True):
    """Import all submodules of a module, recursively."""
    if isinstance(package, str):
        package = importlib.import_module(package)
    results = {}
    for loader, name, is_pkg in pkgutil.walk_packages(package.__path__):
        full_name = package.__name__ + '.' + name
        results[full_name] = importlib.import_module(full_name)
    return results

For each discovered submodule, pkgutil.walk_packages yields:

  • loader: The loader object for the module
  • name: The short name of the submodule or subpackage
  • is_pkg: Boolean indicating whether this is a package with submodules

The function constructs the full dotted name (e.g., holehe.modules.social_media.twitter) and imports it via importlib.import_module. This recursive approach ensures that even deeply nested service categories are automatically discovered.

Phase 2: Function Extraction with get_functions

Once modules are loaded, the get_functions function (lines 50-63 in holehe/core.py) extracts the actual callable implementations. This function filters out internal packages and retrieves the async function that matches each service's name.

def get_functions(modules, args):
    """Extract service functions from imported modules."""
    websites = []
    for module_name in modules:
        parts = module_name.split('.')
        # Skip internal packages (fewer than 4 components)

        if len(parts) < 4:
            continue
        # Extract the service name from the module path

        service_name = parts[-1]
        module_obj = modules[module_name]
        # Get the async function named after the service

        if hasattr(module_obj, service_name):
            websites.append(getattr(module_obj, service_name))
    return websites

The filtering logic (len(parts) < 4) excludes package infrastructure like holehe.modules or holehe.modules.social_media itself, keeping only leaf modules that contain actual service implementations.

Phase 3: Asynchronous Execution with Trio

The extracted functions are executed concurrently using Trio's nursery pattern (lines 66-73 and 79-84 in holehe/core.py). The maincore function orchestrates this execution:

async def maincore(email, args):
    modules = import_submodules("holehe.modules")
    websites = get_functions(modules, args)
    
    client = httpx.AsyncClient()
    out = []
    
    async with trio.open_nursery() as nursery:
        for site_fn in websites:
            nursery.start_soon(launch_module, site_fn, email, client, out)
    
    return out

async def launch_module(site_fn, email, client, out):
    """Wrapper to launch a module and capture its result."""
    await site_fn(email, client, out)

This design ensures that hundreds of service checks run concurrently without blocking, with each module receiving a shared HTTP client and output collection for efficiency.

The Module Structure That Enables Discovery

For dynamic discovery to work, service modules must follow a strict naming convention. Each service module is a Python file located under holehe/modules/<category>/<service>.py that exports:

  1. An async function named identically to the module file
  2. The function signature: async def <service>(email, client, out)

Example from holehe/modules/social_media/twitter.py:

import httpx

async def twitter(email, client, out):
    """Check if email is registered on Twitter."""
    # Implementation details...

    out.append({
        "name": "twitter",
        "domain": "twitter.com",
        "method": "register",
        # ... result data

    })

The module path holehe.modules.social_media.twitter maps directly to the callable twitter, which get_functions extracts automatically. No registration boilerplate or explicit imports are required.

Complete Discovery Workflow Example

Here's how the three phases connect in practice when Holehe runs an email check:


# Phase 1: Discover all modules

>>> modules = import_submodules("holehe.modules")
>>> list(modules.keys())[:3]
[
    'holehe.modules.social_media.twitter',
    'holehe.modules.social_media.instagram',
    'holehe.modules.software.adobe'
]

# Phase 2: Extract service functions

>>> websites = get_functions(modules, args)
>>> [fn.__name__ for fn in websites[:3]]
['twitter', 'instagram', 'adobe']

# Phase 3: Execute concurrently (simplified)

>>> async with trio.open_nursery() as nursery:
...     for fn in websites:
...         nursery.start_soon(fn, "target@example.com", client, results)

Why This Architecture Matters

The dynamic discovery approach in megadose/holehe provides several advantages over static imports:

  • Zero-configuration extensibility: Adding holehe/modules/ecommerce/shopify.py with an async shopify function automatically includes it in scans
  • No merge conflicts: Contributors modify only their new service file, never touching central registry code
  • Lazy loading resilience: Import failures in one module don't crash the entire application
  • Clean dependency boundaries: Service categories map naturally to filesystem directories

According to the holehe/core.py source, this design has remained stable across versions because it relies on Python's standard library (pkgutil, importlib) rather than third-party plugin frameworks.

Summary

  • Dynamic discovery in Holehe relies on pkgutil.walk_packages in holehe/core.py to traverse the holehe.modules package tree
  • import_submodules (lines 37-47) builds a dictionary mapping full module names to imported module objects
  • get_functions (lines 50-63) filters internal packages and extracts async functions matching their module names
  • Trio concurrency executes discovered services without hard-coded orchestration logic
  • New services require only a properly named .py file under holehe/modules/ with an identically named async function

Frequently Asked Questions

How does Holehe avoid importing non-service modules?

get_functions filters modules by counting dot-separated components in the module name. Paths with fewer than 4 parts (e.g., holehe.modules.social_media) are skipped as package infrastructure, keeping only leaf modules like holehe.modules.social_media.twitter.

What happens if a new module has an import error?

importlib.import_module raises ImportError during the import_submodules phase. Holehe doesn't implement explicit error handling in the discovery code, so a broken module will crash discovery unless fixed—a trade-off for simplicity.

Can I add services to nested subdirectories?

Yes. pkgutil.walk_packages recurses automatically, so holehe/modules/productivity/collaboration/notion.py would be discovered as holehe.modules.productivity.collaboration.notion provided you include __init__.py files in each parent directory.

Why use Trio instead of asyncio for concurrent execution?

The megadose/holehe codebase standardizes on Trio for structured concurrency. The launch_module wrapper (lines 79-84) provides a uniform interface for spawning each service check as a cancelable task with proper exception propagation—critical when running dozens of network-bound checks simultaneously.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →