How Holehe Discovers and Loads Site Modules at Runtime
Holehe dynamically discovers and loads site modules at runtime by scanning the holehe/modules directory with pkgutil.walk_packages, importing each submodule via importlib, and extracting check functions by naming convention rather than maintaining a static registry.
Holehe (megadose/holehe) is an OSINT tool that checks whether an email address is registered on hundreds of online services. Instead of hard-coding every service into a central list, it uses a runtime discovery mechanism in holehe/core.py to automatically find and execute site-specific check modules. This architecture allows contributors to add new services simply by dropping a properly named Python file into the modules hierarchy.
The Discovery Mechanism in holehe/core.py
The core orchestration logic resides in holehe/core.py, which implements a two-phase discovery process: first scanning the package structure, then extracting callable functions.
Scanning the Package Tree with import_submodules
The import_submodules function implements recursive package scanning using Python’s standard library utilities. It accepts the root package name "holehe.modules" and traverses the entire directory tree:
def import_submodules(package, recursive=True):
if isinstance(package, str):
package = importlib.import_module(package)
results = {}
for loader, name, is_pkg in pkgutil.walk_packages(package.__path__):
full_name = package.__name__ + '.' + name # e.g., holehe.modules.social_media.twitter
results[full_name] = importlib.import_module(full_name)
if recursive and is_pkg:
results.update(import_submodules(full_name))
return results
pkgutil.walk_packagesenumerates every submodule insideholehe/modules, including nested sub-packages likeholehe/modules/social_media/.importlib.import_moduleloads each discovered module on-the-fly, populating a dictionary that maps full dotted names (e.g.,holehe.modules.social_media.twitter) to module objects.- The
recursive=Trueflag ensures deep traversal of nested directories, enabling categorical organization of services.
Extracting Site Check Functions with get_functions
After loading the modules, get_functions filters the dictionary to extract only the callable check functions. It assumes each leaf module exposes a single public function matching the filename:
def get_functions(modules, args=None):
websites = []
for module in modules:
if len(module.split(".")) > 3: # ignore top-level package entries
modu = modules[module]
site = module.split(".")[-1] # last component is the function name
if args is not None and args.nopasswordrecovery:
# filter out password-recovery-heavy sites when flag is set
if "adobe" not in str(modu.__dict__[site]) and ...
websites.append(modu.__dict__[site])
else:
websites.append(modu.__dict__[site])
return websites
The function identifies valid site modules by checking if the dotted path has more than three components (filtering out holehe.modules itself). It then retrieves the function from the module’s __dict__ using the final component of the module path—meaning twitter.py must define a twitter() function. The optional --no-password-recovery flag removes resource-intensive checks like Adobe from the execution list.
Runtime Execution Flow
In the maincore function, Holehe orchestrates the discovery and execution phases sequentially before launching concurrent checks:
modules = import_submodules("holehe.modules")
websites = get_functions(modules, args)
…
for website in websites:
nursery.start_soon(launch_module, website, email, client, out)
The modules dictionary contains every file found under holehe/modules/, while websites contains only the callable functions ready for execution. Holehe uses Trio to run these checks concurrently, passing each function the target email, an HTTP client, and an output collector.
Adding New Sites Without Code Changes
Because Holehe relies on filesystem discovery rather than a static registry, adding a new service requires no modifications to core.py. To add a check for example.com:
- Create
holehe/modules/social_media/example.py(or any appropriate subdirectory). - Define a function matching the filename:
def example(email, client, out):
# implement the check logic here …
out.append({
"name": "example",
"domain": "example.com",
"rateLimit": False,
"error": False,
"exists": True,
"emailrecovery": None,
"phoneNumber": None,
"others": None,
})
When Holehe runs, the new module is automatically imported as holehe.modules.social_media.example and added to the execution queue via get_functions.
Summary
- No static registry: Holehe does not maintain a hard-coded list of supported sites in
core.py. - Filesystem-driven discovery: The
import_submodulesfunction inholehe/core.pyusespkgutil.walk_packagesto scanholehe/modulesat startup. - Convention over configuration: Each module must define a function matching its filename (e.g.,
twitter.pycontainsdef twitter(...)). - Dynamic import chain:
importlib.import_moduleloads each discovered file into memory during the initialization phase. - Concurrent execution: After discovery,
get_functionscompiles a list of callables that run simultaneously via Trio.
Frequently Asked Questions
How does Holehe know which function to call inside each module?
Holehe expects each site module to follow a strict naming convention according to the source code in holehe/core.py. The get_functions method extracts the last component of the module’s dotted path (e.g., twitter from holehe.modules.social_media.twitter) and looks for a function with that exact name inside the module’s __dict__. If twitter.py defines a function named twitter(), it is added to the execution list.
Can I organize site modules into subdirectories?
Yes. The import_submodules function uses pkgutil.walk_packages with recursive=True, which traverses nested packages automatically. You can create subdirectories like holehe/modules/social_media/ or holehe/modules/entertainment/, and Holehe will discover modules at every depth level. The dotted module path will reflect the directory structure (e.g., holehe.modules.social_media.instagram).
What happens if I add a malformed module to the directory?
If a file exists under holehe/modules/ but does not define the expected function, get_functions will still attempt to access modu.__dict__[site] based on the filename. This will raise a KeyError at runtime when Holehe tries to extract the callable. The discovery mechanism imports all Python files indiscriminately, so every module must implement the standard function signature or be excluded by the length check (len(module.split(".")) > 3).
Does Holehe support excluding specific sites from the scan?
Yes. The get_functions method accepts an args parameter that checks for the nopasswordrecovery flag. When this flag is present, the function filters out specific resource-intensive modules (such as Adobe) by checking if their string representation contains excluded keywords. You can extend this logic to add custom filtering criteria without modifying the discovery mechanism itself.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →