How to Integrate Holehe as a Python Library: Complete Guide with Code Examples

Yes, Holehe can be used as a Python library—its modular architecture exposes async functions for individual services and full-suite scanning through importable entry points in holehe/core.py.

Holehe, the popular email OSINT tool by megadose, is designed for dual use: command-line scanning and programmatic integration. Whether you need to check a single service like Twitter or orchestrate bulk email enumeration within your own application, Holehe's Python API gives you direct access to its async core. This guide shows you exactly how to import and call Holehe from your code, with working examples derived from the actual source.

Core Architecture: How Holehe Works as a Library

Holehe's design centers on dynamic module discovery and async execution. Understanding this architecture helps you choose the right integration pattern.

Dynamic Module Discovery

In holehe/core.py, the import_submodules() function (lines 37–47) walks the package tree using pkgutil.walk_packages() to find every service module:

modules = import_submodules("holehe.modules")
websites = get_functions(modules, args)

This means new services added to holehe/modules/ are automatically available to library users—no manual registration required.

Async Execution Pipeline

The core execution flow in holehe/core.py follows three steps:

  1. Extract callables — get_functions() (lines 50–64) converts modules to async functions
  2. Launch with concurrency — trio.open_nursery() runs all services in parallel
  3. Normalize results — launch_module() (lines 66–78) catches exceptions and formats output

Each service module implements the standard signature: async def <service>(email, client, out).

Method 1: Check a Single Service Directly

For targeted OSINT, import and call individual service functions. This avoids the overhead of scanning all 100+ sites.

import trio
import httpx

# Import the specific service you need

from holehe.modules.social_media.twitter import twitter

async def check_twitter(email: str):
    out = []
    client = httpx.AsyncClient(timeout=10)
    await twitter(email, client, out)
    await client.aclose()
    return out

# Run it

results = trio.run(check_twitter, "target@example.com")
print(results)

# [{'name': 'twitter', 'domain': 'twitter.com', 'exists': True, ...}]

Key files: holehe/modules/social_media/twitter.py contains the Twitter check logic; holehe/core.py provides the client handling patterns you should mirror.

Method 2: Run the Full CLI Pipeline Programmatically

To replicate the complete holehe command-line experience inside Python, use the maincore() entry point:

import sys
import trio
from holehe.core import maincore

# Configure arguments as if running: holehe alice@example.com --no-color

sys.argv = ["holehe", "alice@example.com", "--no-color"]

trio.run(maincore)  # Executes full scan and prints formatted table

This is equivalent to calling holehe.core.main() (lines 32–34), which parses arguments before delegating to maincore() (lines 80–99).

Method 3: Custom Integration with Full Control

For production applications, bypass the CLI wrappers and orchestrate the async pipeline yourself. This pattern—found in holehe/core.py lines 66–99—gives you complete control over result handling:

import trio
import httpx
from holehe.core import import_submodules, get_functions, launch_module

async def enumerate_email(email: str, timeout: int = 10):
    """
    Run Holehe against all services with custom result processing.
    """
    # Discover all available modules

    modules = import_submodules("holehe.modules")
    websites = get_functions(modules)  # No args = use all services

    
    # Configure HTTP client (same defaults as CLI)

    client = httpx.AsyncClient(timeout=timeout)
    results = []
    
    # Execute all services concurrently

    async with trio.open_nursery() as nursery:
        for site in websites:
            nursery.start_soon(launch_module, site, email, client, results)
    
    await client.aclose()
    
    # Sort by domain (matches CLI behavior) and return

    return sorted(results, key=lambda x: x.get("domain", ""))

# Integration example: store results in database

async def process_email(email: str):
    findings = await enumerate_email(email)
    
    for site in findings:
        if site.get("exists"):
            print(f"Found: {site['domain']} ({site.get('rateLimit', 'N/A')})")
        # Add your persistence logic here

    
    return findings

# Execute

trio.run(process_email, "bob@example.org")

Why this matters: You control timeout values, add custom result filtering, integrate with databases, or feed findings into other OSINT pipelines—without spawning subprocesses.

Key Integration Points and Source References

Component Source Location Purpose
Package root holehe/__init__.py Empty marker; enables import holehe
Core engine holehe/core.py import_submodules(), get_functions(), launch_module(), maincore()
Service modules holehe/modules/social_media/twitter.py (example) Individual async check functions
Module package holehe/modules/__init__.py Enables dynamic discovery
Utilities holehe/localuseragent.py ua variable for random User-Agent strings
Progress display holehe/instruments.py TrioProgress CLI progress bar (skip for library use)

Practical Considerations for Library Usage

Async-First Design

Holehe requires an async runtime. Use trio.run() as shown, or integrate with asyncio via trio-asyncio if your application uses a different loop.

HTTP Client Lifecycle

Always close the httpx.AsyncClient with await client.aclose() or use async with. The CLI uses a 10-second default timeout; adjust based on your reliability needs.

Rate Limiting and Ethics

Individual modules return rateLimit status in their result dictionaries. When integrating Holehe at scale, respect these indicators and implement your own throttling—launch_module() catches exceptions but doesn't globally rate-limit across calls.

Result Format

Every service returns a dictionary with these standard keys (see launch_module() implementation in holehe/core.py):

  • name: Service identifier
  • domain: Target website domain
  • exists: Boolean or None for unclear
  • emailrecovery: Partial email if exposed
  • phoneNumber: Partial phone if exposed
  • others: Additional metadata
  • rateLimit: True if hit by rate limiting

Summary

  • Holehe is fully importable — import individual services from holehe.modules.* or use core orchestration functions from holehe/core.py
  • Three integration patterns: single-service calls (finest control), maincore() (CLI equivalent), or custom nursery orchestration (full flexibility)
  • Async-native architecture — requires trio or compatible runtime; services run concurrently for performance
  • Dynamic module discovery — automatically includes new services without code changes
  • Standardized result format — consistent dictionaries across all 100+ supported sites

Frequently Asked Questions

Does Holehe work with asyncio instead of trio?

Holehe uses trio for structured concurrency. You can bridge to asyncio using trio-asyncio or run Holehe in a separate thread with trio.run(), then bridge results back to your asyncio code. The source in holehe/core.py has no direct asyncio support.

How do I check only specific services, not all 100+?

Pass a module subset to get_functions(). Filter the modules dict returned by import_submodules("holehe.modules") before calling get_functions(), or import service functions directly as shown in Method 1.

Can I use Holehe in a synchronous Python script?

Not directly—all service functions are async. Wrap calls in trio.run() or use asyncio.run() with trio-asyncio. For one-off CLI-style usage from sync code, consider subprocess calls to the holehe command instead.

Where does Holehe store its User-Agent strings?

The ua variable in holehe/localuseragent.py provides randomized User-Agent strings. When calling services directly, the AsyncClient inherits default headers; override by passing custom headers to httpx.AsyncClient() if needed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →