How to Use Holehe's Python API to Scan a Single Website

To scan a single website with Holehe's Python API, import the specific service module directly (e.g., holehe.modules.programing.github) and call its async function with an httpx.AsyncClient and a results list.

Holehe is an open-source Python tool for checking whether an email address is registered across hundreds of online services. While the CLI performs bulk scans, the internal architecture exposes modular async functions that you can invoke individually. This guide shows you how to use the Holehe Python API to target a single website without running the full orchestration logic.

How Holehe's Module System Works

Holehe organizes each service check as a standalone Python module under holehe.modules. Every module implements an async function with a consistent signature:

async def <service_name>(email: str, client: httpx.AsyncClient, out: list) -> None

The function populates the out list with a result dictionary containing these fields:

Field Description
name Internal service identifier (e.g., "github").
domain Human-readable domain (e.g., "github.com").
method Request method used ("register" for most services).
exists True if email is registered, False if not, None on error.
rateLimit True if the endpoint throttled the request.
emailrecovery Recovered email address if available (rare).
phoneNumber Phone number recovered if available.
others Additional payload returned by the service.

The bulk scanning logic lives in holehe/core.py, which builds an httpx.AsyncClient, launches all module functions concurrently with Trio, and aggregates results. Bypassing this lets you target one service directly.

Method 1: Direct Module Import for a Single Service

The simplest approach imports the specific module and calls its function directly.

import asyncio
import httpx
import holehe.modules.programing.github as gh

async def scan_github(email: str):
    # Create an async HTTP client with 10-second timeout (matches core.py defaults)

    async with httpx.AsyncClient(timeout=10) as client:
        results = []                     # Module appends results here

        await gh.github(email, client, results)
        return results[0]                # Single entry for single module

# Execute

email = "test@example.com"
result = asyncio.run(scan_github(email))
print(result)

# Output: {'name': 'github', 'domain': 'github.com', 'method': 'register',

#          'exists': False, 'rateLimit': False, 'emailrecovery': None,

#          'phoneNumber': None, 'others': None}

This mirrors the function call defined in holehe/modules/programing/github.py at lines 5-38.

Method 2: Generic Helper for Any Single Service

Use dynamic imports to build a reusable scanner without hardcoding module paths.

import importlib
import asyncio
import httpx

async def scan_one(service: str, email: str):
    """
    Scan a single service by name.
    service: module name under holehe.modules (e.g., "github", "instagram")
    """
    # Dynamically resolve module path

    mod = importlib.import_module(f"holehe.modules.programing.{service}")
    func = getattr(mod, service)          # Function name matches module name

    async with httpx.AsyncClient(timeout=10) as client:
        out = []
        await func(email, client, out)
        return out[0]

# Example usage

result = asyncio.run(scan_one("github", "alice@example.com"))
print(result)

This replicates the module loading approach from import_submodules in holehe/core.py (lines 37-47) but loads only the target module instead of all available services.

Method 3: Integration with Core Utilities

For advanced use cases, combine single-service scanning with Holehe's built-in instrumentation.

import asyncio
import httpx
from holehe.core import import_submodules
from holehe.instruments import TrioProgress
import trio

async def scan_single_with_progress(service: str, email: str):
    # Load modules using Holehe's loader, then extract target

    modules = import_submodules("holehe.modules")
    target_path = f"holehe.modules.programing.{service}"
    target_func = modules[target_path].__dict__[service]

    async with httpx.AsyncClient(timeout=10) as client:
        out = []
        
        # Add progress instrumentation (1 task total)

        instrument = TrioProgress(1)
        trio.lowlevel.add_instrument(instrument)
        
        await target_func(email, client, out)
        
        trio.lowlevel.remove_instrument(instrument)
        return out[0]

# Run with progress indicator

result = asyncio.run(scan_single_with_progress("github", "bob@example.com"))
print(result)

The TrioProgress class defined in holehe/instruments.py provides visual feedback for async task completion. When scanning a single service, pass 1 as the total task count.

Key Source Files Reference

File Purpose
holehe/core.py Bulk orchestration, client setup, and module discovery.
holehe/modules/programing/github.py Example service implementation showing the standard function pattern.
holehe/localuseragent.py Rotating user-agent strings used by the HTTP client.
holehe/instruments.py TrioProgress class for async progress indication.

Summary

  • Holehe's architecture separates each service check into independent async functions under holehe.modules.
  • Scan a single website by importing the specific module and calling its function with httpx.AsyncClient and a results list.
  • Avoid bulk overhead by bypassing the full orchestration in holehe/core.py when you only need one result.
  • Reuse core utilities like TrioProgress and import_submodules for advanced integrations.

Frequently Asked Questions

How do I find the correct module name for a specific website?

Module names correspond to the Python file name without the .py extension. Check holehe/modules/ and its subdirectories (e.g., programing/, shopping/)—the function name inside each file matches the file name. For Instagram, use instagram; for GitHub, use github.

Can I use a custom HTTP client configuration?

Yes. The httpx.AsyncClient is fully configurable before passing it to the module function. Adjust timeout, headers, proxy, or follow_redirects as needed. Holehe's core.py uses a 10-second timeout by default.

What happens if the service rate-limits my request?

The result dictionary sets rateLimit: True and exists: None. Your calling code should handle this case explicitly—consider adding retry logic with exponential backoff or rotating proxies for repeated scans.

Is Trio required for single-service scans?

No. While Holehe's core uses Trio for concurrency management, individual module functions are standard async def coroutines. You can run them with asyncio as shown in the examples. Trio is only needed if you use TrioProgress or other Trio-specific instrumentation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →