# How Holehe Uses BeautifulSoup (bs4) for HTML Parsing in OSINT Investigations

> Discover how Holehe utilizes BeautifulSoup (bs4) to parse HTML, extracting key data from web pages for OSINT investigations and determining email registration status on services.

- Repository: [Palenath/holehe](https://github.com/megadose/holehe)
- Tags: how-to-guide
- Published: 2026-09-09

---

**Holehe leverages BeautifulSoup to transform raw HTTP responses into searchable DOM trees, extracting specific HTML elements like form inputs, script tags, and authentication alerts to determine if an email address is registered on target services.**

Holehe is an asynchronous OSINT tool maintained by megadose that checks whether an email address exists across hundreds of online platforms. The core parsing logic relies on BeautifulSoup (bs4) to process HTML responses, enabling the tool to locate hidden form fields, detect error messages, and extract CSRF tokens from unstructured web pages.

## The BeautifulSoup Parsing Workflow

Every service module in the Holehe repository follows a consistent five-step parsing pattern. After executing an async HTTP request with `httpx`, the response is immediately passed to `BeautifulSoup` with the `html.parser` backend.

The standard workflow includes:

1. **Fetch the target page** using `await client.get(url)` or `await client.post(url)`.
2. **Instantiate the parser** with `BeautifulSoup(response.text, "html.parser")` or `BeautifulSoup(response.content, "html.parser")` for bytes.
3. **Locate DOM markers** via `soup.find()` or `soup.select()` to identify account existence indicators.
4. **Extract relevant data** such as hidden input values, CSRF tokens, or specific error message containers.
5. **Determine the result** based on element presence or absence, appending a dictionary with `exists`, `rateLimit`, and other metadata to the output list.

This approach minimizes memory overhead while providing the flexibility to handle diverse website structures.

## Implementation Examples from the Source Code

### Amazon Module (holehe/modules/shopping/amazon.py)

The Amazon implementation demonstrates two-phase parsing. First, it extracts hidden form fields required for the authentication flow:

```python
body = BeautifulSoup(req.text, 'html.parser')
data = dict([(x["name"], x["value"]) for x in body.select(
    'form input') if ('name' in x.attrs and 'value' in x.attrs)])
data["email"] = email

```

After submitting the POST request, it parses the response again to detect the password-missing alert:

```python
body = BeautifulSoup(req.text, 'html.parser')
if body.find("div", {"id": "auth-password-missing-alert"}):
    out.append({"exists": True, ...})

```

The presence of the `auth-password-missing-alert` div confirms the email is registered in Amazon's system.

### Tumblr Module (holehe/modules/social_media/tumblr.py)

For Tumblr, the parser targets script tags to extract bearer tokens or configuration objects embedded in the page:

```python
p = BeautifulSoup(getBearer.text, "html.parser").find_all("script")

```

This collects all JavaScript blocks, allowing the module to search for authentication markers within inline scripts.

### Odnoklassniki Module (holehe/modules/social_media/odnoklassniki.py)

This module searches for profile-specific containers that appear only on registered user pages:

```python
root_soup = BeautifulSoup(request.content, 'html.parser')
if root_soup.find("div", {"class": "profile-info"}):
    out.append({"exists": True, ...})

```

Note the use of `request.content` (bytes) rather than `request.text`, demonstrating BeautifulSoup's ability to handle both input types.

### SoundCloud Module (holehe/modules/music/soundcloud.py)

SoundCloud requires extracting authentication tokens from specific script elements. The module uses indexed access after parsing:

```python
script = BeautifulSoup(getAuth.text, 'html.parser').find_all('script')[4]

```

This targets the fifth script tag on the page, which typically contains the client ID required for API validation.

## Core Import Structure

All modules inherit the BeautifulSoup dependency from [`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py), which handles the centralized import:

```python
from bs4 import BeautifulSoup

```

This ensures consistency across the codebase, allowing every service module to instantiate parsers without managing individual dependencies or parser configurations.

## Building Custom Modules with BeautifulSoup

When extending Holehe to support new services, replicate the established parsing pattern. The following template matches the repository's architectural conventions:

```python
import httpx
from bs4 import BeautifulSoup

async def custom_service_check(email: str, client: httpx.AsyncClient, out: list):
    """Check email existence on a target service."""
    # Initial request

    resp = await client.get("https://example.com/login")
    
    # Parse and extract form data

    soup = BeautifulSoup(resp.text, "html.parser")
    payload = {
        i["name"]: i["value"] 
        for i in soup.select("form input")
        if "name" in i.attrs and "value" in i.attrs
    }
    payload["email"] = email
    
    # Submit check

    post_resp = await client.post("https://example.com/auth", data=payload)
    post_soup = BeautifulSoup(post_resp.text, "html.parser")
    
    # Detect existence marker

    if post_soup.find("div", {"class": "account-found"}):
        out.append({
            "name": "example",
            "domain": "example.com",
            "exists": True,
            "rateLimit": False
        })

```

Key implementation details include using `"html.parser"` for portability, leveraging dictionary comprehensions with `soup.select()` for form extraction, and targeting specific IDs or classes with `soup.find()`.

## Summary

- Holehe uses **BeautifulSoup** with the `html.parser` backend to process HTTP responses across all service modules.
- The tool extracts **hidden form fields** using `soup.select('form input')` to construct POST payloads for authentication attempts.
- Detection logic relies on `soup.find()` to locate specific DOM elements like `auth-password-missing-alert` or `profile-info` classes.
- Both **text** (`response.text`) and **bytes** (`response.content`) parsing are supported for handling various response types.
- All modules import from **[`holehe/core.py`](https://github.com/megadose/holehe/blob/main/holehe/core.py)**, maintaining a standardized parsing interface throughout the codebase.

## Frequently Asked Questions

### What HTML parser does Holehe use with BeautifulSoup?

Holehe consistently uses the built-in `html.parser` rather than external dependencies like lxml or html5lib. This choice appears in every module as `BeautifulSoup(response.text, "html.parser")`, ensuring portability across environments without requiring additional system libraries or compiled extensions.

### How does Holehe extract CSRF tokens and hidden form fields?

The tool uses CSS selection via `soup.select("form input")` combined with dictionary comprehensions to map name-value pairs. As seen in [`holehe/modules/shopping/amazon.py`](https://github.com/megadose/holehe/blob/main/holehe/modules/shopping/amazon.py), the code iterates over input elements and filters for those containing both `name` and `value` attributes, creating a data dictionary used in subsequent POST requests.

### Can I modify Holehe to use lxml instead of html.parser?

While the codebase currently hardcodes `html.parser`, you could modify individual modules to use `BeautifulSoup(response.text, "lxml")`. However, this would introduce an external dependency on the `lxml` package that is not required by the original project, potentially complicating deployment in minimal environments or containers where compilation tools are unavailable.

### Why does Holehe use soup.find() for some checks and soup.select() for others?

The choice depends on the target website's DOM structure. `soup.find()` is used when searching for unique identifiers like specific `id` attributes (e.g., `auth-password-missing-alert`), while `soup.select()` is preferred for extracting collections of elements such as all inputs within a form. Both are standard BeautifulSoup APIs selected based on whether the target is a single unique element or a group of similar elements.