How Holehe Uses BeautifulSoup (bs4) for HTML Parsing in OSINT Investigations
Holehe leverages BeautifulSoup to transform raw HTTP responses into searchable DOM trees, extracting specific HTML elements like form inputs, script tags, and authentication alerts to determine if an email address is registered on target services.
Holehe is an asynchronous OSINT tool maintained by megadose that checks whether an email address exists across hundreds of online platforms. The core parsing logic relies on BeautifulSoup (bs4) to process HTML responses, enabling the tool to locate hidden form fields, detect error messages, and extract CSRF tokens from unstructured web pages.
The BeautifulSoup Parsing Workflow
Every service module in the Holehe repository follows a consistent five-step parsing pattern. After executing an async HTTP request with httpx, the response is immediately passed to BeautifulSoup with the html.parser backend.
The standard workflow includes:
- Fetch the target page using
await client.get(url)orawait client.post(url). - Instantiate the parser with
BeautifulSoup(response.text, "html.parser")orBeautifulSoup(response.content, "html.parser")for bytes. - Locate DOM markers via
soup.find()orsoup.select()to identify account existence indicators. - Extract relevant data such as hidden input values, CSRF tokens, or specific error message containers.
- Determine the result based on element presence or absence, appending a dictionary with
exists,rateLimit, and other metadata to the output list.
This approach minimizes memory overhead while providing the flexibility to handle diverse website structures.
Implementation Examples from the Source Code
Amazon Module (holehe/modules/shopping/amazon.py)
The Amazon implementation demonstrates two-phase parsing. First, it extracts hidden form fields required for the authentication flow:
body = BeautifulSoup(req.text, 'html.parser')
data = dict([(x["name"], x["value"]) for x in body.select(
'form input') if ('name' in x.attrs and 'value' in x.attrs)])
data["email"] = email
After submitting the POST request, it parses the response again to detect the password-missing alert:
body = BeautifulSoup(req.text, 'html.parser')
if body.find("div", {"id": "auth-password-missing-alert"}):
out.append({"exists": True, ...})
The presence of the auth-password-missing-alert div confirms the email is registered in Amazon's system.
Tumblr Module (holehe/modules/social_media/tumblr.py)
For Tumblr, the parser targets script tags to extract bearer tokens or configuration objects embedded in the page:
p = BeautifulSoup(getBearer.text, "html.parser").find_all("script")
This collects all JavaScript blocks, allowing the module to search for authentication markers within inline scripts.
Odnoklassniki Module (holehe/modules/social_media/odnoklassniki.py)
This module searches for profile-specific containers that appear only on registered user pages:
root_soup = BeautifulSoup(request.content, 'html.parser')
if root_soup.find("div", {"class": "profile-info"}):
out.append({"exists": True, ...})
Note the use of request.content (bytes) rather than request.text, demonstrating BeautifulSoup's ability to handle both input types.
SoundCloud Module (holehe/modules/music/soundcloud.py)
SoundCloud requires extracting authentication tokens from specific script elements. The module uses indexed access after parsing:
script = BeautifulSoup(getAuth.text, 'html.parser').find_all('script')[4]
This targets the fifth script tag on the page, which typically contains the client ID required for API validation.
Core Import Structure
All modules inherit the BeautifulSoup dependency from holehe/core.py, which handles the centralized import:
from bs4 import BeautifulSoup
This ensures consistency across the codebase, allowing every service module to instantiate parsers without managing individual dependencies or parser configurations.
Building Custom Modules with BeautifulSoup
When extending Holehe to support new services, replicate the established parsing pattern. The following template matches the repository's architectural conventions:
import httpx
from bs4 import BeautifulSoup
async def custom_service_check(email: str, client: httpx.AsyncClient, out: list):
"""Check email existence on a target service."""
# Initial request
resp = await client.get("https://example.com/login")
# Parse and extract form data
soup = BeautifulSoup(resp.text, "html.parser")
payload = {
i["name"]: i["value"]
for i in soup.select("form input")
if "name" in i.attrs and "value" in i.attrs
}
payload["email"] = email
# Submit check
post_resp = await client.post("https://example.com/auth", data=payload)
post_soup = BeautifulSoup(post_resp.text, "html.parser")
# Detect existence marker
if post_soup.find("div", {"class": "account-found"}):
out.append({
"name": "example",
"domain": "example.com",
"exists": True,
"rateLimit": False
})
Key implementation details include using "html.parser" for portability, leveraging dictionary comprehensions with soup.select() for form extraction, and targeting specific IDs or classes with soup.find().
Summary
- Holehe uses BeautifulSoup with the
html.parserbackend to process HTTP responses across all service modules. - The tool extracts hidden form fields using
soup.select('form input')to construct POST payloads for authentication attempts. - Detection logic relies on
soup.find()to locate specific DOM elements likeauth-password-missing-alertorprofile-infoclasses. - Both text (
response.text) and bytes (response.content) parsing are supported for handling various response types. - All modules import from
holehe/core.py, maintaining a standardized parsing interface throughout the codebase.
Frequently Asked Questions
What HTML parser does Holehe use with BeautifulSoup?
Holehe consistently uses the built-in html.parser rather than external dependencies like lxml or html5lib. This choice appears in every module as BeautifulSoup(response.text, "html.parser"), ensuring portability across environments without requiring additional system libraries or compiled extensions.
How does Holehe extract CSRF tokens and hidden form fields?
The tool uses CSS selection via soup.select("form input") combined with dictionary comprehensions to map name-value pairs. As seen in holehe/modules/shopping/amazon.py, the code iterates over input elements and filters for those containing both name and value attributes, creating a data dictionary used in subsequent POST requests.
Can I modify Holehe to use lxml instead of html.parser?
While the codebase currently hardcodes html.parser, you could modify individual modules to use BeautifulSoup(response.text, "lxml"). However, this would introduce an external dependency on the lxml package that is not required by the original project, potentially complicating deployment in minimal environments or containers where compilation tools are unavailable.
Why does Holehe use soup.find() for some checks and soup.select() for others?
The choice depends on the target website's DOM structure. soup.find() is used when searching for unique identifiers like specific id attributes (e.g., auth-password-missing-alert), while soup.select() is preferred for extracting collections of elements such as all inputs within a form. Both are standard BeautifulSoup APIs selected based on whether the target is a single unique element or a group of similar elements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →