How Sherlock Handles HTTP Errors vs Application-Level Errors: A Deep Dive into the Request Pipeline
Sherlock distinguishes transport-level failures (network timeouts, DNS errors, HTTP 5xx) from application-level "username not found" signals by first catching requests exceptions in get_response(), then applying site-specific heuristics (errorType, errorMsg, errorCode) to successful HTTP responses.
Sherlock is an open-source username enumeration tool that queries hundreds of social media sites to check username availability. Understanding how Sherlock handles HTTP errors vs application-level errors is crucial for interpreting its results accurately, as the tool must differentiate between a broken network connection and a valid "user not found" page.
Transport-Level Error Handling in get_response()
All network operations in Sherlock flow through the get_response() function in sherlock_project/sherlock.py (lines 13-40). This function acts as a protective barrier that catches transport-level failures before any application logic runs.
The requests Exception Hierarchy
The function wraps the requests-futures call in a comprehensive exception handler that maps the Python requests exception hierarchy to human-readable error contexts:
requests.exceptions.HTTPError– Triggered whenresponse.raise_for_status()encounters 4xx or 5xx status codesrequests.exceptions.ProxyError– Proxy misconfiguration or unreachable proxy serverrequests.exceptions.ConnectionError– DNS failures, refused connections, or network unreachablerequests.exceptions.Timeout– Request exceeded the user-specified timeout thresholdrequests.exceptions.RequestException– Catch-all for any other requests-related errors
Mapping Exceptions to Error Contexts
When get_response() catches an exception, it returns a tuple of (None, error_context, exception_text) where error_context is set to specific strings:
| Exception | Error Context String |
|---|---|
| HTTPError | "HTTP Error" |
| ProxyError | "Proxy Error" |
| ConnectionError | "Error Connecting" |
| Timeout | "Timeout Error" |
| RequestException | "Unknown Error" |
The caller stores this context in the QueryResult.context field defined in sherlock_project/result.py (lines 30-38), ensuring transport failures are clearly distinguished from username availability checks.
Application-Level Error Detection After Successful Responses
Once get_response() returns a non-None response, Sherlock transitions to application-level error detection. This phase determines whether the HTTP response indicates a claimed username or an available one.
Site-Specific Error Type Definitions
Each site in Sherlock's manifest defines an errorType field loaded by sherlock_project/sites.py (lines 165-190). The main loop in sherlock() (lines 390-440 in sherlock_project/sherlock.py) uses this definition to interpret responses.
The Three Detection Strategies
Sherlock supports three application-level detection methods specified in errorType:
1. Message-Based Detection ("message")
When errorType contains "message", Sherlock searches the response body for strings defined in errorMsg. If any error message is found, the username is marked AVAILABLE; otherwise CLAIMED.
2. Status Code Detection ("status_code")
When errorType contains "status_code", Sherlock compares response.status_code against errorCode values. Matching codes indicate AVAILABLE; other 2xx codes indicate CLAIMED. Non-2xx codes are conservatively treated as AVAILABLE to avoid false positives from generic error pages.
3. Response URL Detection ("response_url")
When errorType contains "response_url", Sherlock treats any non-2xx status code as AVAILABLE. This strategy is used when sites disable redirects and return specific status codes for unclaimed usernames.
Code Examples: Inspecting Both Error Types
Example 1: Detecting Transport-Level Errors
This example demonstrates how to programmatically invoke Sherlock and inspect transport-level failures:
from sherlock_project.sherlock import sherlock
from sherlock_project.sites import SitesInformation
from sherlock_project.notify import QueryNotify
# Load the live site manifest
sites = SitesInformation()
# Simple notifier that captures results
class PrintNotify(QueryNotify):
def start(self, username):
print(f"Checking {username}…")
def update(self, result):
print(f"{result.site_name}: {result}")
notify = PrintNotify()
# Run lookup with a deliberately short timeout to trigger transport errors
results = sherlock(
username="testuser123",
site_data=sites.sites,
query_notify=notify,
timeout=0.001, # Aggressive timeout to force Timeout Error
)
# Examine transport errors
for site, data in results.items():
if data["status"].context: # Non-None indicates transport failure
print(f"Transport error on {site}: {data['status'].context}")
The context attribute contains strings like "Timeout Error" or "Error Connecting" that originate from get_response() exception handling.
Example 2: Application-Level Username Detection
This example shows how Sherlock applies site-specific rules to determine username availability:
# Assuming a site manifest entry for "example-forum.com":
# {
# "errorType": ["message"],
# "errorMsg": ["User not found", "This user doesn't exist"],
# ...
# }
# After a successful HTTP response (status 200), Sherlock executes:
def check_username_availability(response, site_data):
error_type = site_data.get("errorType", [])
if "message" in error_type:
error_messages = site_data.get("errorMsg", [])
response_text = response.text
# Check if any error message appears in the response
for msg in error_messages:
if msg in response_text:
return "AVAILABLE" # Username not found
return "CLAIMED" # Username exists (no error message found)
elif "status_code" in error_type:
error_codes = site_data.get("errorCode", [])
if response.status_code in error_codes:
return "AVAILABLE"
elif 200 <= response.status_code < 300:
return "CLAIMED"
else:
return "AVAILABLE" # Conservative fallback
# Usage
# status = check_username_availability(response, site_data)
This logic resides in the main loop of sherlock() (lines 390-440 in sherlock_project/sherlock.py), applying the errorType definitions loaded by sites.py.
Key Files in the Error Handling Architecture
| File | Role in Error Handling |
|---|---|
sherlock_project/sherlock.py |
Contains get_response() for transport-level error catching and the main sherlock() loop (lines 390-440) for application-level detection |
sherlock_project/result.py |
Defines QueryResult class (lines 30-38) that stores both the final status and transport error context |
sherlock_project/sites.py |
Loads site manifests (lines 165-190) containing errorType, errorMsg, and errorCode definitions |
sherlock_project/notify.py |
Interface for propagating results including error contexts to callers |
Summary
-
Transport-level errors (network failures, timeouts, HTTP 5xx) are caught in
get_response()insherlock_project/sherlock.pyand stored in theQueryResult.contextfield as descriptive strings like"Timeout Error"or"Error Connecting". -
Application-level errors are determined only after a successful HTTP response, using site-specific rules defined in the manifest (
errorType,errorMsg,errorCode) loaded bysherlock_project/sites.py. -
The main loop in
sherlock()(lines 390-440) applies three detection strategies: message-based (searching response text), status-code-based (comparing HTTP codes), and response-URL-based (handling redirect behavior). -
This two-phase architecture ensures that network problems do not trigger false positives for "available" usernames, while allowing flexible, site-specific detection of unclaimed accounts.
Frequently Asked Questions
What is the difference between a transport error and an application error in Sherlock?
A transport error occurs when Sherlock cannot successfully complete an HTTP request due to network issues, timeouts, DNS failures, or HTTP protocol errors (5xx status codes). These are caught in get_response() and stored in the context field of the result. An application error occurs when the HTTP request succeeds but the target site returns content indicating the username does not exist (such as "User not found" text or specific 404 status codes). These are detected by the site-specific errorType logic in the main sherlock() loop.
How does Sherlock handle HTTP 404 status codes?
Sherlock's handling of HTTP 404 depends on the site's errorType configuration. If the site uses "status_code" detection and lists 404 in errorCode, Sherlock marks the username as AVAILABLE. If the site uses "message" detection, Sherlock ignores the 404 status and instead searches the response body for errorMsg strings. If the site uses "response_url" detection, any non-2xx status code (including 404) is treated as AVAILABLE. This flexibility allows Sherlock to handle sites that return 404 for missing users versus sites that return 200 with "not found" text.
What happens when Sherlock encounters a timeout?
When a request exceeds the user-specified timeout threshold, the requests library raises requests.exceptions.Timeout. In get_response() (lines 13-40 of sherlock_project/sherlock.py), this exception is caught and mapped to the error context string "Timeout Error". The function returns (None, "Timeout Error", exception_text), and the caller stores this context in the QueryResult object. The username check is marked as UNKNOWN rather than AVAILABLE, ensuring that network timeouts do not generate false positives for available usernames.
Can I customize error detection for a specific site?
Yes, error detection is controlled through the site manifest JSON files loaded by sherlock_project/sites.py (lines 165-190). Each site entry can define errorType as a list containing "message", "status_code", or "response_url". For "message", provide errorMsg (a string or list of strings) that indicate an unavailable username. For "status_code", provide errorCode (integer or list) of HTTP codes indicating unavailability. These definitions are evaluated in the main loop of sherlock() (lines 390-440), allowing per-site customization without modifying the core Python code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →