How Sherlock Handles HTTP Errors vs Application-Level Errors: A Deep Dive into the Request Pipeline

Sherlock distinguishes transport-level failures (network timeouts, DNS errors, HTTP 5xx) from application-level "username not found" signals by first catching requests exceptions in get_response(), then applying site-specific heuristics (errorType, errorMsg, errorCode) to successful HTTP responses.

Sherlock is an open-source username enumeration tool that queries hundreds of social media sites to check username availability. Understanding how Sherlock handles HTTP errors vs application-level errors is crucial for interpreting its results accurately, as the tool must differentiate between a broken network connection and a valid "user not found" page.

Transport-Level Error Handling in get_response()

All network operations in Sherlock flow through the get_response() function in sherlock_project/sherlock.py (lines 13-40). This function acts as a protective barrier that catches transport-level failures before any application logic runs.

The requests Exception Hierarchy

The function wraps the requests-futures call in a comprehensive exception handler that maps the Python requests exception hierarchy to human-readable error contexts:

  • requests.exceptions.HTTPError – Triggered when response.raise_for_status() encounters 4xx or 5xx status codes
  • requests.exceptions.ProxyError – Proxy misconfiguration or unreachable proxy server
  • requests.exceptions.ConnectionError – DNS failures, refused connections, or network unreachable
  • requests.exceptions.Timeout – Request exceeded the user-specified timeout threshold
  • requests.exceptions.RequestException – Catch-all for any other requests-related errors

Mapping Exceptions to Error Contexts

When get_response() catches an exception, it returns a tuple of (None, error_context, exception_text) where error_context is set to specific strings:

Exception Error Context String
HTTPError "HTTP Error"
ProxyError "Proxy Error"
ConnectionError "Error Connecting"
Timeout "Timeout Error"
RequestException "Unknown Error"

The caller stores this context in the QueryResult.context field defined in sherlock_project/result.py (lines 30-38), ensuring transport failures are clearly distinguished from username availability checks.

Application-Level Error Detection After Successful Responses

Once get_response() returns a non-None response, Sherlock transitions to application-level error detection. This phase determines whether the HTTP response indicates a claimed username or an available one.

Site-Specific Error Type Definitions

Each site in Sherlock's manifest defines an errorType field loaded by sherlock_project/sites.py (lines 165-190). The main loop in sherlock() (lines 390-440 in sherlock_project/sherlock.py) uses this definition to interpret responses.

The Three Detection Strategies

Sherlock supports three application-level detection methods specified in errorType:

1. Message-Based Detection ("message")

When errorType contains "message", Sherlock searches the response body for strings defined in errorMsg. If any error message is found, the username is marked AVAILABLE; otherwise CLAIMED.

2. Status Code Detection ("status_code")

When errorType contains "status_code", Sherlock compares response.status_code against errorCode values. Matching codes indicate AVAILABLE; other 2xx codes indicate CLAIMED. Non-2xx codes are conservatively treated as AVAILABLE to avoid false positives from generic error pages.

3. Response URL Detection ("response_url")

When errorType contains "response_url", Sherlock treats any non-2xx status code as AVAILABLE. This strategy is used when sites disable redirects and return specific status codes for unclaimed usernames.

Code Examples: Inspecting Both Error Types

Example 1: Detecting Transport-Level Errors

This example demonstrates how to programmatically invoke Sherlock and inspect transport-level failures:

from sherlock_project.sherlock import sherlock
from sherlock_project.sites import SitesInformation
from sherlock_project.notify import QueryNotify

# Load the live site manifest

sites = SitesInformation()

# Simple notifier that captures results

class PrintNotify(QueryNotify):
    def start(self, username): 
        print(f"Checking {username}…")
    
    def update(self, result):
        print(f"{result.site_name}: {result}")

notify = PrintNotify()

# Run lookup with a deliberately short timeout to trigger transport errors

results = sherlock(
    username="testuser123",
    site_data=sites.sites,
    query_notify=notify,
    timeout=0.001,  # Aggressive timeout to force Timeout Error

)

# Examine transport errors

for site, data in results.items():
    if data["status"].context:  # Non-None indicates transport failure

        print(f"Transport error on {site}: {data['status'].context}")

The context attribute contains strings like "Timeout Error" or "Error Connecting" that originate from get_response() exception handling.

Example 2: Application-Level Username Detection

This example shows how Sherlock applies site-specific rules to determine username availability:


# Assuming a site manifest entry for "example-forum.com":

# {

#   "errorType": ["message"],

#   "errorMsg": ["User not found", "This user doesn't exist"],

#   ...

# }

# After a successful HTTP response (status 200), Sherlock executes:

def check_username_availability(response, site_data):
    error_type = site_data.get("errorType", [])
    
    if "message" in error_type:
        error_messages = site_data.get("errorMsg", [])
        response_text = response.text
        
        # Check if any error message appears in the response

        for msg in error_messages:
            if msg in response_text:
                return "AVAILABLE"  # Username not found

        
        return "CLAIMED"  # Username exists (no error message found)

    
    elif "status_code" in error_type:
        error_codes = site_data.get("errorCode", [])
        if response.status_code in error_codes:
            return "AVAILABLE"
        elif 200 <= response.status_code < 300:
            return "CLAIMED"
        else:
            return "AVAILABLE"  # Conservative fallback

# Usage

# status = check_username_availability(response, site_data)

This logic resides in the main loop of sherlock() (lines 390-440 in sherlock_project/sherlock.py), applying the errorType definitions loaded by sites.py.

Key Files in the Error Handling Architecture

File Role in Error Handling
sherlock_project/sherlock.py Contains get_response() for transport-level error catching and the main sherlock() loop (lines 390-440) for application-level detection
sherlock_project/result.py Defines QueryResult class (lines 30-38) that stores both the final status and transport error context
sherlock_project/sites.py Loads site manifests (lines 165-190) containing errorType, errorMsg, and errorCode definitions
sherlock_project/notify.py Interface for propagating results including error contexts to callers

Summary

  • Transport-level errors (network failures, timeouts, HTTP 5xx) are caught in get_response() in sherlock_project/sherlock.py and stored in the QueryResult.context field as descriptive strings like "Timeout Error" or "Error Connecting".

  • Application-level errors are determined only after a successful HTTP response, using site-specific rules defined in the manifest (errorType, errorMsg, errorCode) loaded by sherlock_project/sites.py.

  • The main loop in sherlock() (lines 390-440) applies three detection strategies: message-based (searching response text), status-code-based (comparing HTTP codes), and response-URL-based (handling redirect behavior).

  • This two-phase architecture ensures that network problems do not trigger false positives for "available" usernames, while allowing flexible, site-specific detection of unclaimed accounts.

Frequently Asked Questions

What is the difference between a transport error and an application error in Sherlock?

A transport error occurs when Sherlock cannot successfully complete an HTTP request due to network issues, timeouts, DNS failures, or HTTP protocol errors (5xx status codes). These are caught in get_response() and stored in the context field of the result. An application error occurs when the HTTP request succeeds but the target site returns content indicating the username does not exist (such as "User not found" text or specific 404 status codes). These are detected by the site-specific errorType logic in the main sherlock() loop.

How does Sherlock handle HTTP 404 status codes?

Sherlock's handling of HTTP 404 depends on the site's errorType configuration. If the site uses "status_code" detection and lists 404 in errorCode, Sherlock marks the username as AVAILABLE. If the site uses "message" detection, Sherlock ignores the 404 status and instead searches the response body for errorMsg strings. If the site uses "response_url" detection, any non-2xx status code (including 404) is treated as AVAILABLE. This flexibility allows Sherlock to handle sites that return 404 for missing users versus sites that return 200 with "not found" text.

What happens when Sherlock encounters a timeout?

When a request exceeds the user-specified timeout threshold, the requests library raises requests.exceptions.Timeout. In get_response() (lines 13-40 of sherlock_project/sherlock.py), this exception is caught and mapped to the error context string "Timeout Error". The function returns (None, "Timeout Error", exception_text), and the caller stores this context in the QueryResult object. The username check is marked as UNKNOWN rather than AVAILABLE, ensuring that network timeouts do not generate false positives for available usernames.

Can I customize error detection for a specific site?

Yes, error detection is controlled through the site manifest JSON files loaded by sherlock_project/sites.py (lines 165-190). Each site entry can define errorType as a list containing "message", "status_code", or "response_url". For "message", provide errorMsg (a string or list of strings) that indicate an unavailable username. For "status_code", provide errorCode (integer or list) of HTTP codes indicating unavailability. These definitions are evaluated in the main loop of sherlock() (lines 390-440), allowing per-site customization without modifying the core Python code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →