# How Sherlock Handles HTTP Errors vs Application-Level Errors: A Deep Dive into the Request Pipeline

> Understand how Sherlock handles HTTP errors versus application errors. Learn how it differentiates transport failures from site-specific signals within its request pipeline.

- Repository: [Sherlock/sherlock](https://github.com/sherlock-project/sherlock)
- Tags: deep-dive
- Published: 2026-03-02

---

**Sherlock distinguishes transport-level failures (network timeouts, DNS errors, HTTP 5xx) from application-level "username not found" signals by first catching requests exceptions in `get_response()`, then applying site-specific heuristics (`errorType`, `errorMsg`, `errorCode`) to successful HTTP responses.**

Sherlock is an open-source username enumeration tool that queries hundreds of social media sites to check username availability. Understanding how Sherlock handles HTTP errors vs application-level errors is crucial for interpreting its results accurately, as the tool must differentiate between a broken network connection and a valid "user not found" page.

## Transport-Level Error Handling in `get_response()`

All network operations in Sherlock flow through the `get_response()` function in [`sherlock_project/sherlock.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sherlock.py) (lines 13-40). This function acts as a protective barrier that catches transport-level failures before any application logic runs.

### The requests Exception Hierarchy

The function wraps the `requests-futures` call in a comprehensive exception handler that maps the Python `requests` exception hierarchy to human-readable error contexts:

- **`requests.exceptions.HTTPError`** – Triggered when `response.raise_for_status()` encounters 4xx or 5xx status codes
- **`requests.exceptions.ProxyError`** – Proxy misconfiguration or unreachable proxy server
- **`requests.exceptions.ConnectionError`** – DNS failures, refused connections, or network unreachable
- **`requests.exceptions.Timeout`** – Request exceeded the user-specified timeout threshold
- **`requests.exceptions.RequestException`** – Catch-all for any other requests-related errors

### Mapping Exceptions to Error Contexts

When `get_response()` catches an exception, it returns a tuple of `(None, error_context, exception_text)` where `error_context` is set to specific strings:

| Exception | Error Context String |
|-----------|---------------------|
| HTTPError | `"HTTP Error"` |
| ProxyError | `"Proxy Error"` |
| ConnectionError | `"Error Connecting"` |
| Timeout | `"Timeout Error"` |
| RequestException | `"Unknown Error"` |

The caller stores this context in the `QueryResult.context` field defined in [`sherlock_project/result.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/result.py) (lines 30-38), ensuring transport failures are clearly distinguished from username availability checks.

## Application-Level Error Detection After Successful Responses

Once `get_response()` returns a non-`None` response, Sherlock transitions to application-level error detection. This phase determines whether the HTTP response indicates a claimed username or an available one.

### Site-Specific Error Type Definitions

Each site in Sherlock's manifest defines an `errorType` field loaded by [`sherlock_project/sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sites.py) (lines 165-190). The main loop in `sherlock()` (lines 390-440 in [`sherlock_project/sherlock.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sherlock.py)) uses this definition to interpret responses.

### The Three Detection Strategies

Sherlock supports three application-level detection methods specified in `errorType`:

**1. Message-Based Detection (`"message"`)**

When `errorType` contains `"message"`, Sherlock searches the response body for strings defined in `errorMsg`. If any error message is found, the username is marked **AVAILABLE**; otherwise **CLAIMED**.

**2. Status Code Detection (`"status_code"`)**

When `errorType` contains `"status_code"`, Sherlock compares `response.status_code` against `errorCode` values. Matching codes indicate **AVAILABLE**; other 2xx codes indicate **CLAIMED**. Non-2xx codes are conservatively treated as **AVAILABLE** to avoid false positives from generic error pages.

**3. Response URL Detection (`"response_url"`)**

When `errorType` contains `"response_url"`, Sherlock treats any non-2xx status code as **AVAILABLE**. This strategy is used when sites disable redirects and return specific status codes for unclaimed usernames.

## Code Examples: Inspecting Both Error Types

### Example 1: Detecting Transport-Level Errors

This example demonstrates how to programmatically invoke Sherlock and inspect transport-level failures:

```python
from sherlock_project.sherlock import sherlock
from sherlock_project.sites import SitesInformation
from sherlock_project.notify import QueryNotify

# Load the live site manifest

sites = SitesInformation()

# Simple notifier that captures results

class PrintNotify(QueryNotify):
    def start(self, username): 
        print(f"Checking {username}…")
    
    def update(self, result):
        print(f"{result.site_name}: {result}")

notify = PrintNotify()

# Run lookup with a deliberately short timeout to trigger transport errors

results = sherlock(
    username="testuser123",
    site_data=sites.sites,
    query_notify=notify,
    timeout=0.001,  # Aggressive timeout to force Timeout Error

)

# Examine transport errors

for site, data in results.items():
    if data["status"].context:  # Non-None indicates transport failure

        print(f"Transport error on {site}: {data['status'].context}")

```

*The `context` attribute contains strings like `"Timeout Error"` or `"Error Connecting"` that originate from `get_response()` exception handling.*

### Example 2: Application-Level Username Detection

This example shows how Sherlock applies site-specific rules to determine username availability:

```python

# Assuming a site manifest entry for "example-forum.com":

# {

#   "errorType": ["message"],

#   "errorMsg": ["User not found", "This user doesn't exist"],

#   ...

# }

# After a successful HTTP response (status 200), Sherlock executes:

def check_username_availability(response, site_data):
    error_type = site_data.get("errorType", [])
    
    if "message" in error_type:
        error_messages = site_data.get("errorMsg", [])
        response_text = response.text
        
        # Check if any error message appears in the response

        for msg in error_messages:
            if msg in response_text:
                return "AVAILABLE"  # Username not found

        
        return "CLAIMED"  # Username exists (no error message found)

    
    elif "status_code" in error_type:
        error_codes = site_data.get("errorCode", [])
        if response.status_code in error_codes:
            return "AVAILABLE"
        elif 200 <= response.status_code < 300:
            return "CLAIMED"
        else:
            return "AVAILABLE"  # Conservative fallback

# Usage

# status = check_username_availability(response, site_data)

```

*This logic resides in the main loop of `sherlock()` (lines 390-440 in [`sherlock_project/sherlock.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sherlock.py)), applying the `errorType` definitions loaded by [`sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sites.py).*

## Key Files in the Error Handling Architecture

| File | Role in Error Handling |
|------|------------------------|
| [`sherlock_project/sherlock.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sherlock.py) | Contains `get_response()` for transport-level error catching and the main `sherlock()` loop (lines 390-440) for application-level detection |
| [`sherlock_project/result.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/result.py) | Defines `QueryResult` class (lines 30-38) that stores both the final status and transport error context |
| [`sherlock_project/sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sites.py) | Loads site manifests (lines 165-190) containing `errorType`, `errorMsg`, and `errorCode` definitions |
| [`sherlock_project/notify.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/notify.py) | Interface for propagating results including error contexts to callers |

## Summary

- **Transport-level errors** (network failures, timeouts, HTTP 5xx) are caught in `get_response()` in [`sherlock_project/sherlock.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sherlock.py) and stored in the `QueryResult.context` field as descriptive strings like `"Timeout Error"` or `"Error Connecting"`.

- **Application-level errors** are determined only after a successful HTTP response, using site-specific rules defined in the manifest (`errorType`, `errorMsg`, `errorCode`) loaded by [`sherlock_project/sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sites.py).

- The main loop in `sherlock()` (lines 390-440) applies three detection strategies: message-based (searching response text), status-code-based (comparing HTTP codes), and response-URL-based (handling redirect behavior).

- This two-phase architecture ensures that network problems do not trigger false positives for "available" usernames, while allowing flexible, site-specific detection of unclaimed accounts.

## Frequently Asked Questions

### What is the difference between a transport error and an application error in Sherlock?

A **transport error** occurs when Sherlock cannot successfully complete an HTTP request due to network issues, timeouts, DNS failures, or HTTP protocol errors (5xx status codes). These are caught in `get_response()` and stored in the `context` field of the result. An **application error** occurs when the HTTP request succeeds but the target site returns content indicating the username does not exist (such as "User not found" text or specific 404 status codes). These are detected by the site-specific `errorType` logic in the main `sherlock()` loop.

### How does Sherlock handle HTTP 404 status codes?

Sherlock's handling of HTTP 404 depends on the site's `errorType` configuration. If the site uses `"status_code"` detection and lists 404 in `errorCode`, Sherlock marks the username as **AVAILABLE**. If the site uses `"message"` detection, Sherlock ignores the 404 status and instead searches the response body for `errorMsg` strings. If the site uses `"response_url"` detection, any non-2xx status code (including 404) is treated as **AVAILABLE**. This flexibility allows Sherlock to handle sites that return 404 for missing users versus sites that return 200 with "not found" text.

### What happens when Sherlock encounters a timeout?

When a request exceeds the user-specified timeout threshold, the `requests` library raises `requests.exceptions.Timeout`. In `get_response()` (lines 13-40 of [`sherlock_project/sherlock.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sherlock.py)), this exception is caught and mapped to the error context string `"Timeout Error"`. The function returns `(None, "Timeout Error", exception_text)`, and the caller stores this context in the `QueryResult` object. The username check is marked as **UNKNOWN** rather than **AVAILABLE**, ensuring that network timeouts do not generate false positives for available usernames.

### Can I customize error detection for a specific site?

Yes, error detection is controlled through the site manifest JSON files loaded by [`sherlock_project/sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sites.py) (lines 165-190). Each site entry can define `errorType` as a list containing `"message"`, `"status_code"`, or `"response_url"`. For `"message"`, provide `errorMsg` (a string or list of strings) that indicate an unavailable username. For `"status_code"`, provide `errorCode` (integer or list) of HTTP codes indicating unavailability. These definitions are evaluated in the main loop of `sherlock()` (lines 390-440), allowing per-site customization without modifying the core Python code.