How Sherlock Detects Username Existence Using Error Types: status_code, response_url, and message
Sherlock determines whether a username exists on a platform by evaluating HTTP responses against three configurable error strategies—status_code, response_url, and message—defined declaratively in sherlock_project/sites.py.
Sherlock, the open-source reconnaissance tool from the sherlock-project/sherlock repository, identifies account existence across hundreds of services without hardcoding site-specific logic. Instead, it relies on a data-driven engine that interprets per-site error type configurations to classify usernames as Claimed (exists) or Available (does not exist). This modular architecture allows the tool to adapt to diverse platform behaviors—from returning 404s to redirecting or displaying custom error messages—using a single, uniform detection loop.
Status Code Detection (status_code)
When a site configuration specifies errorType: "status_code", Sherlock evaluates the numeric HTTP status code returned by the server. In sherlock_project/sherlock.py, the logic first assumes the username is Claimed (exists) and only reverses this verdict if the response matches failure criteria.
The detection works as follows:
- Explicit error codes: If the site defines an
errorCodefield (e.g.,404), and the response status matches any code in that list, the username is marked Available. - Implicit non-2xx handling: If no
errorCodeis explicitly defined, any status code outside the 200–299 range (i.e.,< 200or>= 300) is also interpreted as Available. - Optimization: For this error type, Sherlock prefers lightweight
HEADrequests when no request body is required, reducing bandwidth and latency.
Response URL Detection (response_url)
The response_url strategy handles platforms that indicate missing accounts through URL redirection rather than status codes. Sherlock disables automatic redirect following (allow_redirects = False) to capture the immediate response before any location changes occur.
The verdict logic in sherlock_project/sherlock.py evaluates the final URL behavior:
- If the server returns a 2xx status (200–299), the username is considered Claimed.
- Any other status code—including 3xx redirects or 4xx client errors—results in an Available classification, as it indicates the profile does not resolve to a valid endpoint.
Message Text Detection (message)
For sites that return HTTP 200 OK pages containing "not found" or "user does not exist" text, Sherlock uses the message (or errorMsg) error type. This requires retrieving the full response body, so Sherlock issues GET requests for these sites.
The detection logic inspects r.text from the response:
- The site configuration provides one or more strings in the
errorMsgfield. - Sherlock checks if any of these strings appear in the HTML body.
- If a match is found, the username is Available; otherwise, it is Claimed.
This method is essential for platforms that serve soft 404s—returning successful status codes while displaying error messages in the page content.
Request Execution and Workflow
Sherlock’s detection engine follows a consistent asynchronous workflow regardless of the error type configured. In sherlock_project/sherlock.py, the process unfolds as:
- Request Construction: The engine builds headers, applies optional proxies, and selects the HTTP method (
HEADforstatus_codewhen possible, otherwiseGET). - Async Dispatch: A
Futureobject representing the network call is stored innet_info["request_future"]. - Response Retrieval: The
get_responsefunction awaits the future, capturing theResponseobject, any exception text, and error context. - Verdict Interpretation: Based on the site’s
errorTypelist, Sherlock applies the appropriate logic (status code, URL, or message checks) to setquery_statustoClaimed,Available,WAF, orUnknown. - Result Packaging: A
QueryResultobject is instantiated and passed to theQueryNotifyinterface for reporting.
This design separates network I/O from detection logic, enabling efficient concurrent scanning across hundreds of sites.
Configuring Detection in sites.py
The data-driven configuration resides in sherlock_project/sites.py, where each platform entry declares which detection strategy to use. Here are representative configurations for each error type:
# In sherlock_project/sites.py
sites = {
"GitHub": {
"urlMain": "https://github.com",
"url": "https://github.com/{}",
"errorType": "status_code",
"errorCode": 404, # 404 indicates username is available
},
"Twitter": {
"urlMain": "https://twitter.com",
"url": "https://twitter.com/{}",
"errorType": "response_url", # Absence of redirect indicates existence
},
"Reddit": {
"urlMain": "https://reddit.com",
"url": "https://www.reddit.com/user/{}",
"errorType": "message",
"errorMsg": "Sorry, there isn’t a Reddit account with that name",
},
}
Each entry tells Sherlock exactly which response attribute to evaluate, allowing the core engine to remain generic while supporting platform-specific detection nuances.
Programmatic Usage Examples
You can invoke Sherlock’s detection engine programmatically to scan specific targets or test new site configurations:
from sherlock_project.sherlock import sherlock
from sherlock_project.result import QueryNotify
class SimpleNotify(QueryNotify):
def start(self, username):
print(f"Scanning {username} …")
def update(self, result):
print(f"{result.site_name}: {result.status.name}")
# Execute scan using the sites configuration
results = sherlock(
username="alice",
site_data=sites, # Imported from sherlock_project.sites
query_notify=SimpleNotify(),
)
For manual testing of detection logic on a specific endpoint:
from sherlock_project.sherlock import get_response
# Initiate a HEAD request for status_code detection
future = session.head(url="https://github.com/nonexistentuser", timeout=5)
# Retrieve response and apply Sherlock's logic
response, err_ctx, exc_text = get_response(
request_future=future,
error_type=["status_code"],
social_network="GitHub"
)
# Evaluate based on status code
if response and response.status_code == 404:
print("Username AVAILABLE")
else:
print("Username CLAIMED")
Summary
- Sherlock uses three error types to detect username existence:
status_code(checks HTTP codes),response_url(checks redirect behavior with redirects disabled), andmessage(checks response body text). - Configuration is data-driven via
sherlock_project/sites.py, allowing new platforms to be added without modifying the core detection engine insherlock_project/sherlock.py. - Request methods are optimized based on error type:
HEADforstatus_code(when no payload needed) andGETformessagedetection. - Results are standardized into
QueryResultobjects with statuses ofClaimed,Available,WAF, orUnknown, enabling consistent reporting across all supported services.
Frequently Asked Questions
What is the difference between status_code and response_url detection in Sherlock?
status_code detection examines the numeric HTTP response code (e.g., treating 404 as Available), while response_url detection inspects URL redirection behavior by disabling redirects (allow_redirects=False) and checking if the final URL returns a 2xx status. The former relies on explicit error codes, whereas the latter infers existence from whether the server attempts to redirect the request.
Why does Sherlock use HEAD requests for some sites but GET for others?
Sherlock selects the HTTP method based on the data requirements of the error type. For status_code detection, a lightweight HEAD request suffices because only headers are needed. For message detection, a GET request is mandatory to retrieve the full HTML body and search for specific error strings. This optimization reduces bandwidth and improves scanning speed.
How do I add a new site to Sherlock that uses message-based detection?
Add a new entry to the dictionary in sherlock_project/sites.py with errorType set to "message" and include an errorMsg field containing the exact string or strings that appear when a username does not exist. For example, {"errorType": "message", "errorMsg": "Profile not found"} will mark the username as Available if that text appears in the response body.
What happens if a site returns a non-2xx status code without an explicit errorCode defined?
When using status_code detection, if no specific errorCode list is provided in the site configuration, Sherlock treats any status code outside the 200–299 range (including 3xx, 4xx, and 5xx responses) as indicating the username is Available. This fallback ensures the tool remains functional even for sites with inconsistent error handling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →