How Sherlock's Site Data JSON Structure Defines Detection Rules for Each Social Network
Sherlock stores all social network detection rules in sherlock_project/resources/data.json, where each entry defines URL patterns, HTTP methods, and validation logic that the Python engine interprets to determine if a username exists.
The sherlock-project/sherlock tool uses a declarative approach to profile enumeration, externalizing platform-specific knowledge into a site data JSON structure that separates detection rules from execution logic. This JSON file contains over 600 individual site definitions, each acting as a standalone configuration that tells Sherlock how to construct requests and interpret responses without requiring changes to the core codebase.
Core Detection Rule Fields
Each top-level key in data.json represents a canonical site name (e.g., GitHub, Discord, Codeforces), with its value containing the complete detection rule definition. The SiteInformation class in sherlock_project/sites.py parses these dictionaries during initialization, mapping JSON fields to attributes that drive the enumeration process.
URL and Metadata Configuration
The most fundamental fields define how Sherlock builds and presents profile links:
url– A template string containing{}as a username placeholder that becomesurl_username_formatin theSiteInformationobject. Sherlock substitutes the target username before making HTTP requests.urlMain– The base homepage URL used for display purposes in result listings and theSiteInformation.__str__method.urlProbe– An optional lightweight endpoint for existence checks that avoids loading full profile pages, used by platforms like GitHub and Codeforces that expose dedicated "user exists" APIs.
Detection Strategy Fields
The errorType field determines which validation strategy Sherlock applies, with support for three distinct detection modes:
status_code– Treats HTTP 200 responses as "claimed" profiles while non-200 status codes indicate unclaimed usernames. This is the fastest method requiring no response body parsing.message– Searches the response body for strings defined in theerrorMsgarray; presence of any error message indicates the username is unclaimed.response_url– Compares the final redirected URL againsterrorUrlto detect "not found" pages that redirect to generic error endpoints.
These strategies are implemented in sherlock_project/sherlock.py (lines 291-302, 351), where the engine branches based on the errorType value to interpret results correctly.
Request Configuration Options
Modern platforms often require specific HTTP methods or authentication headers, which the JSON structure accommodates through optional fields:
request_method– Specifies the HTTP verb (defaults toGET), with some platforms like Discord and Holopin requiringPOSTrequests.request_payload– JSON object or form data sent with POST requests, supporting username placeholder substitution within the payload structure.headers– Dictionary of custom HTTP headers (e.g.,Content-Type: application/json) inserted into the request before transmission.regexCheck– A regular expression pattern that preemptively validates usernames before network requests, preventing unnecessary HTTP calls for invalid formats based on platform-specific constraints like character restrictions or length limits.
Safety and Validation Metadata
Additional fields ensure safe and accurate scanning:
isNSFW– Boolean flag indicating adult content that activates theSitesInformation.remove_nsfw_sitesmethod when users opt out of NSFW results.username_claimed– A known-valid username used bytests/test_manifest.pyto validate that detection rules work correctly against live platforms.username_unclaimed– A generated internal placeholder created bySiteInformation.__init__for negative testing, though this is not stored in the JSON file itself.
How Detection Rules Are Executed
The separation between rule declaration and execution allows Sherlock to apply generic logic to platform-specific requirements. When you initiate a scan, Sherlock follows this workflow:
-
Rule Loading – The
SitesInformationclass instantiates by fetchingdata.jsonfrom the repository or a user-provided path, creating aSiteInformationobject for each entry that encapsulates the detection logic. -
Request Construction – For each target site, Sherlock checks
regexCheck(if present), formats theurltemplate with the target username, and constructs an HTTP request using the specifiedrequest_method,headers, andrequest_payload. -
Response Interpretation – Based on the
errorTypevalue, the engine examines status codes, searches response text forerrorMsgstrings, or compares redirect URLs againsterrorUrlto determine profile existence. -
Result Compilation – Sites returning "claimed" status populate the final report with formatted URLs constructed from the original templates.
Practical Implementation Examples
Loading and Inspecting Detection Rules
You can programmatically inspect how specific platforms define their detection criteria:
from sherlock_project.sites import SitesInformation
# Load live data.json from the official repo
sites_info = SitesInformation()
# Access the rule for GitHub
github_rule = sites_info.sites["GitHub"]
print(github_rule) # → GitHub (https://www.github.com/)
print(github_rule.url_username_format) # → https://www.github.com/{}
print(github_rule.information["errorType"]) # → status_code
Manual Execution of POST-Based Detection
Some platforms require complex request configurations that the JSON structure handles declaratively:
import requests
from sherlock_project.sites import SitesInformation
sites = SitesInformation()
site = sites.sites["Discord"] # uses POST + JSON payload
username = "someuser123"
# Build request according to the rule
payload = site.information["request_payload"]
payload = {k: v.replace("{}", username) for k, v in payload.items()}
resp = requests.post(site.url_username_format,
json=payload,
headers=site.information.get("headers", {}))
# Interpret according to errorType == "message"
claimed = not any(err in resp.text for err in site.information["errorMsg"])
print(f"{username} on Discord {'exists' if claimed else 'does not exist'}")
Filtering NSFW Sites Before Scanning
The JSON metadata enables content filtering without modifying the rule set:
from sherlock_project.sites import SitesInformation
sites = SitesInformation()
sites.remove_nsfw_sites(do_not_remove=["BongaCams"]) # keep a specific NSFW site
print(f"Remaining sites: {len(sites)}")
Summary
- Centralized configuration – All detection logic resides in
sherlock_project/resources/data.json, allowing hundreds of social networks to be defined without Python code changes. - Three validation strategies – The
errorTypefield supportsstatus_code,message, andresponse_urldetection methods that accommodate different platform behaviors. - Flexible request building – Optional fields like
request_method,request_payload, andheadershandle APIs requiring POST data or authentication headers. - Pre-validation optimization – The
regexCheckfield prevents unnecessary HTTP requests by filtering invalid usernames before network calls. - Safety controls – Boolean flags like
isNSFWintegrate withSitesInformation.remove_nsfw_sitesto respect content preferences.
Frequently Asked Questions
What is the difference between url and urlProbe in the detection rules?
The url field defines the public profile page template that Sherlock reports to users when a username is found, while urlProbe specifies an internal API endpoint optimized for existence checks. For example, GitHub's urlProbe might query the user API directly, reducing bandwidth and avoiding full HTML page loads when verifying account existence.
How does the regexCheck field improve scanning performance?
regexCheck contains a regular expression that usernames must match before Sherlock sends HTTP requests. Platforms like Twitter or Instagram have specific username requirements (alphanumeric characters, length limits, no special characters). By validating against these patterns first, Sherlock eliminates network overhead for usernames that could never exist on that platform.
Can I create a custom data.json file for internal company platforms?
Yes, the SitesInformation class accepts custom JSON file paths, allowing you to define detection rules for internal corporate networks or private platforms using the same schema. You must define the required fields (url, errorType) and optionally include urlMain, errorMsg, or request_headers depending on your authentication requirements and detection strategy.
How does Sherlock handle platforms that return HTTP 200 for missing profiles?
Platforms using errorType: message handle this scenario by searching the response body for specific text strings defined in errorMsg. Even if the server returns a 200 status code, presence of error text like "user not found" or "profile unavailable" indicates an unclaimed username, while absence of these strings suggests the profile exists.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →