How Sherlock's Site Data JSON Structure Defines Detection Rules for Each Social Network

Sherlock stores all social network detection rules in sherlock_project/resources/data.json, where each entry defines URL patterns, HTTP methods, and validation logic that the Python engine interprets to determine if a username exists.

The sherlock-project/sherlock tool uses a declarative approach to profile enumeration, externalizing platform-specific knowledge into a site data JSON structure that separates detection rules from execution logic. This JSON file contains over 600 individual site definitions, each acting as a standalone configuration that tells Sherlock how to construct requests and interpret responses without requiring changes to the core codebase.

Core Detection Rule Fields

Each top-level key in data.json represents a canonical site name (e.g., GitHub, Discord, Codeforces), with its value containing the complete detection rule definition. The SiteInformation class in sherlock_project/sites.py parses these dictionaries during initialization, mapping JSON fields to attributes that drive the enumeration process.

URL and Metadata Configuration

The most fundamental fields define how Sherlock builds and presents profile links:

  • url – A template string containing {} as a username placeholder that becomes url_username_format in the SiteInformation object. Sherlock substitutes the target username before making HTTP requests.
  • urlMain – The base homepage URL used for display purposes in result listings and the SiteInformation.__str__ method.
  • urlProbe – An optional lightweight endpoint for existence checks that avoids loading full profile pages, used by platforms like GitHub and Codeforces that expose dedicated "user exists" APIs.

Detection Strategy Fields

The errorType field determines which validation strategy Sherlock applies, with support for three distinct detection modes:

  • status_code – Treats HTTP 200 responses as "claimed" profiles while non-200 status codes indicate unclaimed usernames. This is the fastest method requiring no response body parsing.
  • message – Searches the response body for strings defined in the errorMsg array; presence of any error message indicates the username is unclaimed.
  • response_url – Compares the final redirected URL against errorUrl to detect "not found" pages that redirect to generic error endpoints.

These strategies are implemented in sherlock_project/sherlock.py (lines 291-302, 351), where the engine branches based on the errorType value to interpret results correctly.

Request Configuration Options

Modern platforms often require specific HTTP methods or authentication headers, which the JSON structure accommodates through optional fields:

  • request_method – Specifies the HTTP verb (defaults to GET), with some platforms like Discord and Holopin requiring POST requests.
  • request_payload – JSON object or form data sent with POST requests, supporting username placeholder substitution within the payload structure.
  • headers – Dictionary of custom HTTP headers (e.g., Content-Type: application/json) inserted into the request before transmission.
  • regexCheck – A regular expression pattern that preemptively validates usernames before network requests, preventing unnecessary HTTP calls for invalid formats based on platform-specific constraints like character restrictions or length limits.

Safety and Validation Metadata

Additional fields ensure safe and accurate scanning:

  • isNSFW – Boolean flag indicating adult content that activates the SitesInformation.remove_nsfw_sites method when users opt out of NSFW results.
  • username_claimed – A known-valid username used by tests/test_manifest.py to validate that detection rules work correctly against live platforms.
  • username_unclaimed – A generated internal placeholder created by SiteInformation.__init__ for negative testing, though this is not stored in the JSON file itself.

How Detection Rules Are Executed

The separation between rule declaration and execution allows Sherlock to apply generic logic to platform-specific requirements. When you initiate a scan, Sherlock follows this workflow:

  1. Rule Loading – The SitesInformation class instantiates by fetching data.json from the repository or a user-provided path, creating a SiteInformation object for each entry that encapsulates the detection logic.

  2. Request Construction – For each target site, Sherlock checks regexCheck (if present), formats the url template with the target username, and constructs an HTTP request using the specified request_method, headers, and request_payload.

  3. Response Interpretation – Based on the errorType value, the engine examines status codes, searches response text for errorMsg strings, or compares redirect URLs against errorUrl to determine profile existence.

  4. Result Compilation – Sites returning "claimed" status populate the final report with formatted URLs constructed from the original templates.

Practical Implementation Examples

Loading and Inspecting Detection Rules

You can programmatically inspect how specific platforms define their detection criteria:

from sherlock_project.sites import SitesInformation

# Load live data.json from the official repo

sites_info = SitesInformation()

# Access the rule for GitHub

github_rule = sites_info.sites["GitHub"]
print(github_rule)                       # → GitHub (https://www.github.com/)

print(github_rule.url_username_format)   # → https://www.github.com/{}

print(github_rule.information["errorType"])   # → status_code

Manual Execution of POST-Based Detection

Some platforms require complex request configurations that the JSON structure handles declaratively:

import requests
from sherlock_project.sites import SitesInformation

sites = SitesInformation()
site = sites.sites["Discord"]          # uses POST + JSON payload

username = "someuser123"

# Build request according to the rule

payload = site.information["request_payload"]
payload = {k: v.replace("{}", username) for k, v in payload.items()}
resp = requests.post(site.url_username_format,
                     json=payload,
                     headers=site.information.get("headers", {}))

# Interpret according to errorType == "message"

claimed = not any(err in resp.text for err in site.information["errorMsg"])
print(f"{username} on Discord {'exists' if claimed else 'does not exist'}")

Filtering NSFW Sites Before Scanning

The JSON metadata enables content filtering without modifying the rule set:

from sherlock_project.sites import SitesInformation

sites = SitesInformation()
sites.remove_nsfw_sites(do_not_remove=["BongaCams"])   # keep a specific NSFW site

print(f"Remaining sites: {len(sites)}")

Summary

  • Centralized configuration – All detection logic resides in sherlock_project/resources/data.json, allowing hundreds of social networks to be defined without Python code changes.
  • Three validation strategies – The errorType field supports status_code, message, and response_url detection methods that accommodate different platform behaviors.
  • Flexible request building – Optional fields like request_method, request_payload, and headers handle APIs requiring POST data or authentication headers.
  • Pre-validation optimization – The regexCheck field prevents unnecessary HTTP requests by filtering invalid usernames before network calls.
  • Safety controls – Boolean flags like isNSFW integrate with SitesInformation.remove_nsfw_sites to respect content preferences.

Frequently Asked Questions

What is the difference between url and urlProbe in the detection rules?

The url field defines the public profile page template that Sherlock reports to users when a username is found, while urlProbe specifies an internal API endpoint optimized for existence checks. For example, GitHub's urlProbe might query the user API directly, reducing bandwidth and avoiding full HTML page loads when verifying account existence.

How does the regexCheck field improve scanning performance?

regexCheck contains a regular expression that usernames must match before Sherlock sends HTTP requests. Platforms like Twitter or Instagram have specific username requirements (alphanumeric characters, length limits, no special characters). By validating against these patterns first, Sherlock eliminates network overhead for usernames that could never exist on that platform.

Can I create a custom data.json file for internal company platforms?

Yes, the SitesInformation class accepts custom JSON file paths, allowing you to define detection rules for internal corporate networks or private platforms using the same schema. You must define the required fields (url, errorType) and optionally include urlMain, errorMsg, or request_headers depending on your authentication requirements and detection strategy.

How does Sherlock handle platforms that return HTTP 200 for missing profiles?

Platforms using errorType: message handle this scenario by searching the response body for specific text strings defined in errorMsg. Even if the server returns a 200 status code, presence of error text like "user not found" or "profile unavailable" indicates an unclaimed username, while absence of these strings suggests the profile exists.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →