How to Add a New Site to Sherlock's Detection List: Complete Configuration Guide

To add a new site to Sherlock, extend the SITE_DATA dictionary in sherlock_project/sites.py with the URL pattern and validation regex, add corresponding metadata to sherlock_project/resources/data.json, and run the validation scripts to ensure the new entry is correctly interpreted by the tool.

Sherlock discovers user accounts across the web by querying a centralized dictionary of site definitions stored in the sherlock-project/sherlock repository. Each entry contains the URL pattern, request method, username validation regex, and optional API-specific configuration required to probe for account existence. To add a new site to Sherlock's detection list, you must extend the site definition dictionary, update the validation data file, and verify your changes against the built-in test suite.

Extend the Site Definition Dictionary in sites.py

The core site definitions live in sherlock_project/sites.py. Open this file and locate the SITE_DATA dictionary. Add a new key-value pair following the existing structure, using a unique lowercase identifier as the key:


# sherlock_project/sites.py

SITE_DATA = {
    # … existing entries …

    "example": {                     # unique site identifier (lower-case, no spaces)

        "name": "Example",           # Human-readable name

        "url_main": "https://example.com",   # Base URL of the service

        "url_user": "https://example.com/{username}",  # Profile URL pattern

        "regex_check": r"^[a-zA-Z0-9_]{1,30}$",       # Username validation regex

        "request_method": "GET",      # HTTP method: GET, POST, etc.

        "request_headers": {},       # Optional custom headers

        "status_codes": [200, 404],  # Expected HTTP status codes

        "error_message": None,       # Optional error message pattern

    },
}

Key configuration fields to include when you add a new site to Sherlock:

  • url_user – Must include {username} as a placeholder for the target account name
  • regex_check – Validates usernames against the site's own rules (e.g., length, allowed characters)
  • request_method – Most sites use GET, but some APIs require POST or specific headers

Add Validation Metadata to data.json

Sherlock validates site entries against a JSON schema stored in sherlock_project/resources/data.json. Append a matching block that mirrors your Python definition:

{
  "example": {
    "type": "social",
    "name": "Example",
    "urlMain": "https://example.com",
    "urlUser": "https://example.com/{username}",
    "usernameRegex": "^[a-zA-Z0-9_]{1,30}$",
    "requestMethod": "GET",
    "requestHeaders": {}
  }
}

The devel/summarize_site_validation.py script compares these two files to ensure consistency across the codebase.

Run the Validation Script

Execute the built-in validation script from the repository root to check for schema errors, missing keys, or inconsistent regex patterns:

python devel/summarize_site_validation.py

Fix any reported discrepancies between sites.py and data.json before proceeding to testing.

Update and Run the Test Suite

The manifest test in tests/test_manifest.py automatically verifies that every site listed in resources/data.json has a corresponding entry in sites.py. Run the complete test suite to confirm your addition doesn't break existing functionality:

pytest -q

If the new site requires special handling (such as POST request bodies or custom error detection), add specific test cases to tests/test_probes.py to cover these edge cases.

Submit Your Changes

Once validation and tests pass, commit your changes following the project's contribution guidelines:

git checkout -b add-example-site
git add sherlock_project/sites.py sherlock_project/resources/data.json
git commit -m "Add Example.com to Sherlock site list"
git push origin add-example-site

The CI pipeline will automatically re-run validation and tests when you open a pull request on GitHub.

Summary

  • Extend sherlock_project/sites.py by adding a new entry to the SITE_DATA dictionary with URL patterns, regex validation, and HTTP configuration
  • Update sherlock_project/resources/data.json with matching metadata to satisfy the validation schema
  • Run python devel/summarize_site_validation.py to verify consistency between Python definitions and JSON metadata
  • Execute pytest -q to ensure the manifest tests pass and the new site is correctly integrated
  • Commit and push to a feature branch for review via pull request

Frequently Asked Questions

What is the main file for adding a new site to Sherlock?

The primary file is sherlock_project/sites.py, which contains the SITE_DATA dictionary where all supported platforms are defined. Each key represents a unique site identifier used on the command line, and each value is a dictionary containing the URL patterns, validation regexes, and request configuration required to probe for account existence.

How does Sherlock validate new site configurations?

Sherlock uses the devel/summarize_site_validation.py script to cross-reference the Python definitions in sites.py against the JSON schema in sherlock_project/resources/data.json. This validation checks for missing required keys, inconsistent regex patterns, unsupported HTTP methods, and URL placeholder formatting to ensure every new site is properly configured before merging.

What is the purpose of the data.json file in Sherlock?

The sherlock_project/resources/data.json file serves as the canonical schema for site metadata, enabling automated validation and consistency checks. It mirrors the data in sites.py but in JSON format, allowing external tools and the validation script to parse site definitions without executing Python code, and ensuring that all entries conform to the expected structure.

Can I add sites that require POST requests or custom headers?

Yes. When you add a new site to Sherlock's detection list, set request_method to "POST" and populate request_headers with required values such as custom User-Agents or authentication tokens. If the site requires a POST body or complex authentication, add dedicated test cases in tests/test_probes.py to verify the probe behavior works correctly with these custom configurations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →