How to Add a New Site to Sherlock's Detection List: Complete Configuration Guide
To add a new site to Sherlock, extend the SITE_DATA dictionary in sherlock_project/sites.py with the URL pattern and validation regex, add corresponding metadata to sherlock_project/resources/data.json, and run the validation scripts to ensure the new entry is correctly interpreted by the tool.
Sherlock discovers user accounts across the web by querying a centralized dictionary of site definitions stored in the sherlock-project/sherlock repository. Each entry contains the URL pattern, request method, username validation regex, and optional API-specific configuration required to probe for account existence. To add a new site to Sherlock's detection list, you must extend the site definition dictionary, update the validation data file, and verify your changes against the built-in test suite.
Extend the Site Definition Dictionary in sites.py
The core site definitions live in sherlock_project/sites.py. Open this file and locate the SITE_DATA dictionary. Add a new key-value pair following the existing structure, using a unique lowercase identifier as the key:
# sherlock_project/sites.py
SITE_DATA = {
# … existing entries …
"example": { # unique site identifier (lower-case, no spaces)
"name": "Example", # Human-readable name
"url_main": "https://example.com", # Base URL of the service
"url_user": "https://example.com/{username}", # Profile URL pattern
"regex_check": r"^[a-zA-Z0-9_]{1,30}$", # Username validation regex
"request_method": "GET", # HTTP method: GET, POST, etc.
"request_headers": {}, # Optional custom headers
"status_codes": [200, 404], # Expected HTTP status codes
"error_message": None, # Optional error message pattern
},
}
Key configuration fields to include when you add a new site to Sherlock:
url_user– Must include{username}as a placeholder for the target account nameregex_check– Validates usernames against the site's own rules (e.g., length, allowed characters)request_method– Most sites useGET, but some APIs requirePOSTor specific headers
Add Validation Metadata to data.json
Sherlock validates site entries against a JSON schema stored in sherlock_project/resources/data.json. Append a matching block that mirrors your Python definition:
{
"example": {
"type": "social",
"name": "Example",
"urlMain": "https://example.com",
"urlUser": "https://example.com/{username}",
"usernameRegex": "^[a-zA-Z0-9_]{1,30}$",
"requestMethod": "GET",
"requestHeaders": {}
}
}
The devel/summarize_site_validation.py script compares these two files to ensure consistency across the codebase.
Run the Validation Script
Execute the built-in validation script from the repository root to check for schema errors, missing keys, or inconsistent regex patterns:
python devel/summarize_site_validation.py
Fix any reported discrepancies between sites.py and data.json before proceeding to testing.
Update and Run the Test Suite
The manifest test in tests/test_manifest.py automatically verifies that every site listed in resources/data.json has a corresponding entry in sites.py. Run the complete test suite to confirm your addition doesn't break existing functionality:
pytest -q
If the new site requires special handling (such as POST request bodies or custom error detection), add specific test cases to tests/test_probes.py to cover these edge cases.
Submit Your Changes
Once validation and tests pass, commit your changes following the project's contribution guidelines:
git checkout -b add-example-site
git add sherlock_project/sites.py sherlock_project/resources/data.json
git commit -m "Add Example.com to Sherlock site list"
git push origin add-example-site
The CI pipeline will automatically re-run validation and tests when you open a pull request on GitHub.
Summary
- Extend
sherlock_project/sites.pyby adding a new entry to theSITE_DATAdictionary with URL patterns, regex validation, and HTTP configuration - Update
sherlock_project/resources/data.jsonwith matching metadata to satisfy the validation schema - Run
python devel/summarize_site_validation.pyto verify consistency between Python definitions and JSON metadata - Execute
pytest -qto ensure the manifest tests pass and the new site is correctly integrated - Commit and push to a feature branch for review via pull request
Frequently Asked Questions
What is the main file for adding a new site to Sherlock?
The primary file is sherlock_project/sites.py, which contains the SITE_DATA dictionary where all supported platforms are defined. Each key represents a unique site identifier used on the command line, and each value is a dictionary containing the URL patterns, validation regexes, and request configuration required to probe for account existence.
How does Sherlock validate new site configurations?
Sherlock uses the devel/summarize_site_validation.py script to cross-reference the Python definitions in sites.py against the JSON schema in sherlock_project/resources/data.json. This validation checks for missing required keys, inconsistent regex patterns, unsupported HTTP methods, and URL placeholder formatting to ensure every new site is properly configured before merging.
What is the purpose of the data.json file in Sherlock?
The sherlock_project/resources/data.json file serves as the canonical schema for site metadata, enabling automated validation and consistency checks. It mirrors the data in sites.py but in JSON format, allowing external tools and the validation script to parse site definitions without executing Python code, and ensuring that all entries conform to the expected structure.
Can I add sites that require POST requests or custom headers?
Yes. When you add a new site to Sherlock's detection list, set request_method to "POST" and populate request_headers with required values such as custom User-Agents or authentication tokens. If the site requires a POST body or complex authentication, add dedicated test cases in tests/test_probes.py to verify the probe behavior works correctly with these custom configurations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →