Sherlock False Positive Exclusions System: Configuration and Implementation Guide
Sherlock's False Positive Exclusions system automatically suppresses unreliable sites from username searches by downloading a curated exclusion list from the repository's exclusions branch and filtering the site manifest at runtime.
The Sherlock Project maintains a dynamic False Positive Exclusions mechanism to prevent erroneous "found" results from unreliable platforms. This system references a curated list stored in the repository's exclusions branch, automatically suppressing sites known to generate false positives while allowing users to override the behavior via command-line flags or programmatic arguments.
How the False Positive Exclusions List is Stored and Accessed
The exclusion list resides in a plain-text file named false_positive_exclusions.txt on the dedicated exclusions branch of the sherlock-project/sherlock repository. At runtime, Sherlock fetches this file from the raw GitHub URL defined in sherlock_project/sites.py as the constant EXCLUSIONS_URL:
EXCLUSIONS_URL = "https://raw.githubusercontent.com/sherlock-project/sherlock/refs/heads/exclusions/false_positive_exclusions.txt"
This URL is accessed during the initialization of the SitesInformation class to retrieve the current list of sites to suppress.
Runtime Application of False Positive Exclusions
When SitesInformation is instantiated, it conditionally applies the exclusions based on the honor_exclusions parameter, which defaults to True. The logic, found in sherlock_project/sites.py around lines 167-184, performs the following steps:
- Checks if
honor_exclusionsis enabled. - Performs an HTTP GET request to
EXCLUSIONS_URLwith a 10-second timeout. - Parses the response into a list of site names, stripping whitespace.
- Removes each listed site from the loaded site manifest using
site_data.pop(exclusion, None).
if honor_exclusions:
response = requests.get(url=EXCLUSIONS_URL, timeout=10)
if response.status_code == 200:
exclusions = [e.strip() for e in response.text.splitlines()]
for exclusion in exclusions:
site_data.pop(exclusion, None)
If honor_exclusions is set to False, the manifest loads without any filtering, including all sites regardless of their false-positive status.
Configuring False Positive Exclusions Behavior
Users can override the default exclusion behavior through both command-line interfaces and programmatic APIs.
Command-Line Override
To bypass the remote exclusion list entirely, use the --ignore-exclusions flag when running Sherlock. This flag sets honor_exclusions=False internally, forcing the tool to search against all sites in the manifest.
sherlock alice --ignore-exclusions
This argument is defined in sherlock_project/sherlock.py (lines 96-101) and consumed at line 772 during the initialization of the search logic.
Programmatic Override
When using Sherlock as a library, the SitesInformation constructor accepts a do_not_exclude argument. This parameter accepts a list of site names that should remain in the manifest even if they appear in the remote exclusion list.
from sherlock_project.sites import SitesInformation
# Load the manifest but preserve Twitter even if it's in the exclusion list
sites = SitesInformation(do_not_exclude=["Twitter"])
This logic is handled in SitesInformation.__init__ at lines 74-77, where the specified sites are removed from the exclusion set before the filtering is applied.
Automatic Maintenance via GitHub Actions
The exclusion list is not manually curated in isolation; it is automatically generated and updated through a GitHub Actions workflow defined in .github/workflows/exclusions.yml. This workflow triggers on every push to the main branch and performs the following:
- Extracts sites marked with the
falsePositiveflag fromsherlock_project/resources/data.json. - Generates an updated
false_positive_exclusions.txtfile. - Commits the file to the
exclusionsbranch, ensuring the remote URL always serves the current list.
The workflow steps are detailed in lines 10-89 of the configuration file, ensuring the exclusion mechanism remains synchronized with the latest site definitions without requiring manual intervention.
Summary
- Remote Storage: The exclusion list lives in
false_positive_exclusions.txton theexclusionsbranch, fetched fromEXCLUSIONS_URLinsherlock_project/sites.py. - Runtime Filtering: By default,
SitesInformationdownloads the list and removes matching sites from the manifest usingsite_data.pop(exclusion, None). - User Control: Use
--ignore-exclusionsto bypass filtering entirely, or passdo_not_exclude=["SiteName"]to the Python API to protect specific sites. - Auto-Updates: A GitHub Actions workflow (
.github/workflows/exclusions.yml) regenerates the list from sites markedfalsePositiveindata.jsonon every push to main.
Frequently Asked Questions
How do I completely disable the False Positive Exclusions system?
Pass the --ignore-exclusions flag when running Sherlock from the command line. This sets honor_exclusions=False, causing the tool to load the full site manifest without filtering out any sites listed in the remote exclusion file.
Can I keep specific sites in the search even if they are on the exclusion list?
Yes. When using Sherlock programmatically, instantiate SitesInformation with the do_not_exclude parameter set to a list of site names you wish to preserve. These sites are removed from the exclusion set before the manifest is filtered, ensuring they remain available for searching.
Where does the exclusion list come from and how is it updated?
The list is automatically generated from the site manifest (data.json) by a GitHub Actions workflow defined in .github/workflows/exclusions.yml. This workflow runs on every push to the main branch, extracts sites marked with the falsePositive flag, and commits the updated false_positive_exclusions.txt file to the exclusions branch.
What happens if the remote exclusion file is unreachable?
If the HTTP request to EXCLUSIONS_URL fails or returns a non-200 status code, Sherlock skips the exclusion step and proceeds with the full site manifest. The tool does not halt execution; it gracefully degrades by including all sites, effectively behaving as if --ignore-exclusions was set for that session.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →