SitesInformation Class in Sherlock: How It Loads and Manages Site Data
The SitesInformation class is the central data loader in Sherlock that fetches, validates, filters, and instantiates site definitions from JSON manifests to create a queryable collection of SiteInformation objects.
The SitesInformation class serves as the backbone of the Sherlock username enumeration tool, handling the entire lifecycle of site metadata from download to exposure. Defined in sherlock_project/sites.py, this class aggregates thousands of website definitions into a structured format that the probing engine uses to check username availability across platforms. Understanding how this class loads and manages site data is essential for anyone customizing Sherlock or contributing new site definitions to the sherlock-project/sherlock repository.
What Is the SitesInformation Class?
The SitesInformation class (defined at line 78 in sherlock_project/sites.py) acts as the central registry for all supported websites. It transforms raw JSON manifests into typed SiteInformation objects that Sherlock uses to construct probe URLs and interpret HTTP responses. By abstracting the data layer, the class ensures that downstream components receive validated, consistent site metadata without needing to handle file I/O or JSON parsing directly.
How SitesInformation Loads Site Data
The class implements a robust loading pipeline that handles both remote and local data sources, applies security exclusions, and validates data integrity before instantiation.
Source Selection and Path Validation
When instantiated, SitesInformation determines which JSON manifest to load based on the data_file_path parameter. If no path is provided, the class falls back to the live ** MANIFEST_URL** constant defined at lines 11–13 in sites.py. The constructor enforces a strict file extension check: only paths ending with .json are accepted, and any other extension raises a FileNotFoundError (lines 24–26).
Fetching and Parsing JSON Manifests
The class distinguishes between remote and local sources using a simple protocol check:
- Remote JSON: If the path starts with
http, the class performs aGETrequest usingrequests.getwith a 30‑second timeout (lines 29–38). - Local JSON: For filesystem paths, the class opens the file and loads it using
json.load(lines 51–57).
Parsing errors are captured and re‑raised as descriptive ValueError exceptions that identify the offending source file or URL (lines 42–48), preventing Sherlock from running with corrupted site data.
Applying False‑Positive Exclusions
If honor_exclusions=True is passed during initialization, the class downloads the false‑positive exclusions list from ** EXCLUSIONS_URL** (line 12). Each line in this list is stripped and the corresponding entry is removed from the loaded manifest via site_data.pop(exclusion, None) (lines 70–82). Callers can protect specific sites from removal using the do_not_exclude argument, which is useful when you know a flagged site is actually reliable for your specific use case.
Managing Site Data and Site Objects
After loading the raw JSON, SitesInformation converts the dictionary into a collection of strongly‑typed objects and provides convenient filtering methods.
Materializing SiteInformation Instances
The class iterates over every entry in the parsed JSON and creates a SiteInformation instance for each site (lines 92–104). These objects are stored in an internal dictionary self.sites keyed by the site’s canonical name. Missing mandatory fields trigger descriptive ValueError exceptions, while malformed entries are logged and skipped to prevent total pipeline failure (lines 104–110).
Collection API and Utility Methods
SitesInformation implements Python’s collection protocol to simplify downstream usage:
remove_nsfw_sites(do_not_remove=[]): Filters out sites flagged as NSFW unless explicitly exempted (lines 123–131).site_name_list(): Returns a case‑insensitive sorted list of all site names (lines 131–140).- Iterator Support: The class implements
__iter__,__len__, and__getitem__(lines 144–166), allowing you to uselen(sites), iterate withfor site in sites, or access specific sites via dictionary‑style indexing.
Practical Code Examples
The following examples demonstrate common patterns for working with the SitesInformation class in Sherlock:
# Load the default (live) site list from MANIFEST_URL
from sherlock_project.sites import SitesInformation
sites = SitesInformation() # pulls the remote data.json
print(f"Total supported sites: {len(sites)}") # __len__ works
print(sites.site_name_list()[:10]) # first 10 names
# Load a local manifest and bypass exclusion handling
local_path = "my_custom_sites.json" # must end with .json
sites = SitesInformation(data_file_path=local_path, honor_exclusions=False)
# Iterate over the concrete SiteInformation objects
for site in sites:
print(f"{site.name}: {site.url_home} (NSFW={site.is_nsfw})")
# Exclude NSFW sites but protect a specific platform
sites = SitesInformation()
sites.remove_nsfw_sites(do_not_remove=["OnlyFans"]) # keep OnlyFans even if NSFW
print(f"After filtering: {len(sites)} sites remain")
Summary
- The
SitesInformationclass insherlock_project/sites.pyis the central data aggregator for Sherlock, defined at line 78. - It automatically selects between a user‑provided
data_file_pathand the remoteMANIFEST_URL, rejecting non‑JSON files. - Remote manifests are fetched with a 30‑second timeout, while local files use standard
json.load(), with all errors wrapped inValueError. - The
honor_exclusionsflag controls whether sites listed inEXCLUSIONS_URLare filtered out during initialization. - Raw JSON entries are converted to
SiteInformationobjects stored inself.sites, with full collection protocol support (__len__,__iter__,__getitem__). - Utility methods like
remove_nsfw_sites()andsite_name_list()provide convenient filtering and introspection capabilities.
Frequently Asked Questions
Where is the SitesInformation class defined in the Sherlock repository?
The SitesInformation class is defined in sherlock_project/sites.py starting at line 78. This file also contains the SiteInformation data class and the constant definitions for MANIFEST_URL and EXCLUSIONS_URL used during initialization.
How does SitesInformation handle remote versus local JSON manifests?
The class checks if the data_file_path starts with http. If so, it uses requests.get with a 30‑second timeout to fetch the remote JSON. Otherwise, it treats the path as a local filesystem location and loads it using open() and json.load(). Both paths enforce a .json file extension requirement.
What is the purpose of the honor_exclusions parameter?
When honor_exclusions=True (the default), the class downloads the false‑positive exclusions list from EXCLUSIONS_URL and removes those sites from the loaded manifest. This prevents Sherlock from querying platforms known to generate unreliable results. You can protect specific sites from this filter using the do_not_exclude argument.
How can I filter NSFW sites when using SitesInformation?
After instantiation, call the remove_nsfw_sites() method and optionally pass a list of site names to the do_not_remove parameter to preserve specific platforms. The method modifies the internal self.sites dictionary in‑place, removing any sites where the is_nsfw flag is set to True.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →