# SitesInformation Class in Sherlock: How It Loads and Manages Site Data

> Discover the SitesInformation class in Sherlock. Learn how it loads, validates, and manages site data from JSON manifests to build a queryable collection of SiteInformation objects.

- Repository: [Sherlock/sherlock](https://github.com/sherlock-project/sherlock)
- Tags: internals
- Published: 2026-03-02

---

**The `SitesInformation` class is the central data loader in Sherlock that fetches, validates, filters, and instantiates site definitions from JSON manifests to create a queryable collection of `SiteInformation` objects.**

The `SitesInformation` class serves as the backbone of the Sherlock username enumeration tool, handling the entire lifecycle of site metadata from download to exposure. Defined in [`sherlock_project/sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sites.py), this class aggregates thousands of website definitions into a structured format that the probing engine uses to check username availability across platforms. Understanding how this class loads and manages site data is essential for anyone customizing Sherlock or contributing new site definitions to the sherlock-project/sherlock repository.

## What Is the SitesInformation Class?

The `SitesInformation` class (defined at **line 78** in [`sherlock_project/sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sites.py)) acts as the central registry for all supported websites. It transforms raw JSON manifests into typed `SiteInformation` objects that Sherlock uses to construct probe URLs and interpret HTTP responses. By abstracting the data layer, the class ensures that downstream components receive validated, consistent site metadata without needing to handle file I/O or JSON parsing directly.

## How SitesInformation Loads Site Data

The class implements a robust loading pipeline that handles both remote and local data sources, applies security exclusions, and validates data integrity before instantiation.

### Source Selection and Path Validation

When instantiated, `SitesInformation` determines which JSON manifest to load based on the `data_file_path` parameter. If no path is provided, the class falls back to the live ** `MANIFEST_URL`** constant defined at **lines 11–13** in [`sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sites.py). The constructor enforces a strict file extension check: only paths ending with **`.json`** are accepted, and any other extension raises a `FileNotFoundError` (**lines 24–26**).

### Fetching and Parsing JSON Manifests

The class distinguishes between remote and local sources using a simple protocol check:

- **Remote JSON**: If the path starts with `http`, the class performs a `GET` request using `requests.get` with a **30‑second timeout** (**lines 29–38**).
- **Local JSON**: For filesystem paths, the class opens the file and loads it using `json.load` (**lines 51–57**).

Parsing errors are captured and re‑raised as descriptive `ValueError` exceptions that identify the offending source file or URL (**lines 42–48**), preventing Sherlock from running with corrupted site data.

### Applying False‑Positive Exclusions

If `honor_exclusions=True` is passed during initialization, the class downloads the false‑positive exclusions list from ** `EXCLUSIONS_URL`** (**line 12**). Each line in this list is stripped and the corresponding entry is removed from the loaded manifest via `site_data.pop(exclusion, None)` (**lines 70–82**). Callers can protect specific sites from removal using the `do_not_exclude` argument, which is useful when you know a flagged site is actually reliable for your specific use case.

## Managing Site Data and Site Objects

After loading the raw JSON, `SitesInformation` converts the dictionary into a collection of strongly‑typed objects and provides convenient filtering methods.

### Materializing SiteInformation Instances

The class iterates over every entry in the parsed JSON and creates a **`SiteInformation`** instance for each site (**lines 92–104**). These objects are stored in an internal dictionary `self.sites` keyed by the site’s canonical name. Missing mandatory fields trigger descriptive `ValueError` exceptions, while malformed entries are logged and skipped to prevent total pipeline failure (**lines 104–110**).

### Collection API and Utility Methods

`SitesInformation` implements Python’s collection protocol to simplify downstream usage:

- **`remove_nsfw_sites(do_not_remove=[])`**: Filters out sites flagged as NSFW unless explicitly exempted (**lines 123–131**).
- **`site_name_list()`**: Returns a case‑insensitive sorted list of all site names (**lines 131–140**).
- **Iterator Support**: The class implements `__iter__`, `__len__`, and `__getitem__` (**lines 144–166**), allowing you to use `len(sites)`, iterate with `for site in sites`, or access specific sites via dictionary‑style indexing.

## Practical Code Examples

The following examples demonstrate common patterns for working with the `SitesInformation` class in Sherlock:

```python

# Load the default (live) site list from MANIFEST_URL

from sherlock_project.sites import SitesInformation

sites = SitesInformation()                 # pulls the remote data.json

print(f"Total supported sites: {len(sites)}")   # __len__ works

print(sites.site_name_list()[:10])               # first 10 names

```

```python

# Load a local manifest and bypass exclusion handling

local_path = "my_custom_sites.json"         # must end with .json

sites = SitesInformation(data_file_path=local_path, honor_exclusions=False)

# Iterate over the concrete SiteInformation objects

for site in sites:
    print(f"{site.name}: {site.url_home} (NSFW={site.is_nsfw})")

```

```python

# Exclude NSFW sites but protect a specific platform

sites = SitesInformation()
sites.remove_nsfw_sites(do_not_remove=["OnlyFans"])   # keep OnlyFans even if NSFW

print(f"After filtering: {len(sites)} sites remain")

```

## Summary

- The **`SitesInformation`** class in [`sherlock_project/sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sites.py) is the central data aggregator for Sherlock, defined at **line 78**.
- It automatically selects between a user‑provided `data_file_path` and the remote **`MANIFEST_URL`**, rejecting non‑JSON files.
- Remote manifests are fetched with a 30‑second timeout, while local files use standard `json.load()`, with all errors wrapped in `ValueError`.
- The **`honor_exclusions`** flag controls whether sites listed in **`EXCLUSIONS_URL`** are filtered out during initialization.
- Raw JSON entries are converted to **`SiteInformation`** objects stored in `self.sites`, with full collection protocol support (`__len__`, `__iter__`, `__getitem__`).
- Utility methods like **`remove_nsfw_sites()`** and **`site_name_list()`** provide convenient filtering and introspection capabilities.

## Frequently Asked Questions

### Where is the SitesInformation class defined in the Sherlock repository?

The `SitesInformation` class is defined in **[`sherlock_project/sites.py`](https://github.com/sherlock-project/sherlock/blob/main/sherlock_project/sites.py)** starting at **line 78**. This file also contains the `SiteInformation` data class and the constant definitions for `MANIFEST_URL` and `EXCLUSIONS_URL` used during initialization.

### How does SitesInformation handle remote versus local JSON manifests?

The class checks if the `data_file_path` starts with `http`. If so, it uses `requests.get` with a 30‑second timeout to fetch the remote JSON. Otherwise, it treats the path as a local filesystem location and loads it using `open()` and `json.load()`. Both paths enforce a `.json` file extension requirement.

### What is the purpose of the honor_exclusions parameter?

When `honor_exclusions=True` (the default), the class downloads the false‑positive exclusions list from `EXCLUSIONS_URL` and removes those sites from the loaded manifest. This prevents Sherlock from querying platforms known to generate unreliable results. You can protect specific sites from this filter using the `do_not_exclude` argument.

### How can I filter NSFW sites when using SitesInformation?

After instantiation, call the **`remove_nsfw_sites()`** method and optionally pass a list of site names to the `do_not_remove` parameter to preserve specific platforms. The method modifies the internal `self.sites` dictionary in‑place, removing any sites where the `is_nsfw` flag is set to `True`.