# Maigret Site Database Structure: Complete Guide to data.json Fields

> Explore the Maigret site database structure in data.json. Understand URL patterns, detection methods, and metadata tags for efficient username checking with Maigret.

- Repository: [Soxoj/maigret](https://github.com/soxoj/maigret)
- Tags: api-reference
- Published: 2026-04-30

---

**Maigret stores all username checking rules in a single JSON file at [`maigret/resources/data.json`](https://github.com/soxoj/maigret/blob/main/maigret/resources/data.json), where each site entry defines URL patterns, detection methods, and metadata tags that determine how the tool probes for account existence.**

The `soxoj/maigret` repository maintains its entire site database in a structured JSON format that powers the username enumeration engine. Understanding the maigret site database structure is essential for contributing new sites, troubleshooting false positives, or customizing detection logic. This guide examines every field in the [`data.json`](https://github.com/soxoj/maigret/blob/main/data.json) schema according to the source code implementation.

## Top-Level Structure of data.json

The database file contains a single top-level key `"sites"` that maps site names to their configuration objects.

```json
{
  "sites": {
    "Twitter": {
      "url": "https://twitter.com/{username}",
      "urlMain": "https://twitter.com/",
      "checkType": "status_code",
      ...
    },
    "GitHub": {
      ...
    }
  }
}

```

Each value under `"sites"` is a dictionary where fields are optional unless required by the specific `checkType` or detection method being used. The schema interpretation logic resides in [`maigret/utils/site_check.py`](https://github.com/soxoj/maigret/blob/main/maigret/utils/site_check.py).

## Core Detection Fields

### URL and Endpoint Configuration

Three fields control where and how Maigret sends HTTP requests:

- **`url`**: The public profile page template containing the `{username}` placeholder (e.g., `"https://twitter.com/{username}"`). This is always required.
- **`urlMain`**: The base service URL used for reporting and browser links.
- **`urlProbe`**: An optional alternative endpoint (often an API or JSON endpoint) that Maigret fetches instead of `url` when present. The `{username}` placeholder is substituted here as well.

### Detection Method (checkType)

The **`checkType`** field determines the existence logic and supports three values as documented in `docs/source/development.rst`:

- **`status_code`**: Username exists if the HTTP response status equals 200.
- **`message`**: Existence is determined by scanning the response body for specific strings.
- **`response_url`**: Detection relies on the final URL after any redirects.

When `checkType` is set to `message` or `response_url`, the optional **`errorUrl`** field specifies a generic error or login page URL that indicates non-existence.

### Content Matching Strings

For `checkType: "message"`, Maigret evaluates two array fields:

- **`presenseStrs`**: Array of substrings whose presence in the response body indicates the username **exists**.
- **`absenceStrs`**: Array of substrings indicating the username **does not exist** or that the page is a generic "not found" placeholder.

## Request and Validation Configuration

### Custom Headers and Regex Validation

- **`headers`**: An object specifying HTTP headers sent with the probe request, such as custom `User-Agent` strings required by specific services.
- **`regexCheck`**: A regular expression pattern that valid usernames must match before Maigret makes any request. This prevents unnecessary API calls for invalid username formats.

### Account Activation

The **`activation`** object contains fields required for sites needing temporary session creation before probing:

- **`url`**: The login or activation endpoint.
- **`method`**: The HTTP method or action to perform.
- **`marks`**: Marker strings used to confirm successful activation.

This configuration is used when services like Vimeo require session initialization before username checking.

## Optional Metadata Fields

### Site Classification and Ranking

- **`tags`**: Array of human-readable categories (e.g., `"social"`, `"video"`, `"apps"`) used for filtering sites with the `--tag` CLI argument and for report generation.
- **`alexaRank`**: Integer representing historic site popularity (maintained for backward compatibility; now derived from the Majestic Million dataset).
- **`similarSearch`**: Boolean indicating whether Maigret should perform similar search queries to improve detection (used for Adobe Community and similar services).

### Testing and Validation

- **`usernameClaimed`**: A known existing username used for self-tests and documentation.
- **`usernameUnclaimed`**: A known non-existing username used to verify detection accuracy.

### Error Handling

The **`errors`** object maps error identifiers (e.g., `"tiktok-verify-page"`) to human-readable messages, surfacing special failure modes such as captchas or rate limits.

## Programmatically Loading the Database

Maigret loads the site database using standard JSON parsing as implemented in [`maigret/utils/site_check.py`](https://github.com/soxoj/maigret/blob/main/maigret/utils/site_check.py):

```python
from pathlib import Path
import json

db_path = Path(__file__).parent.parent / "maigret" / "resources" / "data.json"

with db_path.open(encoding="utf-8") as f:
    site_db = json.load(f)  # {'sites': { ... }}

```

Once loaded, `site_db["sites"]` provides dictionary access to all site configurations.

## Practical Usage Examples

### Filter Sites by Tag

Load the database and filter for social media platforms:

```python
from pathlib import Path
import json

db_path = Path(__file__).parent / "maigret" / "resources" / "data.json"
with db_path.open(encoding="utf-8") as f:
    data = json.load(f)

social_sites = [
    name for name, cfg in data["sites"].items()
    if "social" in cfg.get("tags", [])
]
print("Social sites:", social_sites)

```

### Simulate a Manual Check

Replicate Maigret's checking logic for a specific site:

```python
import requests

site_cfg = data["sites"]["Twitter"]
probe_url = site_cfg.get("urlProbe", site_cfg["url"]).replace("{username}", "jack")
resp = requests.get(probe_url, headers=site_cfg.get("headers", {}))

if site_cfg["checkType"] == "status_code":
    exists = resp.status_code == 200
elif site_cfg["checkType"] == "message":
    body = resp.text
    exists = any(s in body for s in site_cfg.get("presenseStrs", [])) and \
             not any(s in body for s in site_cfg.get("absenceStrs", []))
print("Twitter user 'jack' exists?", exists)

```

### Generate Rankings Report

Extract the top sites by Alexa rank:

```python
top_five = sorted(
    data["sites"].items(),
    key=lambda kv: kv[1].get("alexaRank", 10**9)
)[:5]

for name, cfg in top_five:
    print(f"{name:<20} rank={cfg.get('alexaRank')}")

```

## Key Files in the Repository

Understanding the maigret site database structure requires familiarity with these files:

- **[`maigret/resources/data.json`](https://github.com/soxoj/maigret/blob/main/maigret/resources/data.json)**: The canonical site database containing all checking rules.
- **[`maigret/utils/site_check.py`](https://github.com/soxoj/maigret/blob/main/maigret/utils/site_check.py)**: Loader and validator that interprets fields in [`data.json`](https://github.com/soxoj/maigret/blob/main/data.json).
- **[`maigret/utils/generate_db_meta.py`](https://github.com/soxoj/maigret/blob/main/maigret/utils/generate_db_meta.py)**: Generates [`db_meta.json`](https://github.com/soxoj/maigret/blob/main/db_meta.json) for the auto-update system.
- **`docs/source/development.rst`**: Documentation of supported `checkType` values and schema specifications.
- **[`maigret/maigret.py`](https://github.com/soxoj/maigret/blob/main/maigret/maigret.py)**: Core CLI entry point that consumes the database.

## Summary

- **[`maigret/resources/data.json`](https://github.com/soxoj/maigret/blob/main/maigret/resources/data.json)** stores all site definitions under a single `"sites"` key mapped as `"SiteName": {config}`.
- **Required fields** include `url`, `urlMain`, and `checkType`, while `urlProbe`, `headers`, and `regexCheck` provide request customization.
- **Detection logic** varies by `checkType`: `status_code` checks HTTP status, `message` scans for `presenseStrs`/`absenceStrs`, and `response_url` evaluates final redirect destinations.
- **Metadata fields** like `tags`, `alexaRank`, and `usernameClaimed` support filtering, ranking, and validation workflows.
- **Programmatic access** follows the loader pattern in [`maigret/utils/site_check.py`](https://github.com/soxoj/maigret/blob/main/maigret/utils/site_check.py) using standard JSON parsing.

## Frequently Asked Questions

### Where is the Maigret site database file located?

The database is stored at **[`maigret/resources/data.json`](https://github.com/soxoj/maigret/blob/main/maigret/resources/data.json)** in the repository root. This JSON file contains the complete schema defining how Maigret checks for usernames across all supported services.

### What is the difference between the `url` and `urlProbe` fields?

The **`url`** field specifies the public profile page template shown to users, while **`urlProbe`** defines an optional alternative endpoint (often an API or lighter HTML page) that Maigret actually requests to minimize bandwidth and avoid bot detection. If `urlProbe` is absent, Maigret falls back to using `url`.

### How does the `checkType` field determine username existence?

The **`checkType`** field specifies the detection algorithm: `status_code` considers a 200 response as existing, `message` searches the response body for strings defined in `presenseStrs` and `absenceStrs`, and `response_url` examines the final URL after redirects to see if it matches `errorUrl` or the expected pattern.

### What are `presenseStrs` and `absenceStrs` used for in the database?

These arrays define substring markers for the `message` check type. **`presenseStrs`** contains text fragments that must appear in the response to confirm existence (e.g., "Followers"), while **`absenceStrs`** contains fragments indicating a negative result or generic "not found" page (e.g., "Sorry, that page doesn't exist").