Sherlock urlProbe vs url: Understanding Site Configuration URL Fields

The urlProbe field specifies an internal API endpoint for checking username existence, while the url field defines the public profile link displayed in Sherlock's final report.

In the sherlock-project/sherlock repository, every site entry in data.json configures two distinct URL fields that serve different purposes during the username investigation process. Understanding the difference between urlProbe and url is essential for interpreting search results correctly and optimizing site configurations for reliable detection.

What is the Difference Between urlProbe and url?

Each site definition in Sherlock's configuration contains two URL-related properties with distinct roles:

  • url – The public profile page where users normally view an account (e.g., https://anilist.co/user/{}). Sherlock displays this link in the final report to users.
  • urlProbe – An optional internal endpoint used specifically for the HTTP request that checks username availability. This often points to a lightweight API or GraphQL endpoint.

When urlProbe is absent from a site's configuration, Sherlock automatically falls back to using the url value for both probing and display purposes. This fallback logic ensures backward compatibility while allowing advanced configurations to separate the user-facing URL from the detection endpoint.

How Sherlock Implements the URL Selection Logic

The distinction between these fields is implemented in sherlock_project/sherlock.py, where the probing logic determines which endpoint to request:


# sherlock_project/sherlock.py – URL selection excerpt

url = interpolate_string(net_info["url"], username.replace(' ', '%20'))

url_probe = net_info.get("urlProbe")
if url_probe is None:
    # Probe URL is normal one seen by people out on the web.

    url_probe = url
else:
    # There is a special URL for probing existence separate

    # from where the user profile normally can be found.

    url_probe = interpolate_string(url_probe, username)

In this implementation, url is always interpolated and stored as results_site["url_user"] for the final report, regardless of which endpoint was probed. The url_probe variable determines the actual HTTP request target, allowing Sherlock to check username existence via a private API while still presenting the standard public profile URL to investigators.

Real-World Configuration Examples

Sites Using Both Fields

The Anilist entry in sherlock_project/resources/data.json demonstrates a complete implementation using both properties:

{
  "Anilist": {
    "errorType": "response_url",
    "url": "https://anilist.co/user/{}/",
    "urlMain": "https://anilist.co/",
    "urlProbe": "https://graphql.anilist.co/",
    "username_claimed": "Josh"
  }
}

Here, url points to the visible user profile page, while urlProbe targets the GraphQL API endpoint that returns lightweight JSON responses for existence checks.

Sites Using Only the url Field

Conversely, the Apple Developer entry omits urlProbe entirely, causing Sherlock to use the same URL for both probing and display:

{
  "Apple Developer": {
    "errorType": "status_code",
    "url": "https://developer.apple.com/forums/profile/{}",
    "urlMain": "https://developer.apple.com",
    "username_claimed": "lio24d"
  }
}

In this case, Sherlock sends the HTTP request directly to the public forums profile URL to detect username availability.

Why Configure a Separate urlProbe?

Separating the probe endpoint from the public URL provides several technical advantages for specific platforms:

  • Performance optimization – API endpoints like GraphQL return minimal JSON payloads rather than full HTML pages, reducing bandwidth and response time during batch investigations.
  • Rate limit avoidance – Lightweight probe endpoints typically have less aggressive anti-scraping defenses than public profile pages.
  • Reliable error detection – Some sites return HTTP redirects for missing usernames; dedicated probe endpoints often return explicit status codes or JSON error objects that simplify "username not found" detection.
  • Method flexibility – Certain probes require POST requests with JSON payloads, which would be inappropriate or impossible when requesting a standard profile page.

Accessing URL Configuration Programmatically

You can inspect these fields directly through Sherlock's SitesInformation class:

from sherlock_project.sites import SitesInformation

sites = SitesInformation()
site = sites.sites["Anilist"]

print("Public URL :", site.information["url"])
print("Probe URL  :", site.information.get("urlProbe"))

This returns:

Public URL : https://anilist.co/user/{}/
Probe URL  : https://graphql.anilist.co/

The get() method safely returns None for sites without a configured probe URL, matching the internal logic that triggers the fallback behavior.

Summary

  • url always represents the public profile link shown to users in Sherlock's output report.
  • urlProbe optionally specifies an internal endpoint for checking username existence, often an API or lightweight probe page.
  • Sherlock falls back to url when urlProbe is undefined, maintaining compatibility across all site configurations.
  • The probing logic resides in sherlock_project/sherlock.py, while site definitions are stored in sherlock_project/resources/data.json.
  • Using separate probe URLs improves detection reliability and performance for platforms with available API endpoints.

Frequently Asked Questions

What happens if urlProbe is not defined in a site configuration?

If urlProbe is omitted, Sherlock uses the url value for both the HTTP existence check and the final report display. This fallback behavior is handled automatically in sherlock_project/sherlock.py when the get("urlProbe") call returns None.

Can urlProbe use a different HTTP method than url?

Yes, while the url typically implies a GET request to a profile page, the urlProbe can be configured to work with different request methods. Sherlock's site configuration supports additional parameters that specify HTTP methods, allowing urlProbe endpoints to accept POST requests with JSON payloads while the public url remains a standard web page.

Why does Sherlock show the url instead of urlProbe in results?

Sherlock displays the url value (stored as url_user in results) because this represents the actual location where a human can view the discovered profile. The urlProbe often points to API endpoints or internal services that return raw data rather than renderable web pages, making them unsuitable for direct user access.

Where are these URL configurations stored?

Both url and urlProbe definitions reside in sherlock_project/resources/data.json, which serves as the canonical site list. The sherlock_project/sites.py module loads and exposes this data through the SitesInformation class, making the configuration accessible to the main probing logic in sherlock_project/sherlock.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →