How Sherlock's Parallel Async Request System Works with FuturesSession
Sherlock leverages a thread-pooled FuturesSession subclass to dispatch non-blocking HTTP requests across hundreds of social networks simultaneously, then aggregates results via Future objects to complete username checks in seconds.
The sherlock-project/sherlock tool achieves rapid username enumeration by implementing a sophisticated Sherlock parallel async request system built on the requests-futures library. This architecture allows the tool to query hundreds of sites concurrently without rewriting core logic for asyncio, demonstrating an efficient pattern for high-throughput HTTP scanning in Python.
Core Architecture Components
SherlockFuturesSession Subclass
The custom session wrapper lives in [sherlock_project/sherlock.py](https://github.com/sherlock-project/sherlock/blob/master/sherlock_project/sherlock.py#L48-L110) (lines 48–110). This class extends FuturesSession.request() to inject a response-time hook into every HTTP call. By overriding the request method, Sherlock records precise timing data via monotonic() without altering the underlying request logic or breaking compatibility with standard requests patterns.
The sherlock() Orchestrator Function
The main dispatch logic resides in [sherlock_project/sherlock.py](https://github.com/sherlock-project/sherlock/blob/master/sherlock_project/sherlock.py#L220-L340) (lines 220–340). This function builds a bounded thread pool (capped at 20 workers or the number of target sites, whichever is smaller), creates a Future for each HTTP request, stores the future in the site's dictionary under net_info["request_future"], and later resolves each result through get_response().
Step-by-Step Execution Flow
The Sherlock parallel async request system operates through a five-phase pipeline:
-
Initialize the underlying session – Creates a standard
requests.session()object to handle connection pooling and cookie persistence. -
Wrap with SherlockFuturesSession – Instantiates the custom session with
max_workerslimited to 20, spawning a background thread pool that executes HTTP calls asynchronously.session = SherlockFuturesSession(max_workers=max_workers, session=underlying_session) -
Dispatch futures for each site – Iterates over
site_data, builds probe URLs viainterpolate_string(), selects the HTTP method (GET,HEAD, etc.), and calls the corresponding session method. Each call returns immediately with aFutureobject stored in the site dictionary:future = request( url=url_probe, headers=headers, allow_redirects=allow_redirects, timeout=timeout, json=request_payload, proxies=proxies if proxy else None, ) net_info["request_future"] = future -
Parallel execution – While the main thread continues building futures for remaining sites, the thread pool executes previously dispatched HTTP requests concurrently.
-
Result collection – A second loop traverses
site_data, retrieves each stored future, and blocks onfuture.result()viaget_response(). The hook installed bySherlockFuturesSessionattaches anelapsedattribute to each response, enabling precise latency reporting.
Technical Advantages of FuturesSession
Sherlock selects FuturesSession over pure asyncio for three specific architectural benefits:
- Thread-pooled concurrency – Eliminates the need for
async/awaitsyntax while achieving true parallelism through background threads. - Back-pressure control – The
max_workerscap prevents spawning hundreds of threads when scanning large site lists, protecting both local resources and target servers. - API transparency – The subclass maintains 100% compatibility with the standard
requestsAPI, preserving existing code for headers, proxies, JSON payloads, and redirect handling.
Implementation Example: Building a Similar System
The following pattern mirrors Sherlock's approach for custom tools requiring parallel HTTP measurement:
from requests_futures.sessions import FuturesSession
import requests
from time import monotonic
class TimedFuturesSession(FuturesSession):
"""Inject a response-time hook like Sherlock does."""
def request(self, method, url, hooks=None, *args, **kwargs):
if hooks is None:
hooks = {}
start = monotonic()
def _timer(resp, *_, **__):
resp.elapsed = monotonic() - start
# Ensure our timer runs first
if "response" in hooks:
if isinstance(hooks["response"], list):
hooks["response"].insert(0, _timer)
else:
hooks["response"] = [_timer, hooks["response"]]
else:
hooks["response"] = [_timer]
return super().request(method, url, hooks=hooks, *args, **kwargs)
# ---- Parallel request block ----
underlying = requests.session()
session = TimedFuturesSession(max_workers=10, session=underlying)
urls = [
"https://api.github.com/users/octocat",
"https://gitlab.com/users/octocat",
# … more URLs …
]
futures = {url: session.get(url, timeout=15) for url in urls}
# ---- Collect results ----
for url, future in futures.items():
resp = future.result()
print(f"{url} → {resp.status_code} (took {resp.elapsed:.2f}s)")
This implementation follows Sherlock's exact structure: create a timed FuturesSession, fire off all requests to obtain future objects, then resolve results later while accessing timing metadata.
Key Source Files and Their Roles
| File | Role in the Async System |
|---|---|
sherlock_project/sherlock.py |
Implements SherlockFuturesSession, builds the session, dispatches futures, and aggregates results. |
sherlock_project/result.py |
Defines QueryStatus and QueryResult objects that store final outcomes including elapsed time captured by the session. |
sherlock_project/sites.py |
Supplies per-site configuration (URLs, headers, request methods, error types) that drives the request loop. |
requirements.txt |
Lists requests-futures as a dependency, ensuring the async wrapper is available for installation. |
Summary
- Sherlock uses a lightweight subclass of
FuturesSessionto inject timing hooks without breaking the standardrequestsAPI. - A fixed thread pool with a maximum of 20 workers prevents resource exhaustion while maximizing throughput.
- The fire-and-collect pattern stores
Futureobjects in site dictionaries during dispatch, then resolves them in a second pass. - Response timing occurs automatically through hooks installed in
SherlockFuturesSession.request().
Frequently Asked Questions
What is FuturesSession and how does it enable parallel requests?
FuturesSession extends the standard requests library to return Future objects immediately while executing HTTP calls in background threads from a configurable pool. This allows the main thread to dispatch hundreds of requests sequentially while the thread pool handles network I/O concurrently, achieving parallelism without requiring async/await syntax or event loops.
Why does Sherlock use threading instead of Python's asyncio?
According to the sherlock-project/sherlock source code, the tool uses thread-pooled concurrency via requests-futures because it maintains complete compatibility with existing synchronous requests code. This avoids the refactoring complexity of converting to asyncio while still achieving sub-second response times across hundreds of sites through parallel thread execution.
How does Sherlock limit concurrent connections to avoid being blocked?
The system enforces a hard cap of 20 workers (or fewer if fewer sites are queried) when instantiating SherlockFuturesSession in sherlock_project/sherlock.py. This bounded thread pool naturally throttles connection attempts, preventing both local resource exhaustion and aggressive traffic patterns that might trigger rate-limiting mechanisms on target servers.
Where does Sherlock measure response time for each request?
The timing logic resides in the SherlockFuturesSession.request() method defined at lines 48–110 of sherlock_project/sherlock.py. This method records a start timestamp via monotonic(), injects a response hook that calculates elapsed time upon completion, and attaches the result to the response object as an elapsed attribute for later reporting.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →