How HTTPX and curl_cffi Sessions Are Cached in the User-Scanner Repository
Both HTTPX and curl_cffi sessions are cached globally using thread-safe singleton patterns—HTTPX via a lazily-initialized httpx.Client in helpers.py, and curl_cffi via a dictionary keyed by (impersonate, proxy) tuples in impersonate.py.
The User-Scanner project performs high-volume username and email validation across dozens of platforms. To avoid the latency and resource exhaustion of establishing new TCP connections for every request, the codebase implements process-wide session caching for both its standard HTTP client (HTTPX) and its TLS-impersonating client (curl_cffi). This design is critical when running with the CLI --concurrency flag, where multiple threads execute validations simultaneously.
HTTPX Session Caching: The Global Client Pattern
In user_scanner/core/helpers.py, the repository maintains a single httpx.Client instance that lives for the entire process lifetime. This provides connection pool reuse transparently to all synchronous validators.
Implementation Details
The global client is stored in _httpx_client and protected by _client_lock to ensure thread-safe initialization:
import threading
import httpx
_httpx_client: httpx.Client | None = None
_client_lock = threading.Lock()
def get_httpx_client() -> httpx.Client:
"""Return a lazily-initialised global `httpx.Client`."""
global _httpx_client
if _httpx_client is None:
with _client_lock:
if _httpx_client is None:
timeout = get_timeout()
proxy = get_proxy()
_httpx_client = httpx.Client(timeout=timeout, proxies=proxy)
return _httpx_client
The double-checked locking pattern ensures that only one thread creates the client even when multiple threads race during startup.
Using the Cached Client
Validators call generic_validate, which retrieves the global client automatically:
def generic_validate(url: str, method: str = "GET", **kwargs: Any) -> httpx.Response:
client = get_httpx_client()
response = client.request(method, url, **kwargs)
response.raise_for_status()
return response
Real Usage: Instagram Validator
The Instagram username checker in user_scanner/user_scan/social/instagram.py demonstrates this pattern:
from ..core.helpers import generic_validate
_PROFILE_URL = "https://www.instagram.com/{user}/?__a=1"
def validate_instagram(user: str):
url = _PROFILE_URL.format(user=user)
try:
response = generic_validate(url) # ← reuses global httpx.Client
except httpx.HTTPStatusError as e:
if e.response.status_code == 404:
return "available"
raise
data = json.loads(response.text)
return "taken" if data.get("graphql", {}).get("user", {}) else "available"
curl_cffi Session Caching: Per-Impersonation-Profile Sessions
For platforms that block generic HTTP clients (e.g., Cloudflare-protected sites like Reddit), User-Scanner uses curl_cffi to mimic browser TLS fingerprints. Since creating these sessions involves a warm-up request to obtain clearance cookies, the repository caches them aggressively.
The Session Cache Design
In user_scanner/core/impersonate.py, sessions are stored in _sessions with a composite key:
import threading
import curl_cffi
_sessions: dict[tuple[str, str | None], curl_cffi.requests.Session] = {}
_lock = threading.Lock()
The cache key combines:
- impersonate: The TLS fingerprint string (e.g.,
"chrome","firefox") - proxy: The proxy URL string or
None
This allows different proxy configurations to maintain separate session pools.
Session Retrieval and Warm-up
The _get_warm_session function implements create-on-demand with warm-up:
def _get_warm_session(
impersonate: str,
proxy: str | None,
warmup_url: str,
) -> curl_cffi.requests.Session:
key = (impersonate, proxy)
with _lock:
session = _sessions.get(key)
if session is None:
session = curl_cffi.requests.Session(impersonate=impersonate, proxy=proxy)
_sessions[key] = session
# Warm-up: fetch clearance page once per session
try:
session.get(warmup_url, timeout=_timeout())
except Exception:
pass # Allow caller to handle retry logic
return session
The warm-up request is crucial—Cloudflare and similar services return challenge pages initially. By performing this once and caching the resulting cookie state, subsequent requests bypass the challenge.
Public API: impersonate_validate
Validators use the high-level wrapper:
def impersonate_validate(
impersonate: str,
url: str,
method: str = "GET",
**kwargs: Any,
) -> requests.Response:
from .helpers import get_proxy
proxy = get_proxy()
session = _get_warm_session(impersonate, proxy, warmup_url=url)
return session.request(method, url, timeout=_timeout(), **kwargs)
Real Usage: Reddit Validator
Reddit's bot wall requires impersonation, as shown in user_scanner/user_scan/social/reddit.py:
from ..core.impersonate import impersonate_validate
_BASE_URL = "https://www.reddit.com/user/{user}/about.json"
_WARMUP_URL = "https://www.reddit.com/"
def validate_reddit(user: str):
url = _BASE_URL.format(user=user)
response = impersonate_validate(
impersonate="chrome",
url=url,
method="GET",
warmup_url=_WARMUP_URL,
)
data = json.loads(response.text)
return "taken" if data.get("is_employee") else "available"
Note that warmup_url is passed as a dedicated parameter—this separates the warm-up target from the actual validation URL, allowing cheaper endpoints to prime the session.
Async Support: Thread Pool Delegation
The email scanning modules are async, but curl_cffi is synchronous. The repository bridges this gap using a thread pool executor without breaking session caching:
import asyncio
from concurrent.futures import ThreadPoolExecutor
_executor = ThreadPoolExecutor(max_workers=4)
def _run_in_thread(fn: Callable[..., Any], *args: Any, **kwargs: Any) -> Any:
loop = asyncio.get_event_loop()
return loop.run_in_executor(_executor, lambda: fn(*args, **kwargs))
async def impersonate_async(
impersonate: str,
url: str,
method: str = "GET",
**kwargs: Any,
) -> requests.Response:
"""Async version of impersonate_validate."""
return await _run_in_thread(impersonate_validate, impersonate, url, method, **kwargs)
The same cached session is accessed from worker threads—threading.Lock in _get_warm_session ensures safety.
Comparison: HTTPX vs. curl_cffi Caching
| Aspect | HTTPX (helpers.py) |
curl_cffi (impersonate.py) |
|---|---|---|
| Cache scope | Single global client | Multiple sessions per (impersonate, proxy) key |
| Key dimension | None (one client) | Tuple of fingerprint + proxy |
| Warm-up needed | No | Yes—primes cookies/TLS state |
| Thread safety | threading.Lock on initialization |
threading.Lock on all access |
| Use case | Standard sites (Instagram, etc.) | Bot-protected sites (Reddit, Cloudflare) |
| Async support | Native AsyncClient not cached; sync client used via thread pool |
Thread pool wrapper around sync sessions |
Both mechanisms prioritize connection reuse to reduce latency under concurrent load.
Summary
- HTTPX sessions are cached via a singleton
httpx.Clientinuser_scaner/core/helpers.py, initialized once per process with thread-safe double-checked locking - curl_cffi sessions are cached in a dictionary keyed by
(impersonate, proxy)tuples inuser_scaner/core/impersonate.py, with mandatory warm-up requests to seed clearance state - Both caches use
threading.Lockfor safe access under the--concurrencyCLI flag - Async modules access curl_cffi through a thread pool without creating duplicate sessions
Frequently Asked Questions
Does User-Scanner create a new HTTPX client for every request?
No. The get_httpx_client() function in user_scaner/core/helpers.py creates the client once and returns the same instance for all subsequent calls. This is implemented with a thread-safe singleton pattern using threading.Lock to handle concurrent initialization safely.
Why does curl_cffi need a warm-up request?
The warm-up request in _get_warm_session obtains clearance cookies from services like Cloudflare that present challenge pages to new clients. By caching the session with these cookies, subsequent requests bypass the challenge. Without this, every new session would trigger expensive and potentially blocking verification flows.
Can I use different proxies with the same impersonation profile?
Yes. The cache key in _sessions is (impersonate, proxy), so different proxy configurations automatically create separate session pools. This prevents proxy-specific state (like IP-based rate limits) from leaking between different proxy endpoints.
Is the async interface for curl_cffi truly non-blocking?
The impersonate_async function uses asyncio.get_event_loop().run_in_executor with a ThreadPoolExecutor to run synchronous curl_cffi calls in background threads. While this doesn't provide true async I/O, it prevents the event loop from blocking while preserving the cached session benefits across async and sync code paths.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →