How HTTPX and curl_cffi Sessions Are Cached in the User-Scanner Repository

Both HTTPX and curl_cffi sessions are cached globally using thread-safe singleton patterns—HTTPX via a lazily-initialized httpx.Client in helpers.py, and curl_cffi via a dictionary keyed by (impersonate, proxy) tuples in impersonate.py.

The User-Scanner project performs high-volume username and email validation across dozens of platforms. To avoid the latency and resource exhaustion of establishing new TCP connections for every request, the codebase implements process-wide session caching for both its standard HTTP client (HTTPX) and its TLS-impersonating client (curl_cffi). This design is critical when running with the CLI --concurrency flag, where multiple threads execute validations simultaneously.


HTTPX Session Caching: The Global Client Pattern

In user_scanner/core/helpers.py, the repository maintains a single httpx.Client instance that lives for the entire process lifetime. This provides connection pool reuse transparently to all synchronous validators.

Implementation Details

The global client is stored in _httpx_client and protected by _client_lock to ensure thread-safe initialization:

import threading
import httpx

_httpx_client: httpx.Client | None = None
_client_lock = threading.Lock()


def get_httpx_client() -> httpx.Client:
    """Return a lazily-initialised global `httpx.Client`."""
    global _httpx_client
    if _httpx_client is None:
        with _client_lock:
            if _httpx_client is None:
                timeout = get_timeout()
                proxy = get_proxy()
                _httpx_client = httpx.Client(timeout=timeout, proxies=proxy)
    return _httpx_client

The double-checked locking pattern ensures that only one thread creates the client even when multiple threads race during startup.

Using the Cached Client

Validators call generic_validate, which retrieves the global client automatically:

def generic_validate(url: str, method: str = "GET", **kwargs: Any) -> httpx.Response:
    client = get_httpx_client()
    response = client.request(method, url, **kwargs)
    response.raise_for_status()
    return response

Real Usage: Instagram Validator

The Instagram username checker in user_scanner/user_scan/social/instagram.py demonstrates this pattern:

from ..core.helpers import generic_validate

_PROFILE_URL = "https://www.instagram.com/{user}/?__a=1"

def validate_instagram(user: str):
    url = _PROFILE_URL.format(user=user)
    try:
        response = generic_validate(url)  # ← reuses global httpx.Client

    except httpx.HTTPStatusError as e:
        if e.response.status_code == 404:
            return "available"
        raise
    data = json.loads(response.text)
    return "taken" if data.get("graphql", {}).get("user", {}) else "available"

curl_cffi Session Caching: Per-Impersonation-Profile Sessions

For platforms that block generic HTTP clients (e.g., Cloudflare-protected sites like Reddit), User-Scanner uses curl_cffi to mimic browser TLS fingerprints. Since creating these sessions involves a warm-up request to obtain clearance cookies, the repository caches them aggressively.

The Session Cache Design

In user_scanner/core/impersonate.py, sessions are stored in _sessions with a composite key:

import threading
import curl_cffi

_sessions: dict[tuple[str, str | None], curl_cffi.requests.Session] = {}
_lock = threading.Lock()

The cache key combines:

  • impersonate: The TLS fingerprint string (e.g., "chrome", "firefox")
  • proxy: The proxy URL string or None

This allows different proxy configurations to maintain separate session pools.

Session Retrieval and Warm-up

The _get_warm_session function implements create-on-demand with warm-up:

def _get_warm_session(
    impersonate: str,
    proxy: str | None,
    warmup_url: str,
) -> curl_cffi.requests.Session:
    key = (impersonate, proxy)
    with _lock:
        session = _sessions.get(key)
        if session is None:
            session = curl_cffi.requests.Session(impersonate=impersonate, proxy=proxy)
            _sessions[key] = session
        # Warm-up: fetch clearance page once per session

        try:
            session.get(warmup_url, timeout=_timeout())
        except Exception:
            pass  # Allow caller to handle retry logic

    return session

The warm-up request is crucial—Cloudflare and similar services return challenge pages initially. By performing this once and caching the resulting cookie state, subsequent requests bypass the challenge.

Public API: impersonate_validate

Validators use the high-level wrapper:

def impersonate_validate(
    impersonate: str,
    url: str,
    method: str = "GET",
    **kwargs: Any,
) -> requests.Response:
    from .helpers import get_proxy
    proxy = get_proxy()
    session = _get_warm_session(impersonate, proxy, warmup_url=url)
    return session.request(method, url, timeout=_timeout(), **kwargs)

Real Usage: Reddit Validator

Reddit's bot wall requires impersonation, as shown in user_scanner/user_scan/social/reddit.py:

from ..core.impersonate import impersonate_validate

_BASE_URL = "https://www.reddit.com/user/{user}/about.json"
_WARMUP_URL = "https://www.reddit.com/"

def validate_reddit(user: str):
    url = _BASE_URL.format(user=user)
    response = impersonate_validate(
        impersonate="chrome",
        url=url,
        method="GET",
        warmup_url=_WARMUP_URL,
    )
    data = json.loads(response.text)
    return "taken" if data.get("is_employee") else "available"

Note that warmup_url is passed as a dedicated parameter—this separates the warm-up target from the actual validation URL, allowing cheaper endpoints to prime the session.


Async Support: Thread Pool Delegation

The email scanning modules are async, but curl_cffi is synchronous. The repository bridges this gap using a thread pool executor without breaking session caching:

import asyncio
from concurrent.futures import ThreadPoolExecutor

_executor = ThreadPoolExecutor(max_workers=4)

def _run_in_thread(fn: Callable[..., Any], *args: Any, **kwargs: Any) -> Any:
    loop = asyncio.get_event_loop()
    return loop.run_in_executor(_executor, lambda: fn(*args, **kwargs))


async def impersonate_async(
    impersonate: str,
    url: str,
    method: str = "GET",
    **kwargs: Any,
) -> requests.Response:
    """Async version of impersonate_validate."""
    return await _run_in_thread(impersonate_validate, impersonate, url, method, **kwargs)

The same cached session is accessed from worker threads—threading.Lock in _get_warm_session ensures safety.


Comparison: HTTPX vs. curl_cffi Caching

Aspect HTTPX (helpers.py) curl_cffi (impersonate.py)
Cache scope Single global client Multiple sessions per (impersonate, proxy) key
Key dimension None (one client) Tuple of fingerprint + proxy
Warm-up needed No Yes—primes cookies/TLS state
Thread safety threading.Lock on initialization threading.Lock on all access
Use case Standard sites (Instagram, etc.) Bot-protected sites (Reddit, Cloudflare)
Async support Native AsyncClient not cached; sync client used via thread pool Thread pool wrapper around sync sessions

Both mechanisms prioritize connection reuse to reduce latency under concurrent load.


Summary

  • HTTPX sessions are cached via a singleton httpx.Client in user_scaner/core/helpers.py, initialized once per process with thread-safe double-checked locking
  • curl_cffi sessions are cached in a dictionary keyed by (impersonate, proxy) tuples in user_scaner/core/impersonate.py, with mandatory warm-up requests to seed clearance state
  • Both caches use threading.Lock for safe access under the --concurrency CLI flag
  • Async modules access curl_cffi through a thread pool without creating duplicate sessions

Frequently Asked Questions

Does User-Scanner create a new HTTPX client for every request?

No. The get_httpx_client() function in user_scaner/core/helpers.py creates the client once and returns the same instance for all subsequent calls. This is implemented with a thread-safe singleton pattern using threading.Lock to handle concurrent initialization safely.

Why does curl_cffi need a warm-up request?

The warm-up request in _get_warm_session obtains clearance cookies from services like Cloudflare that present challenge pages to new clients. By caching the session with these cookies, subsequent requests bypass the challenge. Without this, every new session would trigger expensive and potentially blocking verification flows.

Can I use different proxies with the same impersonation profile?

Yes. The cache key in _sessions is (impersonate, proxy), so different proxy configurations automatically create separate session pools. This prevents proxy-specific state (like IP-based rate limits) from leaking between different proxy endpoints.

Is the async interface for curl_cffi truly non-blocking?

The impersonate_async function uses asyncio.get_event_loop().run_in_executor with a ThreadPoolExecutor to run synchronous curl_cffi calls in background threads. While this doesn't provide true async I/O, it prevents the event loop from blocking while preserving the cached session benefits across async and sync code paths.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →