How User-Scanner Performs Username Scanning: A Deep Dive Into the Plugin Architecture

User-scanner discovers username availability across social platforms by dynamically loading platform-specific modules and invoking a unified validation pipeline that handles bot detection, impersonation, and confidence scoring.

The kaifcodec/user-scanner repository implements a modular, extensible system for checking whether usernames exist on hundreds of websites. This article breaks down the core mechanics of how the scanner operates, from module discovery to final result aggregation.

Core Architecture: The Orchestration Layer

At the heart of user-scanner is the orchestrate_user_scan function in user_scanner/core/orchestrator.py. This component coordinates the entire workflow:

  1. Accepts target usernames from CLI or programmatic input
  2. Invokes load_modules from core/engine.py to discover platform modules
  3. Executes validation functions for each discovered platform
  4. Aggregates results and applies confidence scoring

The orchestrator implements a plugin architecture that requires zero manual registration—new platforms are automatically detected at runtime.

Module Discovery and Loading

The engine walks the directory tree under user_scanner/user_scan/** and imports every module matching the <site>.py naming convention. For each module found, it constructs a callable validate_<site> function (e.g., validate_github, validate_reddit).

This dynamic loading happens in core/engine.py, which serves as the central dispatcher between the orchestrator and individual platform implementations.

Validation Pipeline: Generic vs. Impersonate Transport

Each platform module returns standardized results through one of two transport mechanisms defined in core/engine.py:

  • generic_validate — Uses httpx for standard HTTP requests on sites without bot protection
  • impersonate_validate — Routes through curl_cffi when a module's transport attribute is set to impersonate

The impersonation layer in core/impersonate.py automatically handles Cloudflare challenges, JavaScript challenges, and other anti-scraping shields by mimicking real browser fingerprints.

Result Object Model

Validation functions return Result objects from core/result.py with three possible states:

Result.available()           # Username does not exist

Result.taken(extra=..., media=...)  # Username exists with metadata

Result.error(message)        # Check failed due to network or parsing error

Confidence Scoring and False Positive Prevention

After each check, core/confidence.py computes a confidence score based on:

  • Explicit "found" and "not-found" markers in page content
  • HTTP response status codes
  • Response timing characteristics

The dual-marker approach—checking for both positive and negative indicators—significantly reduces false positives compared to single-indicator checks.

Practical Usage Examples

Command-Line Scanning


# Scan single username across all platforms

python -m user_scanner scan username alice

# Batch scan with JSON output and custom timeout

python -m user_scanner scan username -i usernames.txt -o results.json -t 30

# Enable loud prompts for interactive confirmation

python -m user_scanner scan username bob --allow-loud -C 10

Programmatic Integration

from user_scanner.core.orchestrator import orchestrate_user_scan
from user_scanner.core.result import Result

# Execute scan

results = orchestrate_user_scan(user="bob")

# Process outcomes

for site, outcome in results.items():
    match outcome.status:
        case "available":
            print(f"{site}: Available")
        case "taken":
            print(f"{site}: Taken — {outcome.extra.get('profile_url', 'N/A')}")
        case "error":
            print(f"{site}: Error — {outcome.message}")

Adding Custom Platform Modules

Create a new file under user_scanner/user_scan/<category>/<site>.py:

import httpx
from user_scanner.core.result import Result

TRANSPORT = "generic"  # or "impersonate" for bot-protected sites

def validate_example(user: str) -> Result:
    url = f"https://example.com/u/{user}"
    
    try:
        resp = httpx.get(url, timeout=15, follow_redirects=True)
    except httpx.RequestError as exc:
        return Result.error(f"Request failed: {exc}")
    
    # Explicit dual-marker checking for confidence

    not_found_indicators = ["User not found", "404 - Page not found"]
    found_indicators = ["Profile", "Followers", "(@{user})".format(user=user)]
    
    if any(marker in resp.text for marker in not_found_indicators):
        return Result.available()
    
    if resp.status_code == 200 and any(marker in resp.text for marker in found_indicators):
        return Result.taken(
            extra={"profile_url": str(resp.url)},
            media={"avatar": extract_avatar(resp.text)}
        )
    
    return Result.error("Ambiguous response")

The orchestrator automatically discovers and includes this module on the next run—no registry edits required.

Key Implementation Files

File Responsibility
user_scanner/core/orchestrator.py Workflow coordination and result aggregation
user_scanner/core/engine.py Module loading and transport selection
user_scanner/core/helpers.py HTTP utilities and timeout handling
user_scanner/core/result.py Standardized return value encapsulation
user_scanner/core/impersonate.py curl_cffi integration for bot evasion
user_scanner/core/confidence.py Score computation for result reliability
user_scanner/core/formatter.py Output formatting (table, JSON, CSV, PDF)
user_scanner/cli/__init__.py CLI entry point and argument parsing

Summary

  • Plugin architecture: Add platforms by placing modules in user_scan/ without registry changes
  • Dual transport system: httpx for standard sites, curl_cffi impersonation for protected ones
  • Confidence scoring: Dual-marker validation reduces false positives in core/confidence.py
  • Unified result model: Result.available(), Result.taken(), and Result.error() standardize outcomes
  • CLI and library APIs: Use as command-line tool or import orchestrate_user_scan directly

Frequently Asked Questions

How does user-scanner handle websites that block automated requests?

When a platform module specifies TRANSPORT = "impersonate", the engine routes requests through core/impersonate.py using curl_cffi. This library mimics real browser TLS fingerprints, headers, and behavior to bypass Cloudflare and similar protections.

What makes user-scanner's detection more reliable than simple status code checks?

The validation pipeline in core/helpers.py implements dual-marker verification—checking for explicit "found" indicators and "not-found" indicators in page content. This approach, combined with confidence scoring in core/confidence.py, eliminates false positives from ambiguous responses and redirects.

Can I restrict scanning to specific platforms only?

Yes. The orchestrator accepts an optional platform filter list. When provided, load_modules in core/engine.py skips modules not matching the specified platforms. Use the CLI --platforms flag or pass platforms=["github", "twitter"] to orchestrate_user_scan().

How do I increase scanning speed for large username lists?

Adjust the -C (concurrency) and -t (timeout) CLI flags. The orchestrator uses these values to control parallel execution through asyncio.gather() in core/engine.py, with timeout handling preventing slow platforms from blocking the entire batch.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →