Understanding the Architecture of user-scanner: A Modular OSINT Framework

The user-scanner architecture follows a five-layer modular design separating CLI argument parsing, asynchronous orchestration, HTTP engine utilities, pluggable scan modules, and export formatters, enabling high-concurrency OSINT investigations via Python's asyncio.

The kaifcodec/user-scanner project implements a clean, layered architecture designed for high-performance open-source intelligence (OSINT) gathering. This Python-based framework separates concerns across five distinct layers, allowing security researchers to execute concurrent username and email scans while maintaining extensibility for new platforms. Understanding the user-scanner architecture reveals how it balances modularity with asynchronous performance through a sophisticated orchestration system.

CLI and Configuration Layer

The entry point resides in user_scanner/__main__.py, which handles argument parsing, banner display via user_scanner/cli/banner.py, and configuration initialization. When invoked, the CLI constructs a ScanConfig object containing flags such as allow_loud, show_all, and proxy settings. This configuration object propagates through all downstream layers, ensuring consistent behavior across the scan lifecycle.

The CLI also manages the self-update mechanism through user_scanner/utils/update.py and loads user preferences from user_scanner/config.json. After parsing arguments, the entry point delegates execution to the appropriate orchestrator function based on the scan type (username, email, or cross-scan pivoting).

Asynchronous Orchestration Layer

The core concurrency management lives in user_scanner/core/orchestrator.py (for username scans), user_scanner/core/email_orchestrator.py (for email validation), and user_scanner/core/cross_scan.py (for pivoting operations). These orchestrators coordinate the execution of hundreds of site-specific validators while maintaining system stability.

Key implementation details include:

  • Bounded concurrency: A semaphore limits simultaneous connections to 60 by default, preventing resource exhaustion.
  • Task management: The orchestrator spawns asyncio tasks via _async_worker for each discovered module, executing either coroutines directly or wrapping synchronous functions in a thread-pool executor (_shared_executor).
  • Progress reporting: Live scan status displays through rich.progress.Progress bars, providing real-time visibility into completion rates.
  • Result aggregation: As validators complete, the orchestrator collects Result objects into a unified list for downstream formatting.

Engine and Helper Utilities

The user_scanner/core/helpers.py module provides the foundational infrastructure for module discovery, proxy rotation, and HTTP session management. It exposes critical functions including load_modules for dynamic site discovery, get_scan_func for validator retrieval, and make_request for standardized HTTP operations.

Supporting components include:

  • user_scanner/core/engine.py: Wraps individual validator functions with check, handling timeouts, exceptions, and HTTP retries while returning standardized Result objects.
  • user_scanner/core/result.py: Defines the Result and Status dataclasses that encapsulate site name, availability status, metadata (extra, media), and console formatting logic.
  • user_scanner/core/impersonate.py: Manages realistic User-Agent rotation and bot-wall evasion techniques.
  • Proxy management: A thread-safe ProxyManager class rotates proxies across requests, configurable via CLI flags.

Pluggable Scan Module System

Scan modules follow a convention-based discovery pattern under user_scanner/user_scan/<category>/<site>.py for username validation and user_scanner/email_scan/<category>/<service>.py for email checks. Each module implements a validate_<site> function that accepts a username or email string and returns a status determination.

Modules typically utilize the status_validate helper from user_scanner/core/helpers.py to map HTTP status codes (e.g., 200 for available, 404 for taken) to standardized results. The orchestrator discovers these modules automatically at runtime through load_modules, requiring no registry updates when adding new platforms.

For example, implementing support for a new platform requires only creating a single file with a validator function:


# File: user_scanner/user_scan/social/example.py

from user_scanner.core.helpers import make_request, status_validate

def validate_example(username: str):
    # Example.com returns 200 when the profile exists, 404 otherwise

    return status_validate(
        f"https://example.com/{username}",
        available=404,
        taken=200,
        show_url=True,           # displays URL when --verbose is used

    )

Output Formatting and Export

After orchestration completes, user_scanner/core/formatter.py and user_scanner/core/pdf_generator.py convert the list of Result objects into persistent formats. The formatter supports JSON, CSV, and PDF outputs based on the --format CLI flag.

PDF generation specifically accepts metadata including the target identifier, scan type, total modules executed, and version information from user_scanner/core/version.py. This ensures forensic-grade reporting with timestamps and scan parameters embedded in the output.

The following example demonstrates programmatic PDF generation after a scan:

from user_scanner.core.formatter import into_pdf
from user_scanner.core.version import load_local_version

pdf_bytes = into_pdf(
    results,
    target="alice",
    scan_type="Username",
    total_modules=len(results),
    include_media=False,
    version=load_local_version()[0],
)

with open("alice_report.pdf", "wb") as f:
    f.write(pdf_bytes)

Data Flow: Executing a Username Scan

The architecture processes a typical username query through the following pipeline:

  1. Initialization: The CLI parses --username alice and instantiates ScanConfig, then invokes run_user_full("alice", config) from the orchestrator.
  2. Module Discovery: The orchestrator calls load_categories and load_modules to enumerate all available validators across categories (social, e-commerce, etc.).
  3. Concurrent Execution: For each module, an asyncio task executes _async_worker, which fetches the validator via get_scan_func and executes it through the engine's check wrapper.
  4. Result Processing: Validators return Result objects containing status, URLs, and metadata. Errors and timeouts are caught and converted to Result.error instances.
  5. Presentation: Completed results stream to the console via Result.get_console_output, respecting verbosity flags. Finally, the CLI may invoke formatter.into_json, into_csv, or into_pdf to persist findings.

You can also trigger this flow programmatically without the CLI:

from user_scanner.core.orchestrator import run_user_full
from user_scanner.core.helpers import ScanConfig

cfg = ScanConfig(show_all=True, verbose=True)
results = run_user_full("alice", cfg)

for r in results:
    print(r.get_console_output(cfg))

Summary

  • Modular extensibility: New platforms require only a single Python file implementing validate_<site>; the orchestrator handles discovery automatically.
  • Async-first design: All network I/O runs under asyncio with configurable semaphore limits, maximizing throughput while controlling memory usage.
  • Unified data model: The Result class standardizes output across all 200+ site modules, enabling consistent JSON, CSV, and PDF exports.
  • Safety controls: "Loud" modules that may trigger password-reset emails are disabled by default and require explicit --allow-loud activation.
  • Dual-interface support: Functions both as a command-line tool via __main__.py and as an importable Python library for custom automation.

Frequently Asked Questions

How does user-scanner handle concurrent requests?

The framework utilizes Python's asyncio with a bounded semaphore defaulting to 60 concurrent connections, as implemented in user_scanner/core/orchestrator.py. The orchestrator spawns individual tasks for each site module, executing asynchronous validators directly and wrapping synchronous ones in a thread-pool executor to prevent blocking the event loop.

What is the scan module discovery mechanism?

User-scanner employs dynamic module loading via load_modules in user_scanner/core/helpers.py. The orchestrator scans the user_scanner/user_scan/ and user_scanner/email_scan/ directories at runtime, importing any Python files containing validate_ prefixed functions. This convention-based approach eliminates the need for manual registration when adding new platforms.

How can I extend user-scanner to support a new website?

Create a new Python file under the appropriate category directory (e.g., user_scanner/user_scan/social/newsite.py) and implement a validate_newsite(username) function that returns a Result object. Use the status_validate helper to map HTTP status codes to availability states, or implement custom parsing logic for complex responses. The orchestrator will automatically detect and execute the module on the next scan.

Does user-scanner support proxy rotation and user-agent management?

Yes. The architecture includes a thread-safe ProxyManager class in user_scanner/core/helpers.py that rotates proxies across requests, configurable via the --proxy CLI flag or programmatic ScanConfig. Additionally, user_scanner/core/impersonate.py provides realistic User-Agent rotation and bot-wall evasion capabilities to reduce detection rates during large-scale scans.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →