Common Components Shared Between Username and Email Scanning Engines in kaifcodec/user-scanner

Both the username and email scanning pipelines in the user-scanner project reuse a single cohesive core architecture located in user_scanner/core, including the Result data model, ScanConfig, dynamic module loading, async worker patterns, HTTP wrappers, and export utilities.

The kaifcodec/user-scanner repository implements two parallel scanning engines—one for usernames and one for email addresses—yet both are built on a shared foundation of utilities, data structures, and orchestration logic. This design ensures consistency in behavior, performance, and output formats regardless of whether you are investigating a username across social platforms or verifying email availability. The following sections break down each common component with precise references to source files and function signatures.

Shared Data Model: The Result Class

Every scan module—username or email—returns a standardized Result object defined in [user_scanner/core/result.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/result.py).

The Result class encapsulates:

  • Status enum: TAKEN, AVAILABLE, ERROR, SKIPPED
  • Target value: The username or email being scanned
  • Site metadata: Site name, category, optional URL
  • Extended data: Media links, extra metadata for rich output

Helper methods debug(), to_json(), to_csv(), and show() provide uniform serialization and console formatting. Both orchestrator.py and email_orchestrator.py invoke these methods when aggregating and displaying final results.

Global Configuration via ScanConfig

Runtime flags are centralized in the immutable dataclass ScanConfig, located in [user_scanner/core/helpers.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/helpers.py). This component is common to both username and email scanning engines.

ScanConfig fields include:

  • allow_loud — Whether to execute modules marked as "noisy"
  • no_nsfw — Filter for adult-content categories
  • show_all — Display negative results (unavailable/taken) alongside positive findings
  • verbose — Enable detailed debug output
  • timeout — Per-request timeout in seconds

Both orchestrators receive a ScanConfig instance at initialization and respect these flags when filtering modules, skipping loud sites, or controlling output verbosity.

Dynamic Module Discovery and Loading

The filesystem enumeration logic is implemented once in helpers.py and consumed by both engines:

  • load_categories() — Scans the module directory hierarchy (user_scan/ vs email_scan/) and returns valid category names
  • load_modules(category) — Dynamically imports individual scan modules, filtering retired or NSFW entries based on ScanConfig

Both orchestrator.py and email_orchestrator.py call these helpers to build their execution lists. The directory structure differs (user_scan/ for usernames, email_scan/ for emails), but the loading mechanism is identical.

Concurrency Architecture and Async Workers

Both scanning engines implement the same async worker pattern with tunable concurrency limits:

Engine File MAX_CONCURRENT_REQUESTS
Username orchestrator.py 60
Email email_orchestrator.py 25

The core _async_worker logic is nearly identical across both files:

  1. Resolves site name via get_site_name()
  2. Retrieves the validation function (get_scan_func())
  3. Applies loud-module rules from ScanConfig
  4. Executes the validator (awaited if async, otherwise offloaded to the shared thread pool)
  5. Handles timeouts, exceptions, and result normalization
  6. Updates the Result with common fields (URL, metadata)

Both engines share _shared_executor (a thread-pool executor) and use asyncio.Semaphore to bound concurrent HTTP requests.

HTTP Request Abstraction: make_request

A thin httpx wrapper in orchestrator.py provides consistent request handling:

def make_request(url, **kwargs) -> httpx.Response:
    ...

This wrapper is imported by generic validation helpers—generic_validate() and status_validate()—which are in turn used by individual scan modules in both pipelines. Because the same wrapper serves both engines, request headers, proxy handling, and timeout defaults remain uniform across username and email scans.

The email engine additionally patches httpx.AsyncClient and httpx.Client to inject proxy logic automatically, ensuring parity with the username engine's explicit make_request usage.

Progress Reporting and Rich UI

Both orchestrators employ identical rich console interfaces:

from rich.progress import Progress, SpinnerColumn, BarColumn, TextColumn

Live progress bars update via the on_start_cb callback, which sets the description to the current site name. This creates indistinguishable user experiences whether scanning usernames or emails—same visual structure, same responsiveness metrics.

Site Metadata Utilities

Helper functions in helpers.py abstract common operations for both engines:

  • get_site_name(module_path) — Normalizes module filenames to readable site names
  • find_category(module_path) — Determines which category a module belongs to
  • is_loud(module_path) — Detects modules requiring user confirmation
  • get_scan_func(module) — Extracts the validate_ entry point from each module

These utilities are called identically by orchestrator.py and email_orchestrator.py, eliminating duplication in site-resolution logic.

Proxy and User-Agent Management

Network-level concerns are centralized in helpers.py:

  • ProxyManager — Global proxy rotation and configuration
  • get_random_user_agent() — Rotates user-agent strings to avoid fingerprinting

Both engines leverage these utilities: the username engine through direct calls in make_request, the email engine through client patching that applies the same proxy manager transparently.

Export and Formatting Layer

Final output generation is engine-agnostic. The modules [user_scanner/core/formatter.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/formatter.py) and [user_scanner/core/pdf_generator.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/pdf_generator.py) operate on List[Result] without discriminating between username or email origins.

Supported export formats:

  • JSON — Result.to_json() per object, aggregated by formatter.py
  • CSV — Result.to_csv() for tabular reporting
  • PDF — Visual report generation via pdf_generator.py

Code Examples

Running a Username Scan with Shared Components

from user_scanner.core.orchestrator import run_user_full
from user_scanner.core.helpers import ScanConfig

config = ScanConfig(allow_loud=False, show_all=True, verbose=True)
results = run_user_full("targetuser", config)

# Programmatic access to standardized results

for r in results:
    print(r.to_json())

Running an Email Scan with Identical Infrastructure

from user_scanner.core.email_orchestrator import run_email_full_batch
from user_scanner.core.helpers import ScanConfig

config = ScanConfig(allow_loud=False, show_all=False, verbose=False)
email_results = run_email_full_batch("check@example.com", config)

# Export using shared formatting utilities

csv_output = "\n".join(r.to_csv() for r in email_results)
with open("email_report.csv", "w", encoding="utf-8") as f:
    f.write(csv_output)

Using Generic Validation Helpers Directly

from user_scanner.core.orchestrator import status_validate, generic_validate
from user_scanner.core.result import Result

# Status-code based validation (common pattern)

result = status_validate(
    "https://github.com/nonexistentuser123",
    available=404,
    taken=200,
    show_url=True,
)

# Custom parser injection

def parse_twitter(resp):
    return Result.taken() if "data-testid" in resp.text else Result.available()

custom = generic_validate(
    "https://twitter.com/somehandle",
    parse_twitter,
    show_url=True,
)

Key Source Files

File Purpose Engine Usage
[result.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/result.py) Unified result data model Both
[helpers.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/helpers.py) ScanConfig, loading utilities, proxy management Both
[orchestrator.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/orchestrator.py) Username engine with MAX_CONCURRENT_REQUESTS=60 Username only
[email_orchestrator.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/email_orchestrator.py) Email engine with MAX_CONCURRENT_REQUESTS=25 Email only
[formatter.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/formatter.py) JSON/CSV serialization Both
[pdf_generator.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/pdf_generator.py) PDF report generation Both

Summary

  • Result provides a single data model for all scan outputs, with consistent serialization across username and email pipelines
  • ScanConfig centralizes runtime behavior flags consumed by both orchestrators
  • load_categories() and load_modules() enable dynamic discovery with identical logic for both engine types
  • _async worker pattern with shared thread pool and semaphore controls concurrency, differing only inRequest limits (60 vs 25)
  • make_request() wrapper ensures uniform HTTP behavior, headers, and proxy handling
  • Rich progress UI delivers identical user experience regardless of scan target type
  • Export utilities (formatter.py, pdf_generator.py) operate engine-agnostically on Result collections

The username and email scanning engines in kaifcodec/user-scanner differ primarily in target source and concurrency tuning. Their shared core architecture—encompassing data models, configuration, loading, networking, and output—enables maintainable, consistent expansion of scan capabilities.

Frequently Asked Questions

What is the main difference between the username and email orchestrators?

The primary distinction is the concurrency limit and module directory target. orchestrator.py (usernames) permits 60 concurrent requests and loads from user_scan/, while email_orchestrator.py limits concurrency to 25 and loads from email_scan/. The underlying worker logic, result handling, and configuration systems are otherwise identical.

Can I use the same validation function for both username and email modules?

Yes, if the target site's API accepts both formats. Generic helpers like status_validate() and generic_validate() in orchestrator.py are imported and used by modules in both pipelines. The validation logic itself is decoupled from target type; only the URL construction and parsing logic within each module differs.

How do I adjust global timeout and proxy settings for both engines?

Instantiate ScanConfig from user_scanner/core/helpers.py with your desired timeout value, then pass it to either run_user_full() or run_email_full_batch(). Proxy configuration is managed globally through ProxyManager in the same file; both engines respect these settings automatically.

Why do the two engines have different MAX_CONCURRENT_REQUESTS values?

Email validation endpoints typically implement stricter rate limiting and abuse detection than username lookup endpoints. The lower limit of 25 in email_orchestrator.py reduces the risk of IP blocking or account flagging, while the username engine's higher limit of 60 prioritizes throughput against generally more permissive APIs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →