How to Contribute to the user-scanner Project: A Developer's Guide to Adding OSINT Modules
To contribute to user-scanner, implement a new validator module in user_scanner/user_scan/<category>/ for username lookups or user_scanner/email_scan/<category>/ for email OSINT, using the Result class and helper functions from the core orchestrator.
The user-scanner repository is a modular OSINT suite that discovers email and username profiles across hundreds of web services. Whether you want to add support for a new social platform or improve the core engine, contributions follow a standardized pattern centered on validator modules and explicit result states. This guide walks through the architecture, implementation requirements, and submission workflow based on the current source code structure.
Architecture Overview
Understanding the repository layout is essential before writing code. The project separates concerns between synchronous username scanning and asynchronous email validation, with shared utilities handling HTTP orchestration and result formatting.
Core Engine Components
The engine in user_scanner/core/engine.py handles orchestration, concurrency, and optional AI-agent integration via MCP. It coordinates scans but delegates specific platform logic to individual validator modules.
Key supporting files include:
user_scanner/core/orchestrator.py– Provides standardized HTTP helpers likegeneric_validate,impersonate_validate, andstatus_validatethat reduce boilerplate across modulesuser_scanner/core/result.py– Defines theResultclass that every validator must return, supporting states for available, taken, and error conditionsuser_scanner/mcp/server.py– Optional MCP server enabling LLMs to drive scans autonomously
Validator Module Structure
Modules are organized by scan type and logical category:
- Username modules live in
user_scanner/user_scan/<category>/(e.g.,social/,dev/). These are synchronous and must expose the function signaturedef validate_<site>(user: str) -> Result. - Email modules live in
user_scanner/email_scan/<category>/. These are asynchronous and must exposeasync def validate_<service>(email: str) -> Result.
Each module should be a single Python file named after the platform in lowercase with no spaces.
Step-by-Step Contribution Workflow
1. Read the Contribution Guide
Start by reviewing CONTRIBUTING.md in the repository root. This document contains the full checklist, naming conventions, and style rules specific to the project. Adhering to these guidelines ensures your pull request passes automated checks on the first submission.
2. Choose a Category
Determine whether you are adding a username scanner or an email validator, then select the appropriate subdirectory under user_scanner/user_scan/ or user_scanner/email_scan/. Create a new Python file named after the platform (e.g., example.py for ExampleSocial).
3. Implement the Validator
Use the appropriate helper from user_scanner/core/orchestrator.py to handle HTTP requests:
generic_validate– For most sites. Pass a customprocess(response)callback that explicitly checks both taken and available states. Never rely solely on a raw HTTP 200 status.impersonate_validate– For sites that block standard HTTP clients. This leveragescurl_cffito mimic a real browser fingerprint.
Your validator function must accept the target identifier (username or email) as a string and return a Result object.
4. Return Proper Result Objects
Always return one of the following from user_scanner/core/result.py:
Result.available()– When the handle does not exist on the platformResult.taken(extra={...}, media={...})– When the handle exists, optionally attaching profile metadata and image URLsResult.error("short message")– For network failures, rate limits, or unexpected responses
Explicit verification is critical: your code must distinguish between "not found" and "found" states rather than assuming any successful HTTP response indicates availability.
5. Document Your Changes
Update the documentation in the docs/ directory:
- Add CLI flag references to
docs/FLAGS.mdif your module introduces new command-line options - Document any new pattern syntax in
docs/PATTERNS.md - Create new doc pages if your implementation introduces novel behaviors
6. Run Tests
Validate your changes locally before submitting:
- Execute
ruff check .for linting - Run
mypy user_scannerfor type checking - Run
pytestto execute the test suite
New modules are tested via live scans against both real and dummy handles rather than unit mocks, ensuring they work against actual platform APIs.
Code Examples
Username Module Example
Below is a minimal synchronous username validator following project conventions:
# user_scanner/user_scan/social/example.py
from user_scanner.core.orchestrator import generic_validate
from user_scanner.core.result import Result
def validate_example(user: str) -> Result:
"""Check whether a username exists on ExampleSocial."""
url = f"https://api.examplesocial.com/users/{user}"
show_url = f"https://www.examplesocial.com/{user}"
def process(resp):
# Explicitly detect the "not found" state
if resp.status_code == 404 or "User not found" in resp.text:
return Result.available()
# Detect the "taken" state and extract useful metadata
if resp.status_code == 200 and "profileData" in resp.text:
extra = {}
media = {}
# Example: extract JSON payload embedded in the page
# (implementation omitted for brevity)
return Result.taken(extra=extra, media=media)
# Graceful fallback for unexpected responses
return Result.error(f"Unexpected status {resp.status_code}")
return generic_validate(url, process, show_url=show_url, follow_redirects=True)
Email Module Example
Below is an asynchronous email validator using direct HTTPX:
# user_scanner/email_scan/social/examplemail.py
import httpx
from user_scanner.core.result import Result
async def _check(email: str) -> Result:
url = "https://signup.examplemail.com/check"
payload = {"email": email}
async with httpx.AsyncClient(timeout=15) as client:
resp = await client.post(url, data=payload)
if resp.status_code == 200 and "already registered" in resp.text:
return Result.taken()
if resp.status_code == 200 and "available" in resp.text:
return Result.available()
return Result.error(f"Unexpected response {resp.status_code}")
async def validate_examplemail(email: str) -> Result:
"""Public async validator for ExampleMail."""
return await _check(email)
Summary
- Contributions to user-scanner focus on adding new scan modules for username or email OSINT across web services.
- Username validators are synchronous and live in
user_scanner/user_scan/<category>/, while email validators are asynchronous and live inuser_scanner/email_scan/<category>/. - Use helper functions from
user_scanner/core/orchestrator.pylikegeneric_validateandimpersonate_validateto handle HTTP logic. - Always return explicit
Resultobjects indicatingavailable(),taken(), orerror()states. - Follow the full checklist in
CONTRIBUTING.mdand validate your code withruff,mypy, andpytestbefore submitting.
Frequently Asked Questions
What file should I edit to add a new website for username scanning?
Create a new Python file in user_scanner/user_scan/<category>/ named after the platform (e.g., github.py). The file must contain a function named validate_<site> that accepts a username string and returns a Result object from user_scanner/core/result.py.
Do I need to write my own HTTP request logic for every new module?
No. Import generic_validate or impersonate_validate from user_scanner/core/orchestrator.py to handle HTTP sessions, retries, and browser impersonation. You only need to provide a process callback function that interprets the HTTP response and returns the appropriate Result.
How do I handle websites that return HTTP 200 for both existing and non-existing users?
Never rely on status codes alone. In your process callback, explicitly check response content for strings that indicate "user not found" or "profile data present." Return Result.available() only when you confirm the absence of the profile, and Result.taken() only when you confirm its presence with supporting evidence.
What testing is required before submitting a pull request?
Run ruff check . for linting, mypy user_scanner for type checking, and pytest to execute the test suite. The project tests modules against live platforms using real and dummy handles rather than mocked responses, so ensure your validator works against the actual site before committing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →