What Kind of User Information Can User-Scanner Scan For: Complete OSINT Data Guide
User-scanner is a Python-based OSINT suite that extracts account existence status, public profile metadata (bios, follower counts, karma), media assets (avatars, banners), and canonical profile URLs from over 460 platforms using both username and email-based searches.
User-scanner by kaifcodec is an open-source intelligence (OSINT) framework designed to interrogate digital identities across hundreds of online services. Whether you are investigating a username or verifying an email address, understanding what kind of user information user-scanner can scan for is essential for effective reconnaissance. The tool aggregates data into structured Result objects that expose registration status, rich metadata, and media assets through a unified Python API.
The Core Result Object Structure
Every scan operation returns a Result instance defined in user_scanner/core/result.py. This object standardizes the data schema across all 460+ supported platforms, ensuring consistent access to five primary data fields.
Account Status and Availability
The status field indicates whether a target is taken (Found/Registered), available (Not Found/Not Registered), or resulted in an error/skipped state. According to the Result.Status class implementation, this classification allows automated pipelines to filter for active accounts versus available handles for registration.
Profile Metadata (The extra Dictionary)
The extra field contains a free-form dictionary of metadata extracted by each platform module. As implemented in Result.update() (lines 106-119), this dictionary captures platform-specific fields such as:
- Display names –
extra["name"]stores real names or handles (e.g., Reddit, Instagram) - Biographical data –
extra["bio"]holds profile descriptions - Engagement metrics –
extra["followers"],extra["subscriber_count"], andextra["karma_total"]quantify account reach - Account provenance –
extra["created"]timestamps registration dates - Verification states –
extra["verified"]andextra["has_verified_email"]indicate platform verification - Custom status flags –
extra["status"]reports Reddit-specific states like "deleted" or "suspended"
Media Assets (The media Dictionary)
The media dictionary stores URLs to publicly accessible visual assets. Populated by each module’s _extract_media routine (see user_scanner/user_scan/social/reddit.py for reference), this field typically includes:
- Avatar URLs –
media["avatar"]links to profile pictures - Banner images –
media["banner"]stores cover photos - Platform-specific assets – Such as Reddit’s
snoovatarcustom avatars
Canonical Profile URLs
The url field provides the direct link to the public profile page that was queried. This enables immediate browser pivoting or archival without requiring manual URL construction.
Types of User Information Extracted
User-scanner differentiates between username-based and email-based investigations, with each module populating the Result object according to the data exposed by the target platform.
Existence and Registration Verification
For username modules (e.g., user_scanner/user_scan/social/reddit.py, github.py), the tool reports whether a handle is taken or available. For email modules (e.g., user_scanner/email_scan/social/instagram.py, gmail.py), it determines if an address is registered or not registered on the platform.
Deep Profile Metadata Extraction
Each platform-specific scanner extracts whatever public fields the site exposes through the _extract_profile helper methods. The Reddit module, for instance, surfaces employee status and total karma, while Instagram modules expose private account flags and follower counts. This metadata mining occurs after the initial existence check, providing intelligence beyond simple boolean availability.
Cross-Scan Pivot Data
When invoked with the --cross-scan flag, the engine mines links, secondary usernames, and secondary email addresses from the initial results. This pivot data is recursively scanned and surfaced via extra["links"], enabling network mapping of connected digital identities.
Architecture and Data Flow
Understanding the pipeline clarifies how user-scanner aggregates this information.
Engine Orchestration
The engine.check() function in user_scanner/core/engine.py determines whether the target is a username or email and loads the appropriate module set (user_scan versus email_scan). This high-level orchestration manages the asynchronous execution of platform modules.
Module Validation and Extraction
Each module implements a validate_<site> function (e.g., validate_reddit) that performs HTTP requests—often utilizing the impersonation helpers in user_scanner/core/impersonate.py to bypass bot detection. Based on the response, the module returns a Result instance via Result.taken, Result.available, or Result.error factory methods.
Practical Implementation Examples
Scanning a Single Username
import asyncio
from user_scanner.core import engine
from user_scanner.user_scan.social.reddit import validate_reddit
async def main():
result = await engine.check(validate_reddit, "spez")
print(result.to_json())
result.show()
asyncio.run(main())
This returns a JSON structure containing the status, karma totals, creation dates, and avatar URLs as defined in the Result class.
Scanning a Single Email Address
import asyncio
from user_scanner.core import engine
from user_scanner.email_scan.social import instagram
async def main():
result = await engine.check(instagram, "example@gmail.com")
print(result.to_json())
asyncio.run(main())
Email scans populate the same Result schema, with the extra dictionary containing mapped usernames and privacy flags specific to the platform.
Bulk Scanning and Export
import asyncio
from user_scanner.core import engine
from user_scanner.core.formatter import into_json
async def main():
usernames = ["alice", "bob", "charlie"]
results = await engine.check_all(usernames, is_email=False)
with open("report.json", "w") as out:
out.write(into_json(results))
asyncio.run(main())
The engine.check_all() method aggregates multiple Result objects into JSON, CSV, or PDF formats via the formatter utilities.
Summary
- User-scanner extracts existence status, metadata (bios, karma, followers), media URLs (avatars, banners), and profile links from over 460 platforms.
- Data is standardized in the
Resultclass (user_scanner/core/result.py), providing consistent fields:status,url,extra,media, andreason. - Both usernames and email addresses are supported through modular scanners in
user_scanner/user_scan/anduser_scanner/email_scan/. - Cross-scan capabilities enable recursive investigation of linked accounts found in initial searches.
- Output formats include JSON, CSV, and PDF via the formatter module for integration into automated pipelines.
Frequently Asked Questions
What types of targets can user-scanner investigate?
User-scanner supports both username-based and email-based investigations across social networks, code repositories, gaming platforms, and communication services. The engine automatically routes targets to the appropriate validation modules in user_scanner/user_scan/ or user_scanner/email_scan/ based on input format detection.
How does user-scanner handle bot detection and rate limiting?
The tool implements impersonation techniques in user_scanner/core/impersonate.py, utilizing curl-cffi sessions to mimic legitimate browser fingerprints. This allows modules like validate_reddit to bypass basic bot walls and access public profile data that would otherwise be blocked to automated requests.
Can user-scanner extract private account information?
No. User-scanner only accesses publicly available data exposed by platform APIs or HTML responses. Fields like extra["private"] indicate privacy settings but do not grant access to protected content. The tool respects platform visibility settings, reporting only what is openly accessible without authentication.
What output formats does user-scanner support?
The Result object provides native serialization via to_json() for programmatic use, while the formatter utilities in user_scanner/core/formatter.py enable export to JSON arrays, CSV spreadsheets, and PDF reports. This multi-format support facilitates integration into data science workflows, incident response platforms, and documentation pipelines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →