How Maigret Extracts Profile Information Using socid-extractor
Maigret delegates profile data extraction to the socid-extractor library by calling parse() to fetch pages and extract() to analyze HTML, then normalizes discovered identifiers against a whitelist defined in SUPPORTED_IDS according to the soxoj/maigret source code.
Maigret is an open-source OSINT tool that automates username investigations across hundreds of social platforms. Instead of maintaining site-specific scrapers for every supported service, the codebase leverages the socid-extractor library to perform pattern-based extraction of profile identifiers, user IDs, and biographical data directly from raw HTML responses.
The Extraction Pipeline
The integration between Maigret and socid-extractor follows a three-stage pipeline implemented in maigret/maigret.py and maigret/checking.py.
Fetching Pages with parse()
The workflow begins in maigret/maigret.py (lines 65-66) where the extract_ids_from_page function calls socid_extractor.parse(). This method handles HTTP requests, cookie persistence, header injection, and HTML decoding while supporting URL mutations for alternative endpoints.
# fetch the page (including mutated URLs)
page, _ = parse(url, cookies_str='', headers=headers, timeout=timeout)
logger.debug(page)
# run socid-extractor on the HTML
info = extract(page)
Pattern-Based Data Extraction
Once retrieved, the raw HTML is passed to socid_extractor.extract() (lines 66-68 in maigret/maigret.py). This function analyzes the DOM structure against predefined regex patterns and CSS selectors for supported sites, returning a dictionary containing discovered identifiers such as usernames, numeric IDs, and links to connected accounts.
Normalizing Against SUPPORTED_IDS
After extraction, Maigret normalizes results in maigret/checking.py (line 84) by validating keys against the SUPPORTED_IDS tuple defined in lines 36-46. Only identifiers matching this whitelist are preserved and mapped to standardized types.
SUPPORTED_IDS = (
"username",
"yandex_public_id",
"gaia_id",
"vk_id",
"ok_id",
"wikimapia_uid",
"steam_id",
"uidme_uguid",
"yelp_userid",
)
The normalization logic filters the extracted dictionary:
if k in SUPPORTED_IDS:
results[v] = k
Recursive Profile Discovery
When the --parse or --extracting flag is enabled, Maigret implements recursive extraction by feeding newly discovered profile URLs back into the same extract_ids_from_page pipeline. This allows the tool to uncover additional identities linked from the original page, creating a comprehensive map of connected accounts without requiring manual intervention.
Code Examples
Extracting IDs from a Single URL
You can use Maigret’s helper function directly to extract identifiers from any supported page:
from maigret.maigret import extract_ids_from_page
import logging
logger = logging.getLogger(__name__)
url = "https://twitter.com/example_user"
ids = extract_ids_from_page(url, logger, timeout=10)
print(ids)
# → {'example_user': 'username', '123456789': 'yandex_public_id', ...}
Using the Maigret Class API
For full searches with extraction enabled via the programmatic API:
from maigret.maigret import Maigret
# Initialise with default settings
maig = Maigret()
# Run a search that enables profile parsing (`--parse` behavior)
result = maig.run(
usernames=["example_user"],
timeout=15,
extract=True, # equivalent to --parse / --extracting
)
print(result.ids_data) # populated by socid-extractor
# → {'example_user': {'ids_usernames': {...}, 'ids_links': [...], ...}}
Direct socid-extractor Usage
For custom scripts, invoke the library functions directly without initializing Maigret:
from socid_extractor import parse, extract
html, _ = parse("https://github.com/example_user")
profile_info = extract(html)
print(profile_info.get("username")) # "example_user"
print(profile_info.get("ids_usernames")) # {"example_user": "username"}
Summary
- Maigret uses socid-extractor's
parse()function for HTTP fetching andextract()for HTML analysis, located inmaigret/maigret.py(lines 64-68). - Extracted identifiers are validated against the SUPPORTED_IDS whitelist in
maigret/checking.py(lines 36-46) to ensure standardization. - The
extract_ids_from_pagefunction orchestrates the workflow, handling network timeouts, cookies, and header management. - Recursive extraction is available via the
--parseflag, enabling deep OSINT investigations that follow links between profiles.
Frequently Asked Questions
What is socid-extractor and why does Maigret depend on it?
socid-extractor is a standalone Python library that extracts user identifiers from social media HTML using regex patterns and DOM parsing. Maigret depends on it to avoid maintaining hundreds of site-specific parsers, instead leveraging socid-extractor's pattern database to support new sites without code changes.
How does Maigret validate extracted identifiers before storing them?
Maigret validates extracted data by checking dictionary keys against the SUPPORTED_IDS tuple defined in maigret/checking.py (lines 36-46). Only identifiers matching this whitelist—such as "username", "gaia_id", or "vk_id"—are preserved and mapped to a flat {value: id_type} structure at line 84.
Can I extract profile data without running a full Maigret scan?
Yes. You can import extract_ids_from_page directly from maigret.maigret for URL-specific extraction, or use socid-extractor's parse() and extract() functions independently as shown in Example 3 above. This allows targeted extraction without initializing the full Maigret search machinery.
Where is the socid-extractor dependency declared in the repository?
The dependency is listed in pyproject.toml (line 65) within the project root. This ensures socid-extractor is installed alongside Maigret when using pip install maigret, providing the parse and extract functions that power the profile discovery pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →