# How Maigret Extracts Profile Information Using socid-extractor

> Learn how Maigret extracts profile info with socid-extractor. Discover how it parses pages, analyzes HTML, and normalizes identifiers using SUPPORTED_IDS.

- Repository: [Soxoj/maigret](https://github.com/soxoj/maigret)
- Tags: deep-dive
- Published: 2026-04-30

---

**Maigret delegates profile data extraction to the socid-extractor library by calling `parse()` to fetch pages and `extract()` to analyze HTML, then normalizes discovered identifiers against a whitelist defined in `SUPPORTED_IDS` according to the soxoj/maigret source code.**

Maigret is an open-source OSINT tool that automates username investigations across hundreds of social platforms. Instead of maintaining site-specific scrapers for every supported service, the codebase leverages the **socid-extractor** library to perform pattern-based extraction of profile identifiers, user IDs, and biographical data directly from raw HTML responses.

## The Extraction Pipeline

The integration between Maigret and socid-extractor follows a three-stage pipeline implemented in [`maigret/maigret.py`](https://github.com/soxoj/maigret/blob/main/maigret/maigret.py) and [`maigret/checking.py`](https://github.com/soxoj/maigret/blob/main/maigret/checking.py).

### Fetching Pages with parse()

The workflow begins in [`maigret/maigret.py`](https://github.com/soxoj/maigret/blob/main/maigret/maigret.py) (lines 65-66) where the `extract_ids_from_page` function calls `socid_extractor.parse()`. This method handles HTTP requests, cookie persistence, header injection, and HTML decoding while supporting URL mutations for alternative endpoints.

```python

# fetch the page (including mutated URLs)

page, _ = parse(url, cookies_str='', headers=headers, timeout=timeout)
logger.debug(page)

# run socid-extractor on the HTML

info = extract(page)

```

### Pattern-Based Data Extraction

Once retrieved, the raw HTML is passed to `socid_extractor.extract()` (lines 66-68 in [`maigret/maigret.py`](https://github.com/soxoj/maigret/blob/main/maigret/maigret.py)). This function analyzes the DOM structure against predefined regex patterns and CSS selectors for supported sites, returning a dictionary containing discovered identifiers such as usernames, numeric IDs, and links to connected accounts.

### Normalizing Against SUPPORTED_IDS

After extraction, Maigret normalizes results in [`maigret/checking.py`](https://github.com/soxoj/maigret/blob/main/maigret/checking.py) (line 84) by validating keys against the `SUPPORTED_IDS` tuple defined in lines 36-46. Only identifiers matching this whitelist are preserved and mapped to standardized types.

```python
SUPPORTED_IDS = (
    "username",
    "yandex_public_id",
    "gaia_id",
    "vk_id",
    "ok_id",
    "wikimapia_uid",
    "steam_id",
    "uidme_uguid",
    "yelp_userid",
)

```

The normalization logic filters the extracted dictionary:

```python
if k in SUPPORTED_IDS:
    results[v] = k

```

## Recursive Profile Discovery

When the `--parse` or `--extracting` flag is enabled, Maigret implements recursive extraction by feeding newly discovered profile URLs back into the same `extract_ids_from_page` pipeline. This allows the tool to uncover additional identities linked from the original page, creating a comprehensive map of connected accounts without requiring manual intervention.

## Code Examples

### Extracting IDs from a Single URL

You can use Maigret’s helper function directly to extract identifiers from any supported page:

```python
from maigret.maigret import extract_ids_from_page
import logging

logger = logging.getLogger(__name__)

url = "https://twitter.com/example_user"
ids = extract_ids_from_page(url, logger, timeout=10)

print(ids)

# → {'example_user': 'username', '123456789': 'yandex_public_id', ...}

```

### Using the Maigret Class API

For full searches with extraction enabled via the programmatic API:

```python
from maigret.maigret import Maigret

# Initialise with default settings

maig = Maigret()

# Run a search that enables profile parsing (`--parse` behavior)

result = maig.run(
    usernames=["example_user"],
    timeout=15,
    extract=True,          # equivalent to --parse / --extracting

)

print(result.ids_data)   # populated by socid-extractor

# → {'example_user': {'ids_usernames': {...}, 'ids_links': [...], ...}}

```

### Direct socid-extractor Usage

For custom scripts, invoke the library functions directly without initializing Maigret:

```python
from socid_extractor import parse, extract

html, _ = parse("https://github.com/example_user")
profile_info = extract(html)

print(profile_info.get("username"))          # "example_user"

print(profile_info.get("ids_usernames"))     # {"example_user": "username"}

```

## Summary

- Maigret uses **socid-extractor**'s `parse()` function for HTTP fetching and `extract()` for HTML analysis, located in [`maigret/maigret.py`](https://github.com/soxoj/maigret/blob/main/maigret/maigret.py) (lines 64-68).
- Extracted identifiers are validated against the **SUPPORTED_IDS** whitelist in [`maigret/checking.py`](https://github.com/soxoj/maigret/blob/main/maigret/checking.py) (lines 36-46) to ensure standardization.
- The `extract_ids_from_page` function orchestrates the workflow, handling network timeouts, cookies, and header management.
- **Recursive extraction** is available via the `--parse` flag, enabling deep OSINT investigations that follow links between profiles.

## Frequently Asked Questions

### What is socid-extractor and why does Maigret depend on it?

socid-extractor is a standalone Python library that extracts user identifiers from social media HTML using regex patterns and DOM parsing. Maigret depends on it to avoid maintaining hundreds of site-specific parsers, instead leveraging socid-extractor's pattern database to support new sites without code changes.

### How does Maigret validate extracted identifiers before storing them?

Maigret validates extracted data by checking dictionary keys against the `SUPPORTED_IDS` tuple defined in [`maigret/checking.py`](https://github.com/soxoj/maigret/blob/main/maigret/checking.py) (lines 36-46). Only identifiers matching this whitelist—such as `"username"`, `"gaia_id"`, or `"vk_id"`—are preserved and mapped to a flat `{value: id_type}` structure at line 84.

### Can I extract profile data without running a full Maigret scan?

Yes. You can import `extract_ids_from_page` directly from `maigret.maigret` for URL-specific extraction, or use socid-extractor's `parse()` and `extract()` functions independently as shown in Example 3 above. This allows targeted extraction without initializing the full Maigret search machinery.

### Where is the socid-extractor dependency declared in the repository?

The dependency is listed in [`pyproject.toml`](https://github.com/soxoj/maigret/blob/main/pyproject.toml) (line 65) within the project root. This ensures socid-extractor is installed alongside Maigret when using `pip install maigret`, providing the `parse` and `extract` functions that power the profile discovery pipeline.