How Maigret's Asynchronous Site Checking Engine Operates: Technical Architecture Guide
Maigret uses an asyncio-based worker pool to dispatch thousands of HTTP requests concurrently, coordinating request preparation, parallel execution, and response classification to enable high-speed username enumeration across web services.
The Maigret username discovery tool (soxoj/maigret) achieves rapid multi-site checking through a sophisticated asynchronous site checking engine built on Python's asyncio framework. This architecture enables the tool to verify username existence across thousands of web services simultaneously by orchestrating specialized HTTP checkers through a queue-based worker pool system.
Request Preparation and Checker Selection
The engine begins by preparing individual requests for each target site in maigret/checking.py. The make_site_result function constructs a per-site request configuration and instantiates the appropriate checker class based on the site's protocol requirements.
# maigret/checking.py (lines 11‑15, 44‑63)
options["checkers"][site.protocol] # selects the appropriate checker
checker = SimpleAiohttpChecker(proxy=proxy, cookie_jar=cookie_jar, logger=logger)
# or CurlCffiChecker if TLS fingerprinting is required (lines 50‑55)
future = checker.prepare(
method=request_method,
url=url_probe,
headers=headers,
allow_redirects=allow_redirects,
timeout=options['timeout'],
payload=payload,
)
site_result["future"] = future
The engine supports multiple checker implementations:
- SimpleAiohttpChecker: Uses standard
aiohttpfor HTTP requests - CurlCffiChecker: Employs
curl_cffifor TLS fingerprint impersonation when sites block standard HTTP clients
Each checker object encapsulates the request parameters and exposes a check coroutine that performs the actual HTTP call. The prepared future is stored in a site-specific result dictionary awaiting execution.
Concurrent Execution via AsyncioQueueGeneratorExecutor
Once requests are prepared, maigret/executors.py manages parallel execution through the AsyncioQueueGeneratorExecutor class. This component implements an asynchronous worker pool pattern that feeds prepared checkers into an asyncio.Queue serviced by multiple concurrent workers.
# maigret/executors.py (lines 84‑122)
class AsyncioQueueGeneratorExecutor:
async def worker(self):
f, args, kwargs = await self.queue.get()
query_future = f(*args, **kwargs) # f is the checker's `check` coroutine
query_task = asyncio.create_task(query_future)
result = await asyncio.wait_for(query_task, timeout=self.timeout)
await self._results.put(result)
The executor creates a configurable number of workers (self.workers_count) that continuously pull tasks from the queue. Each worker:
- Retrieves a tuple containing the checker function and arguments
- Creates an
asyncio.Taskfor the checker'scheckcoroutine - Applies per-request timeout using
asyncio.wait_for - Pushes the raw
(html, status, error)result tuple to a results queue
Progress tracking is handled through alive_progress (alive_bar) for real-time monitoring (lines 63‑71).
The high-level maigret function in maigret/maigret.py (lines 46-62) orchestrates this process:
executor = AsyncioQueueGeneratorExecutor(
logger=logger,
in_parallel=max_connections,
timeout=timeout + 0.5,
)
...
async for result in executor.run(list(tasks_dict.values())):
cur_results.append(result)
Response Processing and Result Classification
After execution, raw responses flow into process_site_result in maigret/checking.py (lines 173‑219) for interpretation:
# maigret/checking.py (lines 173‑219)
response = await checker.check() # async HTTP request
...
process_site_result(
response, query_notify, logger, default_result, site
)
The process_site_result function performs several validation steps:
- Error Detection: Identifies generic error pages, CDN blocks, and HTTP 403/429 responses using
detect_error_page(lines 30‑42) - Activation Handling: Manages token refresh requirements (e.g., Twitter guest tokens) through activation helpers (lines 87‑104)
- Presence Verification: Applies site-specific check types (
message,status_code,response_url) to determine if the username is CLAIMED or AVAILABLE (lines 118‑170) - Identity Extraction: Optionally runs
socid_extractorto harvest additional identifiers when parsing is enabled (lines 191‑200)
The final output is encapsulated in a MaigretCheckResult object (defined in maigret/result.py), which stores the check status and metadata in the per-site result dictionary under the "status" key.
Retries and Orchestration
The outer execution loop in maigret/maigret.py (lines 68‑86) implements retry logic that re-queues failed sites until success or exhaustion of retry attempts. Failed sites are identified via get_failed_sites and resubmitted to the executor for additional attempts. A QueryNotifyPrint callback interface allows CLI callers to receive real-time progress updates through start, update, and finish events.
Practical Implementation: Code Examples
Checking a Single Site Asynchronously
For targeted verification of a specific site, use the low-level checker API:
import asyncio
from maigret.checking import SimpleAiohttpChecker, make_site_result, process_site_result
from maigret.utils import get_random_user_agent
import logging
async def check_one(site, username):
logger = logging.getLogger("demo")
logger.setLevel(logging.WARNING)
# Build request options (only the essentials)
options = {
"parsing": False,
"cookie_jar": None,
"timeout": 10,
}
# Prepare the site‑specific dict
result = make_site_result(site, username, options, logger)
# Run the checker that was prepared above
checker = result["checker"]
response = await checker.check() # → (html, status, error)
# Transform raw response into a MaigretCheckResult
processed = process_site_result(
response,
query_notify={}, # dummy notifier for this example
logger=logger,
results_info=result,
site=site,
)
return processed["status"] # MaigretCheckResult object
# Example usage:
# asyncio.run(check_one(my_site_object, "alice"))
This example demonstrates the three-stage flow: prepare → checker.check → process_site_result.
Bulk Username Enumeration with the High-Level API
For scanning across the entire site database, use the high-level maigret function:
import asyncio
from maigret.maigret import maigret
from maigret.sites import MaigretDatabase
async def bulk_check(username):
# Load the full site database (JSON + generated objects)
db = MaigretDatabase().load_from_path("maigret/resources/data.json")
results = await maigret(
username,
site_dict=db.sites_dict, # all known sites
logger=__import__("logging").getLogger("maigret"),
timeout=5,
max_connections=200,
)
return results
# asyncio.run(bulk_check("johnsmith"))
This mirrors the command-line interface, automatically building the executor, managing the worker pool, and aggregating results into a dictionary keyed by site name.
Key Source Files and Components
The asynchronous site checking engine spans several critical files in the soxoj/maigret repository:
maigret/checking.py: Core request preparation viamake_site_result, checker classes (SimpleAiohttpChecker,CurlCffiChecker), and response processing throughprocess_site_resultmaigret/executors.py: ImplementsAsyncioQueueGeneratorExecutormanaging the async worker pool and concurrent execution queuemaigret/maigret.py: High-level orchestrator containing themaigretfunction that coordinates retries, progress reporting, and executor lifecyclemaigret/result.py: DefinesMaigretCheckResultandMaigretCheckStatusdata structures representing final check outcomesmaigret/sites.py: Database management throughMaigretDatabaseandMaigretSiteconfiguration objectsmaigret/activation.py: Token refresh and activation helpers for sites requiring session initialization (e.g., Twitter guest tokens)
Summary
- Maigret's asynchronous site checking engine uses asyncio to dispatch thousands of concurrent HTTP requests through specialized checker classes
- The
AsyncioQueueGeneratorExecutorinmaigret/executors.pyimplements a queue-based worker pool pattern with configurable concurrency limits and timeout handling make_site_resultprepares per-site requests selecting betweenSimpleAiohttpCheckerandCurlCffiCheckerbased on protocol requirementsprocess_site_resultinterprets raw HTTP responses applying site-specific detection rules for error pages, activation requirements, and presence indicators- Results are encapsulated in
MaigretCheckResultobjects stored with full metadata about the check outcome - The high-level
maigretfunction orchestrates retries, progress reporting, and result aggregation across the entire site database
Frequently Asked Questions
How does Maigret handle rate limiting during concurrent checks?
Maigret detects rate limiting through detect_error_page logic in maigret/checking.py that identifies HTTP 403/429 responses and generic CDN block pages. When detected, these sites are marked as failed and can be re-queued through the retry mechanism in maigret/maigret.py. However, the engine does not implement automatic rate limit backoff; users should configure appropriate max_connections and timeout values to avoid triggering aggressive rate limits.
What is the difference between SimpleAiohttpChecker and CurlCffiChecker?
SimpleAiohttpChecker uses standard aiohttp for HTTP requests and works for most sites without anti-bot protection. CurlCffiChecker leverages curl_cffi to perform TLS fingerprint impersonation, masquerading as a standard browser to bypass services that block non-browser HTTP clients. The engine automatically selects the appropriate checker based on the site's protocol configuration in the database.
How does Maigret retry failed site checks?
The outer execution loop in maigret/maigret.py maintains a retry counter that re-submits failed sites to the AsyncioQueueGeneratorExecutor. Failed sites are identified using get_failed_sites and re-queued until either all checks succeed or the retry limit is exhausted. This process runs within the same async event loop, preserving connection pools and cookie jars across retry attempts.
Can the concurrent worker pool size be customized?
Yes, the worker pool size is controlled through the max_connections parameter passed to the high-level maigret function. This value is forwarded to AsyncioQueueGeneratorExecutor as the in_parallel argument, which determines how many async workers are spawned to service the queue. Higher values increase throughput but may trigger rate limits on target services; the default balances speed with polite request patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →