How the Flask Web Interface Works in Maigret: Architecture and Implementation

The Flask web interface in maigret orchestrates asynchronous username searches across hundreds of sites using background threads, returning results through a three-stage workflow while automatically generating CSV, JSON, PDF, and HTML reports.

The open-source maigret repository by soxoj provides a comprehensive OSINT tool for username investigations. While the command-line interface handles bulk operations, the Flask web interface in maigret offers a browser-based alternative that visualizes search progress and aggregates reports. This lightweight application, defined entirely in maigret/web/app.py, wraps the core async search engine with a user-friendly HTML frontend.

Application Bootstrap and Configuration

The Flask application initializes with standard patterns but includes custom configuration for the maigret search engine. In maigret/web/app.py (lines 31-35), the app sets critical paths required for operation:

app = Flask(__name__)
app.secret_key = os.getenv('FLASK_SECRET_KEY', os.urandom(24).hex())
app.config['MAIGRET_DB_FILE'] = os.path.join(os.path.dirname(__file__), '../resources/data.json')
app.config['REPORTS_FOLDER'] = '/tmp/maigret_reports'

The configuration keys include:

  • MAIGRET_DB_FILE: Points to resources/data.json containing the supported sites database
  • COOKIES_FILE: Optional authentication cookies for authenticated lookups
  • REPORTS_FOLDER: Temporary directory (/tmp/maigret_reports) where generated reports are stored
  • UPLOAD_FOLDER: Directory for file uploads (currently reserved for future use)

Background Job Architecture

Searching across hundreds of services is I/O-intensive, so the Flask web interface in maigret handles each request asynchronously using Python threads. This prevents HTTP timeouts and allows immediate feedback to users.

The system maintains two global dictionaries:

  • background_jobs: Maps timestamps to thread objects and completion status
  • job_results: Stores final search results and report metadata

The heavy lifting occurs in process_search_task() (lines 101-192), which:

  1. Creates a fresh asyncio event loop
  2. Calls search_multiple_usernames() to execute maigret_search() across all targets
  3. Generates reports via maigret.report.* helpers
  4. Populates job_results with session folders, graph files, and metadata
  5. Sets background_jobs[timestamp]['completed'] = True upon finish (line 192)

Route Handling and Request Flow

The application exposes five primary routes defined in app.py:

Route Method Purpose
/ GET Renders the search form with site autocomplete
/search POST Accepts usernames and options, starts background job
/status/<timestamp> GET Polls job status; redirects to results when complete
/results/<session_id> GET Displays the combined graph and download links
/reports/<path:filename> GET Safely serves generated files from REPORTS_FOLDER

The request flow follows a specific pattern:

  1. index() (lines 95-112) loads MaigretDatabase().load_from_path and renders templates/index.html
  2. search() (lines 146-166) parses form data, generates a unique timestamp, and spawns process_search_task in a daemon thread
  3. status() (lines 169-197) polls every few seconds; when completed becomes True, it redirects to the results page
  4. results() (lines 200-224) reads job_results to display the graph and per-user reports
  5. download_report() (lines 226-336) validates file paths and streams requested reports

Asynchronous Search Execution

The actual username lookups leverage maigret's core async API through maigret_search() (lines 43-84). This function constructs a filtered site dictionary using MaigretDatabase().ranked_sites_dict, then awaits the core search coroutine:

results = await maigret.search(
    username=username,
    site_dict=sites,
    timeout=int(options.get('timeout', 30)),
    logger=logger,
    id_type='username',
    cookies=app.config["COOKIES_FILE"] if options.get('use_cookies') else None,
    is_parsing_enabled=(not options.get('disable_extracting', False)),
    recursive_search_enabled=(not options.get('disable_recursive_search', False))
)

The function returns a dictionary mapping site names to MaigretResult objects containing status codes, discovered URLs, and extracted metadata.

Template Rendering and Static Assets

The UI uses three Jinja2 templates stored under maigret/web/templates/:

  • index.html: Search form with multi-select tag controls and site autocomplete
  • status.html: Auto-refreshing waiting page displayed during background processing
  • results.html: Final view showing the combined network graph, username list, and format-specific download links

Static assets including logos and screenshots reside in maigret/web/static/. All templates are rendered via Flask's render_template() function in their respective route handlers.

Security and Deployment Considerations

The Flask web interface in maigret implements several safety measures:

  • Secret Key Management: Uses FLASK_SECRET_KEY environment variable or generates a random 24-byte hex string at startup
  • Path Traversal Protection: The download_report() route validates requested paths using os.path.normpath to ensure files remain within REPORTS_FOLDER
  • Network Binding: Defaults to 127.0.0.1 unless FLASK_HOST is explicitly overridden, minimizing exposure on production systems

Running the Web Interface Locally

To start the development server locally:

export FLASK_SECRET_KEY=$(openssl rand -hex 12)
export FLASK_DEBUG=1
python -m maigret.web.app

The if __name__ == '__main__' block (lines 41-53) configures logging and launches the server on http://127.0.0.1:5000 by default.

You can submit searches programmatically using curl:

curl -X POST -d "usernames=alice,bob" -d "top_sites=100" \
     http://127.0.0.1:5000/search -L

After completion, download specific report formats:

curl -O http://127.0.0.1:5000/reports/search_20241012_153045/report_alice.pdf

Summary

  • The Flask web interface in maigret is defined in maigret/web/app.py and provides a complete browser-based alternative to the CLI
  • Background threads handle I/O-heavy searches via process_search_task(), preventing HTTP timeouts
  • The application uses a three-stage workflow: search submission (/search), status polling (/status), and results display (/results)
  • Generated reports include CSV, JSON, PDF, and HTML formats stored temporarily in /tmp/maigret_reports
  • Security features include path validation for downloads, random secret key generation, and localhost-only binding by default

Frequently Asked Questions

How do I start the maigret Flask web interface locally?

Set an optional FLASK_SECRET_KEY environment variable and run python -m maigret.web.app from the repository root. The server starts on 127.0.0.1:5000 by default, configurable via the FLASK_HOST and FLASK_PORT environment variables as implemented in the main execution block (lines 41-53).

What file formats does the web interface generate?

The interface automatically generates four report formats for each username searched: CSV for spreadsheet analysis, JSON for programmatic processing, PDF for portable documents, and HTML for browser viewing. These are written to the REPORTS_FOLDER directory (default /tmp/maigret_reports) by the report generation utilities called within process_search_task().

How does the maigret web interface prevent path traversal attacks?

The download_report() route (lines 226-336) validates requested filenames using os.path.normpath and verifies the resolved path starts with the configured REPORTS_FOLDER. This prevents attackers from accessing files outside the intended directory using ../ sequences or other traversal techniques.

Why does maigret use background threads for web searches instead of synchronous processing?

Username searches require concurrent HTTP requests across hundreds of sites, making them I/O-bound operations that can take minutes to complete. Running these in process_search_task() within a daemon thread allows the Flask route to return immediately with a status page, preventing Gateway Timeout errors and allowing users to monitor progress via the /status/<timestamp> polling endpoint.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →