Using MediaCrawler FastAPI WebUI for Programmatic Crawler Control

The MediaCrawler FastAPI WebUI exposes REST endpoints that let you start, monitor, and stop crawler processes programmatically without modifying the core command-line logic.

The MediaCrawler project includes a lightweight FastAPI-based web interface that transforms the standalone Python crawler into an API-driven service. By leveraging the MediaCrawler FastAPI WebUI for programmatic crawler control, developers can automate data collection workflows, integrate with external systems, and monitor execution status through standardized HTTP endpoints.

Core Architecture Components

The FastAPI layer acts as a bridge between the browser (or any HTTP client) and the underlying command-line crawler. It isolates the heavy crawling logic from the API server, ensuring the web interface remains responsive while the crawler runs in a separate subprocess.

FastAPI Application Layer

In main/api/main.py, the FastAPI instance is created and configured with CORS policies to allow cross-origin requests. This file also mounts static assets for the React-based frontend and registers API routers under the /api prefix. Health check endpoints (/api/health, /api/env/check) validate that the environment can execute crawler commands, specifically verifying that uv is installed.

Crawler Manager Service

The api/services/crawler_manager.py module implements a singleton service that orchestrates the crawler process. Key responsibilities include:

  • Process Spawning: Builds and executes command lines using uv run python main.py with platform-specific flags
  • Log Aggregation: Captures stdout lines via CrawlerManager._read_output(), parses log levels, and stores entries in an in-memory FIFO queue (limited to 500 entries)
  • Lifecycle Management: Sends SIGTERM for graceful shutdown, falling back to SIGKILL after 15 seconds if the process hangs
  • State Tracking: Maintains current status (idle, running, stopping, error) and metadata about the active crawl

Request Validation and Routing

The api/routers/crawler.py defines the /api/crawler/* endpoints, while api/schemas/crawler.py contains the Pydantic models. The CrawlerStartRequest schema validates incoming JSON payloads, ensuring parameters like platform, login_type, and crawler_type conform to expected values before the manager spawns the subprocess.

API Control Flow

When you interact with the MediaCrawler FastAPI WebUI programmatically, requests follow a predictable pipeline:

  1. Validation: Client POSTs to /api/crawler/start with a JSON payload matching the CrawlerStartRequest schema
  2. Execution: The router forwards parameters to crawler_manager.start(), which constructs the command line and launches a non-blocking subprocess via uv run
  3. Streaming: As the crawler writes to stdout, _read_output() captures lines, populates the internal log store, and pushes updates via WebSocket for live UI consumption
  4. Polling: Clients query /api/crawler/status for state (running, idle, etc.) and /api/crawler/logs?limit=N for recent entries
  5. Termination: POST to /api/crawler/stop triggers the manager's shutdown sequence

All endpoints are automatically documented at /docs (Swagger UI) when the server is running.

Programmatic Control Examples

Starting a Crawl via Python

Submit crawl jobs using standard HTTP libraries. The payload must match the CrawlerStartRequest schema supported by api/schemas/crawler.py:

import requests

url = "http://localhost:8080/api/crawler/start"
payload = {
    "platform": "dy",                # Douyin

    "login_type": "qrcode",
    "crawler_type": "search",
    "keywords": "travel vlog",
    "start_page": 1,
    "enable_comments": True,
    "enable_sub_comments": False,
    "save_option": "jsonl",
    "headless": True
}

resp = requests.post(url, json=payload)
print(resp.json())

Checking Crawler Status

Query the current execution state to determine if the manager is idle, running, stopping, or in an error state:

curl http://localhost:8080/api/crawler/status

Response:

{
  "status": "running",
  "platform": "dy",
  "crawler_type": "search",
  "started_at": "2026-07-31T12:34:56.789012",
  "error_message": null
}

Retrieving Execution Logs

Fetch recent log entries from the in-memory FIFO queue. Use the limit parameter to control how many of the maximum 500 stored entries to return:

import requests

r = requests.get("http://localhost:8080/api/crawler/logs?limit=50")
logs = r.json()["logs"]

for entry in logs:
    print(f"[{entry['timestamp']}] {entry['level'].upper()}: {entry['message']}")

Stopping the Crawler

Initiate graceful shutdown. The manager sends SIGTERM and escalates to SIGKILL if the process does not terminate within 15 seconds:

curl -X POST http://localhost:8080/api/crawler/stop

Discovering Supported Platforms

Query configuration endpoints to populate your application with valid platform options before initiating crawls:

curl http://localhost:8080/api/config/platforms

Response includes metadata for supported targets:

{
  "platforms": [
    {"value":"xhs","label":"Xiaohongshu","icon":"book-open"},
    {"value":"dy","label":"Douyin","icon":"music"},
    {"value":"ks","label":"Kuaishou","icon":"video"},
    {"value":"bili","label":"Bilibili","icon":"tv"},
    {"value":"wb","label":"Weibo","icon":"message-circle"},
    {"value":"tieba","label":"Baidu Tieba","icon":"messages-square"},
    {"value":"zhihu","label":"Zhihu","icon":"help-circle"}
  ]
}

Key Implementation Files

Understanding these source files helps when extending the API or debugging behavior:

Summary

The MediaCrawler FastAPI WebUI provides robust programmatic control through a clean REST interface:

  • Process Isolation: The CrawlerManager service spawns crawlers as separate subprocesses via uv run, keeping the API responsive
  • In-Memory Logging: Real-time logs are buffered in a 500-entry FIFO queue accessible via HTTP or WebSocket
  • Graceful Shutdown: Automatic SIGTERM escalation to SIGKILL ensures processes terminate cleanly or are forcibly stopped after 15 seconds
  • Schema Validation: Pydantic models in api/schemas/crawler.py ensure type safety and prevent malformed crawler configurations

Frequently Asked Questions

Does the MediaCrawler API support concurrent crawler instances?

No. The CrawlerManager is implemented as a singleton service that maintains state for a single subprocess. Attempting to start a new crawl while another is running will return an error status indicating the system is busy. Queue multiple jobs externally or wait for the current process to reach idle status before launching another.

How long are logs preserved in the FastAPI WebUI?

Logs exist only in memory within the CrawlerManager instance and are lost when the API server restarts. The internal FIFO queue stores a maximum of 500 entries, discarding older lines as new ones arrive. For persistent logging, consume the WebSocket stream or poll the /api/crawler/logs endpoint frequently to export data to external storage.

Can I use the API without the React frontend?

Yes. The FastAPI backend in main/api/main.py serves static files only when the built React assets exist in api/webui/. If these files are absent, the REST endpoints remain fully functional. Any HTTP client—Python requests, curl, Postman, or custom scripts—can interact with the crawler control endpoints directly.

What authentication protects the crawler endpoints?

As implemented in the current source code, the MediaCrawler FastAPI WebUI does not include built-in authentication middleware. When deployed in production, place the service behind a reverse proxy (Nginx, Traefik) with basic auth, OAuth2, or IP whitelisting, or extend main/api/main.py to add FastAPI dependency injection for API key validation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →