Using MediaCrawler FastAPI WebUI for Programmatic Crawler Control
The MediaCrawler FastAPI WebUI exposes REST endpoints that let you start, monitor, and stop crawler processes programmatically without modifying the core command-line logic.
The MediaCrawler project includes a lightweight FastAPI-based web interface that transforms the standalone Python crawler into an API-driven service. By leveraging the MediaCrawler FastAPI WebUI for programmatic crawler control, developers can automate data collection workflows, integrate with external systems, and monitor execution status through standardized HTTP endpoints.
Core Architecture Components
The FastAPI layer acts as a bridge between the browser (or any HTTP client) and the underlying command-line crawler. It isolates the heavy crawling logic from the API server, ensuring the web interface remains responsive while the crawler runs in a separate subprocess.
FastAPI Application Layer
In main/api/main.py, the FastAPI instance is created and configured with CORS policies to allow cross-origin requests. This file also mounts static assets for the React-based frontend and registers API routers under the /api prefix. Health check endpoints (/api/health, /api/env/check) validate that the environment can execute crawler commands, specifically verifying that uv is installed.
Crawler Manager Service
The api/services/crawler_manager.py module implements a singleton service that orchestrates the crawler process. Key responsibilities include:
- Process Spawning: Builds and executes command lines using
uv run python main.pywith platform-specific flags - Log Aggregation: Captures stdout lines via
CrawlerManager._read_output(), parses log levels, and stores entries in an in-memory FIFO queue (limited to 500 entries) - Lifecycle Management: Sends
SIGTERMfor graceful shutdown, falling back toSIGKILLafter 15 seconds if the process hangs - State Tracking: Maintains current status (
idle,running,stopping,error) and metadata about the active crawl
Request Validation and Routing
The api/routers/crawler.py defines the /api/crawler/* endpoints, while api/schemas/crawler.py contains the Pydantic models. The CrawlerStartRequest schema validates incoming JSON payloads, ensuring parameters like platform, login_type, and crawler_type conform to expected values before the manager spawns the subprocess.
API Control Flow
When you interact with the MediaCrawler FastAPI WebUI programmatically, requests follow a predictable pipeline:
- Validation: Client POSTs to
/api/crawler/startwith a JSON payload matching theCrawlerStartRequestschema - Execution: The router forwards parameters to
crawler_manager.start(), which constructs the command line and launches a non-blocking subprocess viauv run - Streaming: As the crawler writes to stdout,
_read_output()captures lines, populates the internal log store, and pushes updates via WebSocket for live UI consumption - Polling: Clients query
/api/crawler/statusfor state (running,idle, etc.) and/api/crawler/logs?limit=Nfor recent entries - Termination: POST to
/api/crawler/stoptriggers the manager's shutdown sequence
All endpoints are automatically documented at /docs (Swagger UI) when the server is running.
Programmatic Control Examples
Starting a Crawl via Python
Submit crawl jobs using standard HTTP libraries. The payload must match the CrawlerStartRequest schema supported by api/schemas/crawler.py:
import requests
url = "http://localhost:8080/api/crawler/start"
payload = {
"platform": "dy", # Douyin
"login_type": "qrcode",
"crawler_type": "search",
"keywords": "travel vlog",
"start_page": 1,
"enable_comments": True,
"enable_sub_comments": False,
"save_option": "jsonl",
"headless": True
}
resp = requests.post(url, json=payload)
print(resp.json())
Checking Crawler Status
Query the current execution state to determine if the manager is idle, running, stopping, or in an error state:
curl http://localhost:8080/api/crawler/status
Response:
{
"status": "running",
"platform": "dy",
"crawler_type": "search",
"started_at": "2026-07-31T12:34:56.789012",
"error_message": null
}
Retrieving Execution Logs
Fetch recent log entries from the in-memory FIFO queue. Use the limit parameter to control how many of the maximum 500 stored entries to return:
import requests
r = requests.get("http://localhost:8080/api/crawler/logs?limit=50")
logs = r.json()["logs"]
for entry in logs:
print(f"[{entry['timestamp']}] {entry['level'].upper()}: {entry['message']}")
Stopping the Crawler
Initiate graceful shutdown. The manager sends SIGTERM and escalates to SIGKILL if the process does not terminate within 15 seconds:
curl -X POST http://localhost:8080/api/crawler/stop
Discovering Supported Platforms
Query configuration endpoints to populate your application with valid platform options before initiating crawls:
curl http://localhost:8080/api/config/platforms
Response includes metadata for supported targets:
{
"platforms": [
{"value":"xhs","label":"Xiaohongshu","icon":"book-open"},
{"value":"dy","label":"Douyin","icon":"music"},
{"value":"ks","label":"Kuaishou","icon":"video"},
{"value":"bili","label":"Bilibili","icon":"tv"},
{"value":"wb","label":"Weibo","icon":"message-circle"},
{"value":"tieba","label":"Baidu Tieba","icon":"messages-square"},
{"value":"zhihu","label":"Zhihu","icon":"help-circle"}
]
}
Key Implementation Files
Understanding these source files helps when extending the API or debugging behavior:
main/api/main.py: FastAPI application factory, CORS configuration, static file mounting, and health endpointsapi/routers/crawler.py: Route definitions for/api/crawler/start,/stop,/status, and/logsapi/services/crawler_manager.py: Subprocess orchestration, log line parsing, and signal managementapi/schemas/crawler.py: Pydantic models (CrawlerStartRequest,CrawlerStatusResponse,LogEntry) defining the API contractapi/routers/websocket.py: WebSocket handler for real-time log streaming to the frontendmain/main.py: The core CLI crawler entry point executed by the manager service
Summary
The MediaCrawler FastAPI WebUI provides robust programmatic control through a clean REST interface:
- Process Isolation: The
CrawlerManagerservice spawns crawlers as separate subprocesses viauv run, keeping the API responsive - In-Memory Logging: Real-time logs are buffered in a 500-entry FIFO queue accessible via HTTP or WebSocket
- Graceful Shutdown: Automatic
SIGTERMescalation toSIGKILLensures processes terminate cleanly or are forcibly stopped after 15 seconds - Schema Validation: Pydantic models in
api/schemas/crawler.pyensure type safety and prevent malformed crawler configurations
Frequently Asked Questions
Does the MediaCrawler API support concurrent crawler instances?
No. The CrawlerManager is implemented as a singleton service that maintains state for a single subprocess. Attempting to start a new crawl while another is running will return an error status indicating the system is busy. Queue multiple jobs externally or wait for the current process to reach idle status before launching another.
How long are logs preserved in the FastAPI WebUI?
Logs exist only in memory within the CrawlerManager instance and are lost when the API server restarts. The internal FIFO queue stores a maximum of 500 entries, discarding older lines as new ones arrive. For persistent logging, consume the WebSocket stream or poll the /api/crawler/logs endpoint frequently to export data to external storage.
Can I use the API without the React frontend?
Yes. The FastAPI backend in main/api/main.py serves static files only when the built React assets exist in api/webui/. If these files are absent, the REST endpoints remain fully functional. Any HTTP client—Python requests, curl, Postman, or custom scripts—can interact with the crawler control endpoints directly.
What authentication protects the crawler endpoints?
As implemented in the current source code, the MediaCrawler FastAPI WebUI does not include built-in authentication middleware. When deployed in production, place the service behind a reverse proxy (Nginx, Traefik) with basic auth, OAuth2, or IP whitelisting, or extend main/api/main.py to add FastAPI dependency injection for API key validation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →