# Using MediaCrawler FastAPI WebUI for Programmatic Crawler Control

> Control MediaCrawler programmatically with its FastAPI WebUI. Start, monitor, and stop crawlers via REST endpoints without altering core logic. Integrate seamless automation today.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**The MediaCrawler FastAPI WebUI exposes REST endpoints that let you start, monitor, and stop crawler processes programmatically without modifying the core command-line logic.**

The MediaCrawler project includes a lightweight FastAPI-based web interface that transforms the standalone Python crawler into an API-driven service. By leveraging the MediaCrawler FastAPI WebUI for programmatic crawler control, developers can automate data collection workflows, integrate with external systems, and monitor execution status through standardized HTTP endpoints.

## Core Architecture Components

The FastAPI layer acts as a bridge between the browser (or any HTTP client) and the underlying command-line crawler. It isolates the heavy crawling logic from the API server, ensuring the web interface remains responsive while the crawler runs in a separate subprocess.

### FastAPI Application Layer

In [`main/api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/api/main.py), the FastAPI instance is created and configured with CORS policies to allow cross-origin requests. This file also mounts static assets for the React-based frontend and registers API routers under the `/api` prefix. Health check endpoints (`/api/health`, `/api/env/check`) validate that the environment can execute crawler commands, specifically verifying that `uv` is installed.

### Crawler Manager Service

The [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py) module implements a **singleton service** that orchestrates the crawler process. Key responsibilities include:

- **Process Spawning**: Builds and executes command lines using `uv run python main.py` with platform-specific flags
- **Log Aggregation**: Captures stdout lines via `CrawlerManager._read_output()`, parses log levels, and stores entries in an in-memory FIFO queue (limited to 500 entries)
- **Lifecycle Management**: Sends `SIGTERM` for graceful shutdown, falling back to `SIGKILL` after 15 seconds if the process hangs
- **State Tracking**: Maintains current status (`idle`, `running`, `stopping`, `error`) and metadata about the active crawl

### Request Validation and Routing

The [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py) defines the `/api/crawler/*` endpoints, while [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py) contains the Pydantic models. The `CrawlerStartRequest` schema validates incoming JSON payloads, ensuring parameters like `platform`, `login_type`, and `crawler_type` conform to expected values before the manager spawns the subprocess.

## API Control Flow

When you interact with the MediaCrawler FastAPI WebUI programmatically, requests follow a predictable pipeline:

1. **Validation**: Client POSTs to `/api/crawler/start` with a JSON payload matching the `CrawlerStartRequest` schema
2. **Execution**: The router forwards parameters to `crawler_manager.start()`, which constructs the command line and launches a non-blocking subprocess via `uv run`
3. **Streaming**: As the crawler writes to stdout, `_read_output()` captures lines, populates the internal log store, and pushes updates via WebSocket for live UI consumption
4. **Polling**: Clients query `/api/crawler/status` for state (`running`, `idle`, etc.) and `/api/crawler/logs?limit=N` for recent entries
5. **Termination**: POST to `/api/crawler/stop` triggers the manager's shutdown sequence

All endpoints are automatically documented at `/docs` (Swagger UI) when the server is running.

## Programmatic Control Examples

### Starting a Crawl via Python

Submit crawl jobs using standard HTTP libraries. The payload must match the `CrawlerStartRequest` schema supported by [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py):

```python
import requests

url = "http://localhost:8080/api/crawler/start"
payload = {
    "platform": "dy",                # Douyin

    "login_type": "qrcode",
    "crawler_type": "search",
    "keywords": "travel vlog",
    "start_page": 1,
    "enable_comments": True,
    "enable_sub_comments": False,
    "save_option": "jsonl",
    "headless": True
}

resp = requests.post(url, json=payload)
print(resp.json())

```

### Checking Crawler Status

Query the current execution state to determine if the manager is `idle`, `running`, `stopping`, or in an `error` state:

```bash
curl http://localhost:8080/api/crawler/status

```

Response:

```json
{
  "status": "running",
  "platform": "dy",
  "crawler_type": "search",
  "started_at": "2026-07-31T12:34:56.789012",
  "error_message": null
}

```

### Retrieving Execution Logs

Fetch recent log entries from the in-memory FIFO queue. Use the `limit` parameter to control how many of the maximum 500 stored entries to return:

```python
import requests

r = requests.get("http://localhost:8080/api/crawler/logs?limit=50")
logs = r.json()["logs"]

for entry in logs:
    print(f"[{entry['timestamp']}] {entry['level'].upper()}: {entry['message']}")

```

### Stopping the Crawler

Initiate graceful shutdown. The manager sends `SIGTERM` and escalates to `SIGKILL` if the process does not terminate within 15 seconds:

```bash
curl -X POST http://localhost:8080/api/crawler/stop

```

### Discovering Supported Platforms

Query configuration endpoints to populate your application with valid platform options before initiating crawls:

```bash
curl http://localhost:8080/api/config/platforms

```

Response includes metadata for supported targets:

```json
{
  "platforms": [
    {"value":"xhs","label":"Xiaohongshu","icon":"book-open"},
    {"value":"dy","label":"Douyin","icon":"music"},
    {"value":"ks","label":"Kuaishou","icon":"video"},
    {"value":"bili","label":"Bilibili","icon":"tv"},
    {"value":"wb","label":"Weibo","icon":"message-circle"},
    {"value":"tieba","label":"Baidu Tieba","icon":"messages-square"},
    {"value":"zhihu","label":"Zhihu","icon":"help-circle"}
  ]
}

```

## Key Implementation Files

Understanding these source files helps when extending the API or debugging behavior:

- **[`main/api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/api/main.py)**: FastAPI application factory, CORS configuration, static file mounting, and health endpoints
- **[`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py)**: Route definitions for `/api/crawler/start`, `/stop`, `/status`, and `/logs`
- **[`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py)**: Subprocess orchestration, log line parsing, and signal management
- **[`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py)**: Pydantic models (`CrawlerStartRequest`, `CrawlerStatusResponse`, `LogEntry`) defining the API contract
- **[`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py)**: WebSocket handler for real-time log streaming to the frontend
- **[`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py)**: The core CLI crawler entry point executed by the manager service

## Summary

The MediaCrawler FastAPI WebUI provides robust programmatic control through a clean REST interface:

- **Process Isolation**: The `CrawlerManager` service spawns crawlers as separate subprocesses via `uv run`, keeping the API responsive
- **In-Memory Logging**: Real-time logs are buffered in a 500-entry FIFO queue accessible via HTTP or WebSocket
- **Graceful Shutdown**: Automatic `SIGTERM` escalation to `SIGKILL` ensures processes terminate cleanly or are forcibly stopped after 15 seconds
- **Schema Validation**: Pydantic models in [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py) ensure type safety and prevent malformed crawler configurations

## Frequently Asked Questions

### Does the MediaCrawler API support concurrent crawler instances?

No. The `CrawlerManager` is implemented as a singleton service that maintains state for a single subprocess. Attempting to start a new crawl while another is running will return an error status indicating the system is busy. Queue multiple jobs externally or wait for the current process to reach `idle` status before launching another.

### How long are logs preserved in the FastAPI WebUI?

Logs exist only in memory within the `CrawlerManager` instance and are lost when the API server restarts. The internal FIFO queue stores a maximum of 500 entries, discarding older lines as new ones arrive. For persistent logging, consume the WebSocket stream or poll the `/api/crawler/logs` endpoint frequently to export data to external storage.

### Can I use the API without the React frontend?

Yes. The FastAPI backend in [`main/api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/api/main.py) serves static files only when the built React assets exist in `api/webui/`. If these files are absent, the REST endpoints remain fully functional. Any HTTP client—Python `requests`, `curl`, Postman, or custom scripts—can interact with the crawler control endpoints directly.

### What authentication protects the crawler endpoints?

As implemented in the current source code, the MediaCrawler FastAPI WebUI does not include built-in authentication middleware. When deployed in production, place the service behind a reverse proxy (Nginx, Traefik) with basic auth, OAuth2, or IP whitelisting, or extend [`main/api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/api/main.py) to add FastAPI dependency injection for API key validation.