# MediaCrawler API: FastAPI HTTP Interface for Orchestrating Crawl Jobs

> Explore the MediaCrawler API, a FastAPI HTTP interface. Start, stop, and monitor crawl jobs with RESTful endpoints & WebSocket streams. Manage downloaded files efficiently.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: api-reference
- Published: 2026-07-29

---

**The MediaCrawler API is a FastAPI-based HTTP interface that provides RESTful endpoints and WebSocket streams to start and stop crawl jobs, monitor real-time status, and manage downloaded data files.**

The MediaCrawler API transforms the `NanmiCoder/MediaCrawler` repository from a command-line tool into a programmable service. Built on FastAPI, it exposes three distinct router groups that handle crawler lifecycle management, data file operations, and real-time event streaming, all validated through Pydantic models defined in [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py).

## MediaCrawler API Architecture

The API is organized into three core router groups, each registered in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) and handling specific operational domains:

- **Crawler Router** (`/api/crawler`) – Controls execution lifecycle and retrieves runtime information
- **Data Router** (`/api/data`) – Manages file listing, preview, download, and aggregate statistics
- **WebSocket Router** (`/api/ws`) – Pushes real-time logs and status updates without polling

All requests and responses utilize strict Pydantic validation, ensuring type safety and automatic OpenAPI documentation generation.

## Crawler Control Endpoints

The crawler control interface, implemented in [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py), provides the primary mechanism for initiating and managing crawl operations.

### Start a Crawl

To begin crawling, send a `POST` request to `/api/crawler/start` with a JSON payload matching the `CrawlerStartRequest` schema:

```json
{
  "platform": "zhihu",
  "login_type": "qrcode",
  "crawler_type": "search",
  "keywords": "AI",
  "start_page": 1,
  "enable_comments": true,
  "save_option": "jsonl"
}

```

The `start_crawler` handler forwards this payload to `crawler_manager.start()` and returns a success confirmation. If a crawl is already active, the endpoint returns a 400 error. Source: [[`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L27-L38).

### Stop a Crawl

Send a `POST` request to `/api/crawler/stop` to terminate the current operation. This invokes `crawler_manager.stop()` and returns an error if no process is currently running. Source: [[`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L40-L50).

### Check Runtime Status

Query `GET /api/crawler/status` to receive a `CrawlerStatusResponse` indicating the current state. Valid states include `idle`, `running`, `stopping`, or `error`, optionally accompanied by platform-specific metadata. Source: [[`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L53-L56).

### Retrieve Log Entries

Access recent crawl logs via `GET /api/crawler/logs?limit=100`. This returns an array of `LogEntry` records, with the `limit` parameter controlling the number of entries returned. Source: [[`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L59-L63).

## Data Management Endpoints

The data router in [`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py) provides filesystem access to crawler output without requiring direct server access.

### List Generated Files

Query `GET /api/data/files` with optional filters to inspect the `data/` directory:

```python
params = {
    "platform": "zhihu",
    "file_type": "jsonl"
}

# Returns metadata including size, modification time, and record count

```

The endpoint walks the data directory and returns file statistics including record counts. Source: [[`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L61-L95).

### Preview File Contents

Access `GET /api/data/files/{file_path}?preview=true&limit=100` to retrieve a slice of a specific file's contents. The endpoint supports JSON, CSV, and Excel formats, returning both the preview data and total record count. Source: [[`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L98-L156).

### Download Raw Files

Stream complete files using `GET /api/data/download/{file_path}`. This endpoint streams the raw file bytes, suitable for client-side downloads of large datasets. Source: [[`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L166-L187).

### Aggregate Statistics

Retrieve summaries via `GET /api/data/stats`, which returns total file counts, cumulative size, and breakdowns by platform and file type. Source: [[`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L90-L130).

## Real-Time Monitoring with WebSocket

For applications requiring live updates, the MediaCrawler API offers WebSocket endpoints defined in [`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py).

### Live Log Streaming

Connect to `ws://<host>/api/ws/logs` to receive log entries as they are generated. This pushes `LogEntry` objects in real-time, eliminating the need for polling during long-running crawls. Source: [[`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py#L89-L102).

### Status Change Notifications

Subscribe to `ws://<host>/api/ws/status` to receive immediate notifications when the crawler transitions between states (`idle`, `running`, `stopping`, `error`). Source: [[`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py#L119-L132).

## Practical Code Examples

### Start a Crawler with Python

```python
import httpx

payload = {
    "platform": "zhihu",
    "login_type": "qrcode",
    "crawler_type": "search",
    "keywords": "人工智能",
    "start_page": 1,
    "save_option": "jsonl"
}
resp = httpx.post("http://localhost:8000/api/crawler/start", json=payload)
print(resp.json())  # {"status":"ok","message":"Crawler started successfully"}

```

### Poll Until Completion

```python
import time
import httpx

while True:
    r = httpx.get("http://localhost:8000/api/crawler/status")
    status = r.json()["status"]
    print("Current:", status)
    if status == "idle":
        break
    time.sleep(5)

```

### List and Filter Data Files

```python
import httpx

r = httpx.get(
    "http://localhost:8000/api/data/files",
    params={"platform": "zhihu", "file_type": "jsonl"}
)
for f in r.json()["files"]:
    print(f["name"], f["size"], "records:", f["record_count"])

```

### Stream File Download

```python
import httpx

file_path = "zhihu/2024-09-01_posts.jsonl"
url = f"http://localhost:8000/api/data/download/{file_path}"

with httpx.stream("GET", url) as stream:
    with open("posts.jsonl", "wb") as out:
        for chunk in stream.iter_bytes():
            out.write(chunk)

```

### Monitor Logs via WebSocket

```python
import asyncio
import json
import websockets

async def watch_logs():
    async with websockets.connect("ws://localhost:8000/api/ws/logs") as ws:
        async for msg in ws:
            entry = json.loads(msg)
            print(f"[{entry['level']}] {entry['message']}")

asyncio.run(watch_logs())

```

## Summary

- The **MediaCrawler API** provides a **FastAPI-based control plane** over the MediaCrawler tool, exposing endpoints in [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py), [`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py), and [`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py).
- **Crawler control** allows starting and stopping jobs via `/api/crawler/start` and `/api/crawler/stop`, with status monitoring through `/api/crawler/status`.
- **Data access** endpoints in `/api/data` enable listing, previewing, and downloading files from the `data/` directory without server shell access.
- **WebSocket streams** at `/api/ws/logs` and `/api/ws/status` push real-time telemetry for monitoring long-running operations.
- All interactions use **Pydantic models** defined in [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py) for automatic validation and documentation.

## Frequently Asked Questions

### What is the base URL path for the MediaCrawler API?

All endpoints are prefixed with `/api`. For example, crawler control endpoints are accessed at `/api/crawler/start`, `/api/crawler/stop`, and `/api/crawler/status`, while data endpoints use `/api/data/files`. The WebSocket endpoints follow the same pattern at `/api/ws/logs` and `/api/ws/status`.

### How can I check if a crawl is currently running?

Send a `GET` request to `/api/crawler/status`. The endpoint returns a JSON object containing a `status` field with values of `idle`, `running`, `stopping`, or `error`. When `status` is `idle`, no crawl is active and you may safely initiate a new job via the start endpoint.

### Can I download crawled data directly through the API?

Yes. Use `GET /api/data/download/{file_path}` where `file_path` is the relative path from the data directory (for example, `zhihu/posts.jsonl`). This endpoint streams the raw file bytes, allowing you to save large datasets directly without accessing the server filesystem.

### What real-time monitoring options does the MediaCrawler API support?

The API provides two WebSocket endpoints for live monitoring: `/api/ws/logs` streams log entries as they are generated during crawling, and `/api/ws/status` pushes state changes immediately when the crawler transitions between idle, running, stopping, or error states. These connections require WebSocket client support and remain open until explicitly closed.