# MediaCrawler API: FastAPI HTTP & WebSocket Interface Explained

> Explore the MediaCrawler API with FastAPI. Control crawls, monitor status, and access data programmatically via its HTTP and WebSocket interfaces.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: api-reference
- Published: 2026-08-12

---

**Yes, MediaCrawler provides a complete FastAPI-based HTTP API with WebSocket support for controlling crawls, monitoring status, and accessing harvested data programmatically.**

MediaCrawler, the open-source media platform scraping framework by **NanmiCoder/MediaCrawler**, exposes a **RESTful control plane** that lets you orchestrate crawling tasks without modifying source code. The API is implemented in the `api/` directory and provides both synchronous endpoints for commands and asynchronous WebSocket streams for real-time telemetry.

## How the MediaCrawler API Is Structured

The API is organized into three router groups registered in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py). Each group handles a distinct concern:

| Router | Base Path | Core Function |
|--------|-----------|---------------|
| **Crawler** | `/crawler` | Start, stop, and monitor crawl jobs |
| **Data** | `/data` | Browse, preview, and download output files |
| **WebSocket** | `/ws` | Receive live log feeds and status changes |

All request and response payloads use **Pydantic models** defined in [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py), ensuring strict validation and automatic OpenAPI documentation generation.

## Crawler Control Endpoints (`/api/crawler/`)

These endpoints manage the lifecycle of a crawling process through the `crawler_manager` service.

### Start a Crawl: `POST /api/crawler/start`

The `start_crawler` handler in [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py) (lines 27-38) accepts a `CrawlerStartRequest` and launches the crawl as a subprocess.

**Required payload structure:**

```json
{
  "platform": "zhihu",
  "login_type": "qrcode",
  "crawler_type": "search",
  "keywords": "AI",
  "start_page": 1,
  "enable_comments": true,
  "save_option": "jsonl"
}

```

**Python client example:**

```python
import httpx

payload = {
    "platform": "zhihu",
    "login_type": "qrcode",
    "crawler_type": "search",
    "keywords": "人工智能",
    "start_page": 1,
    "save_option": "jsonl"
}
resp = httpx.post("http://localhost:8000/api/crawler/start", json=payload)
print(resp.json())

# {"status": "ok", "message": "Crawler started successfully"}

```

Returns **400 Bad Request** if a crawl is already active.

### Stop a Crawl: `POST /api/crawler/stop`

Calls `crawler_manager.stop()` to terminate the running subprocess. Source: [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py) lines 40-50. Returns **400** if no crawl is running.

### Check Status: `GET /api/crawler/status`

Returns a `CrawlerStatusResponse` with the current state: `idle`, `running`, `stopping`, or `error`. Source: [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py) lines 53-56.

**Polling example:**

```python
import time, httpx

while True:
    r = httpx.get("http://localhost:8000/api/crawler/status")
    status = r.json()["status"]
    if status == "idle":
        break
    time.sleep(5)

```

### Retrieve Logs: `GET /api/crawler/logs?limit=100`

Returns recent `LogEntry` records from the internal log buffer. Source: [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py) lines 59-63. The `limit` parameter caps the result set (default varies by implementation).

## Data Access Endpoints (`/api/data/`)

These endpoints interact with files written to the repository's `data/` directory. Implemented in [`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py).

### List Files: `GET /api/data/files`

Walks the data directory and returns metadata for each file: name, size, modification time, and estimated record count. Supports filtering by `platform` and `file_type` query parameters. Source: lines 61-95.

```python
import httpx

r = httpx.get("http://localhost:8000/api/data/files",
              params={"platform": "zhihu", "file_type": "jsonl"})
for f in r.json()["files"]:
    print(f["name"], f["size"], "records:", f["record_count"])

```

### Preview File Contents: `GET /api/data/files/{file_path}`

Returns a slice of a file's contents with `?preview=true&limit=100`. Supports JSONL, CSV, and Excel formats. Source: lines 98-156. Returns both the preview rows and total record count.

### Download Raw Files: `GET /api/data/download/{file_path}`

Streams the complete file for client-side download. Source: lines 166-187.

```python
import httpx

file_path = "zhihu/2024-09-01_posts.jsonl"
url = f"http://localhost:8000/api/data/download/{file_path}"

with httpx.stream("GET", url) as stream:
    with open("posts.jsonl", "wb") as out:
        for chunk in stream.iter_bytes():
            out.write(chunk)

```

### Aggregate Statistics: `GET /api/data/stats`

Summarizes total files, total storage size, and breakdowns by platform and file type. Source: lines 90-130.

## WebSocket Real-Time Streams (`/api/ws/`)

For automation pipelines that need push-based updates, MediaCrawler's API provides WebSocket endpoints in [`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py).

### Live Logs: `ws://localhost:8000/api/ws/logs`

Pushes each log entry as a JSON message immediately when generated. Source: lines 89-102.

```python
import websockets, asyncio, json

async def watch_logs():
    async with websockets.connect("ws://localhost:8000/api/ws/logs") as ws:
        async for msg in ws:
            entry = json.loads(msg)
            print(f"[{entry['level']}] {entry['message']}")

asyncio.run(watch_logs())

```

### Live Status: `ws://localhost:8000/api/ws/status`

Streams status change events (`idle` → `running` → `stopping` → `idle`) as they occur. Source: lines 119-132.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) | FastAPI app instantiation and router registration |
| [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py) | Crawler control endpoints (start/stop/status/logs) |
| [`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py) | Data file browsing, preview, and download |
| [`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py) | WebSocket handlers for real-time streams |
| [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py) | Pydantic models: `CrawlerStartRequest`, `CrawlerStatusResponse`, `LogEntry`, etc. |
| [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py) | Core service spawning crawl subprocesses and managing state |

## Summary

- MediaCrawler exposes a **production-ready FastAPI interface** with REST endpoints and WebSocket support
- **Three router groups** cover crawler control (`/crawler`), data access (`/data`), and real-time telemetry (`/ws`)
- All payloads use **Pydantic validation** with models defined in [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py)
- The `crawler_manager` service handles subprocess lifecycle, making the API **stateful but thread-safe**
- File operations support **filtering, pagination, preview, and streaming download**

## Frequently Asked Questions

### What port does the MediaCrawler API use?

The default port is **8000**, consistent with standard Uvicorn/FastAPI deployments. Configure this via the Uvicorn command or environment variable when starting [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py).

### Can I run multiple crawls simultaneously through the API?

**No.** The `crawler_manager` enforces a single active crawl. Attempting to `POST /crawler/start` while another crawl runs returns HTTP 400 with an error message. Queue multiple jobs externally or run separate API instances.

### Does the API require authentication?

The open-source implementation in **NanmiCoder/MediaCrawler** does not include built-in authentication. Deploy behind a reverse proxy (nginx, Traefik) with API key or JWT validation for production use.

### Which platforms can I target through the API?

Any platform supported by MediaCrawler's base crawler architecture: **Zhihu**, **Xiaohongshu**, **Bilibili**, and others. The `platform` field in `CrawlerStartRequest` maps directly to crawler implementations inheriting from [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py).