MediaCrawler API: FastAPI HTTP & WebSocket Interface Explained

Yes, MediaCrawler provides a complete FastAPI-based HTTP API with WebSocket support for controlling crawls, monitoring status, and accessing harvested data programmatically.

MediaCrawler, the open-source media platform scraping framework by NanmiCoder/MediaCrawler, exposes a RESTful control plane that lets you orchestrate crawling tasks without modifying source code. The API is implemented in the api/ directory and provides both synchronous endpoints for commands and asynchronous WebSocket streams for real-time telemetry.

How the MediaCrawler API Is Structured

The API is organized into three router groups registered in api/main.py. Each group handles a distinct concern:

Router Base Path Core Function
Crawler /crawler Start, stop, and monitor crawl jobs
Data /data Browse, preview, and download output files
WebSocket /ws Receive live log feeds and status changes

All request and response payloads use Pydantic models defined in api/schemas/crawler.py, ensuring strict validation and automatic OpenAPI documentation generation.

Crawler Control Endpoints (/api/crawler/)

These endpoints manage the lifecycle of a crawling process through the crawler_manager service.

Start a Crawl: POST /api/crawler/start

The start_crawler handler in api/routers/crawler.py (lines 27-38) accepts a CrawlerStartRequest and launches the crawl as a subprocess.

Required payload structure:

{
  "platform": "zhihu",
  "login_type": "qrcode",
  "crawler_type": "search",
  "keywords": "AI",
  "start_page": 1,
  "enable_comments": true,
  "save_option": "jsonl"
}

Python client example:

import httpx

payload = {
    "platform": "zhihu",
    "login_type": "qrcode",
    "crawler_type": "search",
    "keywords": "人工智能",
    "start_page": 1,
    "save_option": "jsonl"
}
resp = httpx.post("http://localhost:8000/api/crawler/start", json=payload)
print(resp.json())

# {"status": "ok", "message": "Crawler started successfully"}

Returns 400 Bad Request if a crawl is already active.

Stop a Crawl: POST /api/crawler/stop

Calls crawler_manager.stop() to terminate the running subprocess. Source: api/routers/crawler.py lines 40-50. Returns 400 if no crawl is running.

Check Status: GET /api/crawler/status

Returns a CrawlerStatusResponse with the current state: idle, running, stopping, or error. Source: api/routers/crawler.py lines 53-56.

Polling example:

import time, httpx

while True:
    r = httpx.get("http://localhost:8000/api/crawler/status")
    status = r.json()["status"]
    if status == "idle":
        break
    time.sleep(5)

Retrieve Logs: GET /api/crawler/logs?limit=100

Returns recent LogEntry records from the internal log buffer. Source: api/routers/crawler.py lines 59-63. The limit parameter caps the result set (default varies by implementation).

Data Access Endpoints (/api/data/)

These endpoints interact with files written to the repository's data/ directory. Implemented in api/routers/data.py.

List Files: GET /api/data/files

Walks the data directory and returns metadata for each file: name, size, modification time, and estimated record count. Supports filtering by platform and file_type query parameters. Source: lines 61-95.

import httpx

r = httpx.get("http://localhost:8000/api/data/files",
              params={"platform": "zhihu", "file_type": "jsonl"})
for f in r.json()["files"]:
    print(f["name"], f["size"], "records:", f["record_count"])

Preview File Contents: GET /api/data/files/{file_path}

Returns a slice of a file's contents with ?preview=true&limit=100. Supports JSONL, CSV, and Excel formats. Source: lines 98-156. Returns both the preview rows and total record count.

Download Raw Files: GET /api/data/download/{file_path}

Streams the complete file for client-side download. Source: lines 166-187.

import httpx

file_path = "zhihu/2024-09-01_posts.jsonl"
url = f"http://localhost:8000/api/data/download/{file_path}"

with httpx.stream("GET", url) as stream:
    with open("posts.jsonl", "wb") as out:
        for chunk in stream.iter_bytes():
            out.write(chunk)

Aggregate Statistics: GET /api/data/stats

Summarizes total files, total storage size, and breakdowns by platform and file type. Source: lines 90-130.

WebSocket Real-Time Streams (/api/ws/)

For automation pipelines that need push-based updates, MediaCrawler's API provides WebSocket endpoints in api/routers/websocket.py.

Live Logs: ws://localhost:8000/api/ws/logs

Pushes each log entry as a JSON message immediately when generated. Source: lines 89-102.

import websockets, asyncio, json

async def watch_logs():
    async with websockets.connect("ws://localhost:8000/api/ws/logs") as ws:
        async for msg in ws:
            entry = json.loads(msg)
            print(f"[{entry['level']}] {entry['message']}")

asyncio.run(watch_logs())

Live Status: ws://localhost:8000/api/ws/status

Streams status change events (idle → running → stopping → idle) as they occur. Source: lines 119-132.

Key Implementation Files

File Purpose
api/main.py FastAPI app instantiation and router registration
api/routers/crawler.py Crawler control endpoints (start/stop/status/logs)
api/routers/data.py Data file browsing, preview, and download
api/routers/websocket.py WebSocket handlers for real-time streams
api/schemas/crawler.py Pydantic models: CrawlerStartRequest, CrawlerStatusResponse, LogEntry, etc.
api/services/crawler_manager.py Core service spawning crawl subprocesses and managing state

Summary

  • MediaCrawler exposes a production-ready FastAPI interface with REST endpoints and WebSocket support
  • Three router groups cover crawler control (/crawler), data access (/data), and real-time telemetry (/ws)
  • All payloads use Pydantic validation with models defined in api/schemas/crawler.py
  • The crawler_manager service handles subprocess lifecycle, making the API stateful but thread-safe
  • File operations support filtering, pagination, preview, and streaming download

Frequently Asked Questions

What port does the MediaCrawler API use?

The default port is 8000, consistent with standard Uvicorn/FastAPI deployments. Configure this via the Uvicorn command or environment variable when starting api/main.py.

Can I run multiple crawls simultaneously through the API?

No. The crawler_manager enforces a single active crawl. Attempting to POST /crawler/start while another crawl runs returns HTTP 400 with an error message. Queue multiple jobs externally or run separate API instances.

Does the API require authentication?

The open-source implementation in NanmiCoder/MediaCrawler does not include built-in authentication. Deploy behind a reverse proxy (nginx, Traefik) with API key or JWT validation for production use.

Which platforms can I target through the API?

Any platform supported by MediaCrawler's base crawler architecture: Zhihu, Xiaohongshu, Bilibili, and others. The platform field in CrawlerStartRequest maps directly to crawler implementations inheriting from base/base_crawler.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →