MediaCrawler API: FastAPI HTTP Interface for Orchestrating Crawl Jobs

The MediaCrawler API is a FastAPI-based HTTP interface that provides RESTful endpoints and WebSocket streams to start and stop crawl jobs, monitor real-time status, and manage downloaded data files.

The MediaCrawler API transforms the NanmiCoder/MediaCrawler repository from a command-line tool into a programmable service. Built on FastAPI, it exposes three distinct router groups that handle crawler lifecycle management, data file operations, and real-time event streaming, all validated through Pydantic models defined in api/schemas/crawler.py.

MediaCrawler API Architecture

The API is organized into three core router groups, each registered in api/main.py and handling specific operational domains:

  • Crawler Router (/api/crawler) – Controls execution lifecycle and retrieves runtime information
  • Data Router (/api/data) – Manages file listing, preview, download, and aggregate statistics
  • WebSocket Router (/api/ws) – Pushes real-time logs and status updates without polling

All requests and responses utilize strict Pydantic validation, ensuring type safety and automatic OpenAPI documentation generation.

Crawler Control Endpoints

The crawler control interface, implemented in api/routers/crawler.py, provides the primary mechanism for initiating and managing crawl operations.

Start a Crawl

To begin crawling, send a POST request to /api/crawler/start with a JSON payload matching the CrawlerStartRequest schema:

{
  "platform": "zhihu",
  "login_type": "qrcode",
  "crawler_type": "search",
  "keywords": "AI",
  "start_page": 1,
  "enable_comments": true,
  "save_option": "jsonl"
}

The start_crawler handler forwards this payload to crawler_manager.start() and returns a success confirmation. If a crawl is already active, the endpoint returns a 400 error. Source: [api/routers/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L27-L38).

Stop a Crawl

Send a POST request to /api/crawler/stop to terminate the current operation. This invokes crawler_manager.stop() and returns an error if no process is currently running. Source: [api/routers/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L40-L50).

Check Runtime Status

Query GET /api/crawler/status to receive a CrawlerStatusResponse indicating the current state. Valid states include idle, running, stopping, or error, optionally accompanied by platform-specific metadata. Source: [api/routers/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L53-L56).

Retrieve Log Entries

Access recent crawl logs via GET /api/crawler/logs?limit=100. This returns an array of LogEntry records, with the limit parameter controlling the number of entries returned. Source: [api/routers/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L59-L63).

Data Management Endpoints

The data router in api/routers/data.py provides filesystem access to crawler output without requiring direct server access.

List Generated Files

Query GET /api/data/files with optional filters to inspect the data/ directory:

params = {
    "platform": "zhihu",
    "file_type": "jsonl"
}

# Returns metadata including size, modification time, and record count

The endpoint walks the data directory and returns file statistics including record counts. Source: [api/routers/data.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L61-L95).

Preview File Contents

Access GET /api/data/files/{file_path}?preview=true&limit=100 to retrieve a slice of a specific file's contents. The endpoint supports JSON, CSV, and Excel formats, returning both the preview data and total record count. Source: [api/routers/data.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L98-L156).

Download Raw Files

Stream complete files using GET /api/data/download/{file_path}. This endpoint streams the raw file bytes, suitable for client-side downloads of large datasets. Source: [api/routers/data.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L166-L187).

Aggregate Statistics

Retrieve summaries via GET /api/data/stats, which returns total file counts, cumulative size, and breakdowns by platform and file type. Source: [api/routers/data.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L90-L130).

Real-Time Monitoring with WebSocket

For applications requiring live updates, the MediaCrawler API offers WebSocket endpoints defined in api/routers/websocket.py.

Live Log Streaming

Connect to ws://<host>/api/ws/logs to receive log entries as they are generated. This pushes LogEntry objects in real-time, eliminating the need for polling during long-running crawls. Source: [api/routers/websocket.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py#L89-L102).

Status Change Notifications

Subscribe to ws://<host>/api/ws/status to receive immediate notifications when the crawler transitions between states (idle, running, stopping, error). Source: [api/routers/websocket.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py#L119-L132).

Practical Code Examples

Start a Crawler with Python

import httpx

payload = {
    "platform": "zhihu",
    "login_type": "qrcode",
    "crawler_type": "search",
    "keywords": "人工智能",
    "start_page": 1,
    "save_option": "jsonl"
}
resp = httpx.post("http://localhost:8000/api/crawler/start", json=payload)
print(resp.json())  # {"status":"ok","message":"Crawler started successfully"}

Poll Until Completion

import time
import httpx

while True:
    r = httpx.get("http://localhost:8000/api/crawler/status")
    status = r.json()["status"]
    print("Current:", status)
    if status == "idle":
        break
    time.sleep(5)

List and Filter Data Files

import httpx

r = httpx.get(
    "http://localhost:8000/api/data/files",
    params={"platform": "zhihu", "file_type": "jsonl"}
)
for f in r.json()["files"]:
    print(f["name"], f["size"], "records:", f["record_count"])

Stream File Download

import httpx

file_path = "zhihu/2024-09-01_posts.jsonl"
url = f"http://localhost:8000/api/data/download/{file_path}"

with httpx.stream("GET", url) as stream:
    with open("posts.jsonl", "wb") as out:
        for chunk in stream.iter_bytes():
            out.write(chunk)

Monitor Logs via WebSocket

import asyncio
import json
import websockets

async def watch_logs():
    async with websockets.connect("ws://localhost:8000/api/ws/logs") as ws:
        async for msg in ws:
            entry = json.loads(msg)
            print(f"[{entry['level']}] {entry['message']}")

asyncio.run(watch_logs())

Summary

  • The MediaCrawler API provides a FastAPI-based control plane over the MediaCrawler tool, exposing endpoints in api/routers/crawler.py, api/routers/data.py, and api/routers/websocket.py.
  • Crawler control allows starting and stopping jobs via /api/crawler/start and /api/crawler/stop, with status monitoring through /api/crawler/status.
  • Data access endpoints in /api/data enable listing, previewing, and downloading files from the data/ directory without server shell access.
  • WebSocket streams at /api/ws/logs and /api/ws/status push real-time telemetry for monitoring long-running operations.
  • All interactions use Pydantic models defined in api/schemas/crawler.py for automatic validation and documentation.

Frequently Asked Questions

What is the base URL path for the MediaCrawler API?

All endpoints are prefixed with /api. For example, crawler control endpoints are accessed at /api/crawler/start, /api/crawler/stop, and /api/crawler/status, while data endpoints use /api/data/files. The WebSocket endpoints follow the same pattern at /api/ws/logs and /api/ws/status.

How can I check if a crawl is currently running?

Send a GET request to /api/crawler/status. The endpoint returns a JSON object containing a status field with values of idle, running, stopping, or error. When status is idle, no crawl is active and you may safely initiate a new job via the start endpoint.

Can I download crawled data directly through the API?

Yes. Use GET /api/data/download/{file_path} where file_path is the relative path from the data directory (for example, zhihu/posts.jsonl). This endpoint streams the raw file bytes, allowing you to save large datasets directly without accessing the server filesystem.

What real-time monitoring options does the MediaCrawler API support?

The API provides two WebSocket endpoints for live monitoring: /api/ws/logs streams log entries as they are generated during crawling, and /api/ws/status pushes state changes immediately when the crawler transitions between idle, running, stopping, or error states. These connections require WebSocket client support and remain open until explicitly closed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →