MediaCrawler API: FastAPI HTTP Interface for Orchestrating Crawl Jobs
The MediaCrawler API is a FastAPI-based HTTP interface that provides RESTful endpoints and WebSocket streams to start and stop crawl jobs, monitor real-time status, and manage downloaded data files.
The MediaCrawler API transforms the NanmiCoder/MediaCrawler repository from a command-line tool into a programmable service. Built on FastAPI, it exposes three distinct router groups that handle crawler lifecycle management, data file operations, and real-time event streaming, all validated through Pydantic models defined in api/schemas/crawler.py.
MediaCrawler API Architecture
The API is organized into three core router groups, each registered in api/main.py and handling specific operational domains:
- Crawler Router (
/api/crawler) – Controls execution lifecycle and retrieves runtime information - Data Router (
/api/data) – Manages file listing, preview, download, and aggregate statistics - WebSocket Router (
/api/ws) – Pushes real-time logs and status updates without polling
All requests and responses utilize strict Pydantic validation, ensuring type safety and automatic OpenAPI documentation generation.
Crawler Control Endpoints
The crawler control interface, implemented in api/routers/crawler.py, provides the primary mechanism for initiating and managing crawl operations.
Start a Crawl
To begin crawling, send a POST request to /api/crawler/start with a JSON payload matching the CrawlerStartRequest schema:
{
"platform": "zhihu",
"login_type": "qrcode",
"crawler_type": "search",
"keywords": "AI",
"start_page": 1,
"enable_comments": true,
"save_option": "jsonl"
}
The start_crawler handler forwards this payload to crawler_manager.start() and returns a success confirmation. If a crawl is already active, the endpoint returns a 400 error. Source: [api/routers/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L27-L38).
Stop a Crawl
Send a POST request to /api/crawler/stop to terminate the current operation. This invokes crawler_manager.stop() and returns an error if no process is currently running. Source: [api/routers/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L40-L50).
Check Runtime Status
Query GET /api/crawler/status to receive a CrawlerStatusResponse indicating the current state. Valid states include idle, running, stopping, or error, optionally accompanied by platform-specific metadata. Source: [api/routers/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L53-L56).
Retrieve Log Entries
Access recent crawl logs via GET /api/crawler/logs?limit=100. This returns an array of LogEntry records, with the limit parameter controlling the number of entries returned. Source: [api/routers/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py#L59-L63).
Data Management Endpoints
The data router in api/routers/data.py provides filesystem access to crawler output without requiring direct server access.
List Generated Files
Query GET /api/data/files with optional filters to inspect the data/ directory:
params = {
"platform": "zhihu",
"file_type": "jsonl"
}
# Returns metadata including size, modification time, and record count
The endpoint walks the data directory and returns file statistics including record counts. Source: [api/routers/data.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L61-L95).
Preview File Contents
Access GET /api/data/files/{file_path}?preview=true&limit=100 to retrieve a slice of a specific file's contents. The endpoint supports JSON, CSV, and Excel formats, returning both the preview data and total record count. Source: [api/routers/data.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L98-L156).
Download Raw Files
Stream complete files using GET /api/data/download/{file_path}. This endpoint streams the raw file bytes, suitable for client-side downloads of large datasets. Source: [api/routers/data.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L166-L187).
Aggregate Statistics
Retrieve summaries via GET /api/data/stats, which returns total file counts, cumulative size, and breakdowns by platform and file type. Source: [api/routers/data.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py#L90-L130).
Real-Time Monitoring with WebSocket
For applications requiring live updates, the MediaCrawler API offers WebSocket endpoints defined in api/routers/websocket.py.
Live Log Streaming
Connect to ws://<host>/api/ws/logs to receive log entries as they are generated. This pushes LogEntry objects in real-time, eliminating the need for polling during long-running crawls. Source: [api/routers/websocket.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py#L89-L102).
Status Change Notifications
Subscribe to ws://<host>/api/ws/status to receive immediate notifications when the crawler transitions between states (idle, running, stopping, error). Source: [api/routers/websocket.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py#L119-L132).
Practical Code Examples
Start a Crawler with Python
import httpx
payload = {
"platform": "zhihu",
"login_type": "qrcode",
"crawler_type": "search",
"keywords": "人工智能",
"start_page": 1,
"save_option": "jsonl"
}
resp = httpx.post("http://localhost:8000/api/crawler/start", json=payload)
print(resp.json()) # {"status":"ok","message":"Crawler started successfully"}
Poll Until Completion
import time
import httpx
while True:
r = httpx.get("http://localhost:8000/api/crawler/status")
status = r.json()["status"]
print("Current:", status)
if status == "idle":
break
time.sleep(5)
List and Filter Data Files
import httpx
r = httpx.get(
"http://localhost:8000/api/data/files",
params={"platform": "zhihu", "file_type": "jsonl"}
)
for f in r.json()["files"]:
print(f["name"], f["size"], "records:", f["record_count"])
Stream File Download
import httpx
file_path = "zhihu/2024-09-01_posts.jsonl"
url = f"http://localhost:8000/api/data/download/{file_path}"
with httpx.stream("GET", url) as stream:
with open("posts.jsonl", "wb") as out:
for chunk in stream.iter_bytes():
out.write(chunk)
Monitor Logs via WebSocket
import asyncio
import json
import websockets
async def watch_logs():
async with websockets.connect("ws://localhost:8000/api/ws/logs") as ws:
async for msg in ws:
entry = json.loads(msg)
print(f"[{entry['level']}] {entry['message']}")
asyncio.run(watch_logs())
Summary
- The MediaCrawler API provides a FastAPI-based control plane over the MediaCrawler tool, exposing endpoints in
api/routers/crawler.py,api/routers/data.py, andapi/routers/websocket.py. - Crawler control allows starting and stopping jobs via
/api/crawler/startand/api/crawler/stop, with status monitoring through/api/crawler/status. - Data access endpoints in
/api/dataenable listing, previewing, and downloading files from thedata/directory without server shell access. - WebSocket streams at
/api/ws/logsand/api/ws/statuspush real-time telemetry for monitoring long-running operations. - All interactions use Pydantic models defined in
api/schemas/crawler.pyfor automatic validation and documentation.
Frequently Asked Questions
What is the base URL path for the MediaCrawler API?
All endpoints are prefixed with /api. For example, crawler control endpoints are accessed at /api/crawler/start, /api/crawler/stop, and /api/crawler/status, while data endpoints use /api/data/files. The WebSocket endpoints follow the same pattern at /api/ws/logs and /api/ws/status.
How can I check if a crawl is currently running?
Send a GET request to /api/crawler/status. The endpoint returns a JSON object containing a status field with values of idle, running, stopping, or error. When status is idle, no crawl is active and you may safely initiate a new job via the start endpoint.
Can I download crawled data directly through the API?
Yes. Use GET /api/data/download/{file_path} where file_path is the relative path from the data directory (for example, zhihu/posts.jsonl). This endpoint streams the raw file bytes, allowing you to save large datasets directly without accessing the server filesystem.
What real-time monitoring options does the MediaCrawler API support?
The API provides two WebSocket endpoints for live monitoring: /api/ws/logs streams log entries as they are generated during crawling, and /api/ws/status pushes state changes immediately when the crawler transitions between idle, running, stopping, or error states. These connections require WebSocket client support and remain open until explicitly closed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →