How to Integrate MediaCrawler with Other Tools: A Complete Integration Guide
You can integrate MediaCrawler via direct Python imports, REST API calls, or the built-in WebUI—each suited to different automation and embedding scenarios.
MediaCrawler is a modular Python library designed for scraping content from major Chinese social platforms. Whether you're building a data pipeline, adding scraping capabilities to an existing service, or creating a custom front-end, this guide explains every MediaCrawler integration method based on the actual source code implementation in NanmiCoder/MediaCrawler.
Python Library Integration
The most direct way to integrate MediaCrawler with other tools is to import and use its core classes as a Python library. The CLI entry point in main.py exposes the same objects you can use programmatically.
Basic Programmatic Usage
# my_integration.py
from main import CrawlerFactory, config
from var import crawler_type_var
# Configure the crawler at runtime
config.PLATFORM = "dy" # Douyin
config.SAVE_DATA_OPTION = "json" # Storage format
config.ENABLE_GET_COMMENTS = True
# Create and run crawler instance
crawler = CrawlerFactory.create_crawler(platform=config.PLATFORM)
import asyncio
asyncio.run(crawler.start())
Key integration points:
CrawlerFactory(main.py) – Maps platform strings (xhs,dy,ks,bili,wb,tieba,zhihu) to concrete crawler classesconfig(config/base_config.py) – Central configuration object with all runtime optionsAbstractCrawlersubclasses (media_platform/*.py) – Platform-specific implementations you can extend
This approach works well when embedding MediaCrawler into Django, Flask, Celery, or custom async services.
FastAPI HTTP API Integration
For language-agnostic integration or remote control scenarios, run MediaCrawler as a standalone HTTP service. The FastAPI application in api/main.py exposes full crawler control via REST endpoints.
Starting the API Server
uv run uvicorn api.main:app --port 8080 --reload
Core REST Endpoints
| Method | Endpoint | Purpose |
|---|---|---|
POST |
/api/crawler/start |
Launch crawl with JSON configuration |
POST |
/api/crawler/stop |
Gracefully stop running crawler |
GET |
/api/crawler/status |
Check crawler state and PID |
GET |
/api/crawler/logs?limit=50 |
Retrieve recent log entries |
GET |
/api/config/platforms |
List supported platforms |
GET |
/api/config/options |
Get login types, modes, storage options |
Example API Call
import requests
payload = {
"platform": "xhs",
"login_type": "qrcode",
"crawler_type": "search",
"keyword": "机器学习",
"max_pages": 5
}
response = requests.post(
"http://localhost:8080/api/crawler/start",
json=payload
)
print(response.json()) # Returns job status
The request schema is defined in api/schemas/crawler.py (CrawlerStartRequest), with fields for platform selection, authentication method, crawl mode, and platform-specific parameters.
Router implementations:
api/routers/crawler.py– Start/stop/status/log endpointsapi/routers/data.py– Data export utilitiesapi/routers/websocket.py– Real-time progress streaming (optional)
WebUI Integration and Embedding
MediaCrawler includes a Vue 3 + Vite front-end in the webui/ directory that communicates with the FastAPI backend.
Development Setup
# Terminal 1: API server
uv run uvicorn api.main:app --port 8080 --reload
# Terminal 2: WebUI
cd webui
npm install
npm run dev # http://localhost:5173
Production Deployment
After building (npm run build), FastAPI serves the static files at / and /static. This enables two WebUI integration patterns:
- Standalone application – Users access the crawler through the provided interface
- Embedded iframe – Integrate the visual crawler control into existing dashboards or admin panels
Since the WebUI uses the same REST API documented above, you can also build custom front-ends that call these endpoints directly.
Extending MediaCrawler for Custom Platforms
To integrate a new social platform into MediaCrawler's architecture:
- Create crawler class – Subclass
AbstractCrawlerinmedia_platform/your_platform.py - Register in factory – Add mapping in
CrawlerFactory.CRAWLERS(main.py) - Add configuration – Create
config/your_platform_config.py(optional) - Expose via API – Update
/api/config/platformsresponse inapi/routers/crawler.py
Reference implementations in media_platform/douyin.py and media_platform/xhs.py demonstrate the required interface methods.
Utility Helpers and Post-Processing
MediaCrawler's tools/ directory contains reusable components for data processing integration:
tools.async_file_writer.AsyncFileWriter– Async file operations for result handling- Word cloud generators and text analysis utilities
Import these directly into your pipeline for custom result processing without running the full crawler.
Summary
- Python embedding – Import
CrawlerFactoryandconfigfrommain.pyfor in-process integration - HTTP orchestration – Run
api/main.pywith Uvicorn to expose REST control endpoints - Visual interface – Use or embed the Vue-based
webui/for user-facing crawler management - Platform extension – Follow the
AbstractCrawlerpattern to add new data sources
Frequently Asked Questions
Can I run MediaCrawler inside a Docker container?
Yes. Start the FastAPI server with uvicorn api.main:app --host 0.0.0.0 --port 8080 and expose that port. For headed browser automation (QR code login), use Docker with browser support or mount a local browser profile.
How do I authenticate programmatically without QR codes?
Set config.LOGIN_TYPE = "cookie" and provide valid cookies via config.COOKIES or the API's cookies field in CrawlerStartRequest. Cookie strings must match the target platform's format.
Is there a rate limiting or scheduling mechanism built in?
The core crawler does not include scheduler logic, but you can implement it externally: use Python's asyncio.sleep() between calls, wrap in Celery tasks, or trigger the /api/crawler/start endpoint from cron jobs or Airflow DAGs.
Can I stream crawl progress to my application?
Yes. The api/routers/websocket.py module provides WebSocket endpoints for real-time updates. Alternatively, poll /api/crawler/logs or /api/crawler/status for status-based progress tracking.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →