How to Integrate MediaCrawler with Other Applications: Python and HTTP API Guide
You can integrate MediaCrawler via direct Python imports using CrawlerFactory and config from main.py, or by running the FastAPI server in api/main.py and calling its REST endpoints to control crawlers remotely.
MediaCrawler is a modular Python library designed for scraping content from Chinese social media platforms (Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, Zhihu). According to the NanmiCoder/MediaCrawler source code, the architecture supports two primary integration patterns: embedding as a Python library for programmatic control, or deploying as a standalone HTTP service with FastAPI endpoints.
Direct Python Integration
For applications that require tight coupling with the crawler logic, import the same factory classes used by the CLI. This method provides full control over configuration and execution flow.
Importing Core Components
The entry point for programmatic integration is main.py, which exposes CrawlerFactory and the global config object. The factory maps platform identifiers to concrete implementations located in media_platform/*.
# my_app.py
from main import CrawlerFactory, config # source: main.py
from var import crawler_type_var # source: var.py
# Configure the crawler at runtime
config.PLATFORM = "dy" # e.g., Douyin (dy)
config.SAVE_DATA_OPTION = "json" # storage format: json, csv, db
config.ENABLE_GET_COMMENTS = True # enable comment extraction
Launching the Crawler Programmatically
After configuration, instantiate the crawler via CrawlerFactory.create_crawler() and execute it using the async start() method. This is the same API used internally by the command-line interface.
# Create platform-specific crawler instance
crawler = CrawlerFactory.create_crawler(platform=config.PLATFORM)
# Run the async crawler
import asyncio
asyncio.run(crawler.start())
Key implementation files referenced:
main.py– ContainsCrawlerFactoryand entry-point logicconfig/base_config.py– Central configuration object (config.PLATFORM,config.SAVE_DATA_OPTION)media_platform/douyin.py,media_platform/xhs.py– Concrete crawler implementations inheriting fromAbstractCrawler
HTTP API Integration
For microservices architectures or polyglot environments, run MediaCrawler as an HTTP server using the FastAPI application defined in api/main.py.
Starting the FastAPI Server
Launch the server using Uvicorn with automatic reloading for development:
uv run uvicorn api.main:app --port 8080 --reload
REST Endpoints Reference
External applications can control crawling operations via the following endpoints implemented in api/routers/crawler.py:
POST /api/crawler/start– Initiates a crawl job with parameters defined inCrawlerStartRequest(schema located inapi/schemas/crawler.py)POST /api/crawler/stop– Gracefully terminates the current crawler processGET /api/crawler/status– ReturnsCrawlerStatusResponsewith running state and PIDGET /api/crawler/logs?limit=50– Retrieves recent log entries for monitoringGET /api/config/platforms– Lists supported platforms:xhs,dy,ks,bili,wb,tieba,zhihuGET /api/config/options– Returns available login types (qrcode,cookie) and crawl modes (search,detail,creator)
Sample HTTP Requests
Start a crawler using standard HTTP tools or libraries:
curl -X POST http://localhost:8080/api/crawler/start \
-H "Content-Type: application/json" \
-d '{"platform":"xhs","login_type":"qrcode","crawler_type":"search","keyword":"AI"}'
For Python-based integrations, use the requests library:
import requests
payload = {
"platform": "xhs",
"login_type": "qrcode",
"crawler_type": "search",
"keyword": "机器学习",
"max_pages": 5
}
response = requests.post(
"http://localhost:8080/api/crawler/start",
json=payload
)
print(response.json())
WebUI Integration
MediaCrawler includes a Vue.js front-end located in the webui/ directory that communicates with the FastAPI backend. After building (npm run build), the interface is served at the root path (/) and /static. You can embed this dashboard into existing applications via an <iframe> or run it independently during development:
cd webui
npm install
npm run dev # Serves on http://localhost:5173
Extending MediaCrawler for Custom Platforms
To integrate a new social platform, extend the AbstractCrawler class in media_platform/your_platform.py and register it in CrawlerFactory.CRAWLERS within main.py. Add platform-specific configuration in config/your_platform_config.py if needed, and update the router's platform list in /api/config/platforms to expose it via the HTTP API.
Summary
- Library integration: Import
CrawlerFactoryfrommain.py, configureconfigoptions, and executeawait crawler.start()for embedded Python workflows. - API integration: Deploy
api/main.pyvia Uvicorn to expose REST endpoints for starting, stopping, and monitoring crawlers remotely. - Configuration: All settings are centralized in
config/base_config.py, supporting storage backends including CSV, JSON, Excel, SQLite, MySQL, and MongoDB. - Extensibility: New platforms are added by subclassing
AbstractCrawlerand registering with the factory pattern.
Frequently Asked Questions
How do I change storage options when integrating MediaCrawler?
Set the SAVE_DATA_OPTION attribute on the config object before creating the crawler instance. Supported values include "csv", "json", "excel", and database options like "db" or "mongodb", as defined in config/base_config.py.
Can I run MediaCrawler without the QR code login?
Yes. When calling the /api/crawler/start endpoint or configuring config.LOGIN_TYPE, set the value to "cookie" instead of "qrcode". Ensure you have valid authentication cookies configured in your environment or config file before starting the crawl.
Is the FastAPI server suitable for production use?
The FastAPI application in api/main.py includes health checks and environment validation, but you should deploy it behind a reverse proxy (like Nginx) with proper authentication. Use uvicorn with multiple workers (--workers 4) rather than --reload for production workloads.
How do I access real-time logs from the crawler?
Use the GET /api/crawler/logs endpoint with a limit query parameter (e.g., ?limit=100) to retrieve recent entries. For real-time updates, the repository includes WebSocket support in api/routers/websocket.py that streams log messages as they are generated.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →