How to Integrate MediaCrawler with Other Tools: A Complete Integration Guide

You can integrate MediaCrawler via direct Python imports, REST API calls, or the built-in WebUI—each suited to different automation and embedding scenarios.

MediaCrawler is a modular Python library designed for scraping content from major Chinese social platforms. Whether you're building a data pipeline, adding scraping capabilities to an existing service, or creating a custom front-end, this guide explains every MediaCrawler integration method based on the actual source code implementation in NanmiCoder/MediaCrawler.

Python Library Integration

The most direct way to integrate MediaCrawler with other tools is to import and use its core classes as a Python library. The CLI entry point in main.py exposes the same objects you can use programmatically.

Basic Programmatic Usage


# my_integration.py

from main import CrawlerFactory, config
from var import crawler_type_var

# Configure the crawler at runtime

config.PLATFORM = "dy"                 # Douyin

config.SAVE_DATA_OPTION = "json"       # Storage format

config.ENABLE_GET_COMMENTS = True

# Create and run crawler instance

crawler = CrawlerFactory.create_crawler(platform=config.PLATFORM)

import asyncio
asyncio.run(crawler.start())

Key integration points:

  • CrawlerFactory (main.py) – Maps platform strings (xhs, dy, ks, bili, wb, tieba, zhihu) to concrete crawler classes
  • config (config/base_config.py) – Central configuration object with all runtime options
  • AbstractCrawler subclasses (media_platform/*.py) – Platform-specific implementations you can extend

This approach works well when embedding MediaCrawler into Django, Flask, Celery, or custom async services.

FastAPI HTTP API Integration

For language-agnostic integration or remote control scenarios, run MediaCrawler as a standalone HTTP service. The FastAPI application in api/main.py exposes full crawler control via REST endpoints.

Starting the API Server

uv run uvicorn api.main:app --port 8080 --reload

Core REST Endpoints

Method Endpoint Purpose
POST /api/crawler/start Launch crawl with JSON configuration
POST /api/crawler/stop Gracefully stop running crawler
GET /api/crawler/status Check crawler state and PID
GET /api/crawler/logs?limit=50 Retrieve recent log entries
GET /api/config/platforms List supported platforms
GET /api/config/options Get login types, modes, storage options

Example API Call

import requests

payload = {
    "platform": "xhs",
    "login_type": "qrcode",
    "crawler_type": "search",
    "keyword": "机器学习",
    "max_pages": 5
}

response = requests.post(
    "http://localhost:8080/api/crawler/start",
    json=payload
)
print(response.json())  # Returns job status

The request schema is defined in api/schemas/crawler.py (CrawlerStartRequest), with fields for platform selection, authentication method, crawl mode, and platform-specific parameters.

Router implementations:

WebUI Integration and Embedding

MediaCrawler includes a Vue 3 + Vite front-end in the webui/ directory that communicates with the FastAPI backend.

Development Setup


# Terminal 1: API server

uv run uvicorn api.main:app --port 8080 --reload

# Terminal 2: WebUI

cd webui
npm install
npm run dev  # http://localhost:5173

Production Deployment

After building (npm run build), FastAPI serves the static files at / and /static. This enables two WebUI integration patterns:

  • Standalone application – Users access the crawler through the provided interface
  • Embedded iframe – Integrate the visual crawler control into existing dashboards or admin panels

Since the WebUI uses the same REST API documented above, you can also build custom front-ends that call these endpoints directly.

Extending MediaCrawler for Custom Platforms

To integrate a new social platform into MediaCrawler's architecture:

  1. Create crawler class – Subclass AbstractCrawler in media_platform/your_platform.py
  2. Register in factory – Add mapping in CrawlerFactory.CRAWLERS (main.py)
  3. Add configuration – Create config/your_platform_config.py (optional)
  4. Expose via API – Update /api/config/platforms response in api/routers/crawler.py

Reference implementations in media_platform/douyin.py and media_platform/xhs.py demonstrate the required interface methods.

Utility Helpers and Post-Processing

MediaCrawler's tools/ directory contains reusable components for data processing integration:

  • tools.async_file_writer.AsyncFileWriter – Async file operations for result handling
  • Word cloud generators and text analysis utilities

Import these directly into your pipeline for custom result processing without running the full crawler.

Summary

  • Python embedding – Import CrawlerFactory and config from main.py for in-process integration
  • HTTP orchestration – Run api/main.py with Uvicorn to expose REST control endpoints
  • Visual interface – Use or embed the Vue-based webui/ for user-facing crawler management
  • Platform extension – Follow the AbstractCrawler pattern to add new data sources

Frequently Asked Questions

Can I run MediaCrawler inside a Docker container?

Yes. Start the FastAPI server with uvicorn api.main:app --host 0.0.0.0 --port 8080 and expose that port. For headed browser automation (QR code login), use Docker with browser support or mount a local browser profile.

How do I authenticate programmatically without QR codes?

Set config.LOGIN_TYPE = "cookie" and provide valid cookies via config.COOKIES or the API's cookies field in CrawlerStartRequest. Cookie strings must match the target platform's format.

Is there a rate limiting or scheduling mechanism built in?

The core crawler does not include scheduler logic, but you can implement it externally: use Python's asyncio.sleep() between calls, wrap in Celery tasks, or trigger the /api/crawler/start endpoint from cron jobs or Airflow DAGs.

Can I stream crawl progress to my application?

Yes. The api/routers/websocket.py module provides WebSocket endpoints for real-time updates. Alternatively, poll /api/crawler/logs or /api/crawler/status for status-based progress tracking.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →