# How to Integrate MediaCrawler with Other Tools: A Complete Integration Guide

> Learn how to integrate MediaCrawler with other tools using Python imports, REST API calls, or the WebUI for seamless automation and embedding. Explore the complete integration guide.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-12

---

**You can integrate MediaCrawler via direct Python imports, REST API calls, or the built-in WebUI—each suited to different automation and embedding scenarios.**

MediaCrawler is a modular Python library designed for scraping content from major Chinese social platforms. Whether you're building a data pipeline, adding scraping capabilities to an existing service, or creating a custom front-end, this guide explains every **MediaCrawler integration** method based on the actual source code implementation in `NanmiCoder/MediaCrawler`.

## Python Library Integration

The most direct way to **integrate MediaCrawler with other tools** is to import and use its core classes as a Python library. The CLI entry point in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) exposes the same objects you can use programmatically.

### Basic Programmatic Usage

```python

# my_integration.py

from main import CrawlerFactory, config
from var import crawler_type_var

# Configure the crawler at runtime

config.PLATFORM = "dy"                 # Douyin

config.SAVE_DATA_OPTION = "json"       # Storage format

config.ENABLE_GET_COMMENTS = True

# Create and run crawler instance

crawler = CrawlerFactory.create_crawler(platform=config.PLATFORM)

import asyncio
asyncio.run(crawler.start())

```

**Key integration points:**

- **`CrawlerFactory`** ([`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)) – Maps platform strings (`xhs`, `dy`, `ks`, `bili`, `wb`, `tieba`, `zhihu`) to concrete crawler classes
- **`config`** ([`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)) – Central configuration object with all runtime options
- **`AbstractCrawler`** subclasses (`media_platform/*.py`) – Platform-specific implementations you can extend

This approach works well when embedding MediaCrawler into **Django, Flask, Celery, or custom async services**.

## FastAPI HTTP API Integration

For **language-agnostic integration** or remote control scenarios, run MediaCrawler as a standalone HTTP service. The FastAPI application in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) exposes full crawler control via REST endpoints.

### Starting the API Server

```bash
uv run uvicorn api.main:app --port 8080 --reload

```

### Core REST Endpoints

| Method | Endpoint | Purpose |
|--------|----------|---------|
| `POST` | `/api/crawler/start` | Launch crawl with JSON configuration |
| `POST` | `/api/crawler/stop` | Gracefully stop running crawler |
| `GET` | `/api/crawler/status` | Check crawler state and PID |
| `GET` | `/api/crawler/logs?limit=50` | Retrieve recent log entries |
| `GET` | `/api/config/platforms` | List supported platforms |
| `GET` | `/api/config/options` | Get login types, modes, storage options |

### Example API Call

```python
import requests

payload = {
    "platform": "xhs",
    "login_type": "qrcode",
    "crawler_type": "search",
    "keyword": "机器学习",
    "max_pages": 5
}

response = requests.post(
    "http://localhost:8080/api/crawler/start",
    json=payload
)
print(response.json())  # Returns job status

```

The **request schema** is defined in [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py) (`CrawlerStartRequest`), with fields for platform selection, authentication method, crawl mode, and platform-specific parameters.

**Router implementations:**

- [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py) – Start/stop/status/log endpoints
- [`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py) – Data export utilities
- [`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py) – Real-time progress streaming (optional)

## WebUI Integration and Embedding

MediaCrawler includes a **Vue 3 + Vite front-end** in the `webui/` directory that communicates with the FastAPI backend.

### Development Setup

```bash

# Terminal 1: API server

uv run uvicorn api.main:app --port 8080 --reload

# Terminal 2: WebUI

cd webui
npm install
npm run dev  # http://localhost:5173

```

### Production Deployment

After building (`npm run build`), FastAPI serves the static files at `/` and `/static`. This enables two **WebUI integration** patterns:

- **Standalone application** – Users access the crawler through the provided interface
- **Embedded iframe** – Integrate the visual crawler control into existing dashboards or admin panels

Since the WebUI uses the same REST API documented above, you can also **build custom front-ends** that call these endpoints directly.

## Extending MediaCrawler for Custom Platforms

To **integrate a new social platform** into MediaCrawler's architecture:

1. **Create crawler class** – Subclass `AbstractCrawler` in [`media_platform/your_platform.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/your_platform.py)
2. **Register in factory** – Add mapping in `CrawlerFactory.CRAWLERS` ([`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py))
3. **Add configuration** – Create [`config/your_platform_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/your_platform_config.py) (optional)
4. **Expose via API** – Update `/api/config/platforms` response in [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py)

Reference implementations in [`media_platform/douyin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin.py) and [`media_platform/xhs.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs.py) demonstrate the required interface methods.

## Utility Helpers and Post-Processing

MediaCrawler's `tools/` directory contains reusable components for **data processing integration**:

- `tools.async_file_writer.AsyncFileWriter` – Async file operations for result handling
- Word cloud generators and text analysis utilities

Import these directly into your pipeline for custom result processing without running the full crawler.

## Summary

- **Python embedding** – Import `CrawlerFactory` and `config` from [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) for in-process integration
- **HTTP orchestration** – Run [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) with Uvicorn to expose REST control endpoints
- **Visual interface** – Use or embed the Vue-based `webui/` for user-facing crawler management
- **Platform extension** – Follow the `AbstractCrawler` pattern to add new data sources

## Frequently Asked Questions

### Can I run MediaCrawler inside a Docker container?

Yes. Start the FastAPI server with `uvicorn api.main:app --host 0.0.0.0 --port 8080` and expose that port. For headed browser automation (QR code login), use Docker with browser support or mount a local browser profile.

### How do I authenticate programmatically without QR codes?

Set `config.LOGIN_TYPE = "cookie"` and provide valid cookies via `config.COOKIES` or the API's `cookies` field in `CrawlerStartRequest`. Cookie strings must match the target platform's format.

### Is there a rate limiting or scheduling mechanism built in?

The core crawler does not include scheduler logic, but you can implement it externally: use Python's `asyncio.sleep()` between calls, wrap in Celery tasks, or trigger the `/api/crawler/start` endpoint from cron jobs or Airflow DAGs.

### Can I stream crawl progress to my application?

Yes. The [`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py) module provides WebSocket endpoints for real-time updates. Alternatively, poll `/api/crawler/logs` or `/api/crawler/status` for status-based progress tracking.