# How to Integrate MediaCrawler with Other Applications: Python and HTTP API Guide

> Learn to integrate MediaCrawler with other applications using Python imports or its HTTP API. Control crawlers remotely and streamline your media management with this comprehensive guide.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-29

---

**You can integrate MediaCrawler via direct Python imports using `CrawlerFactory` and `config` from [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py), or by running the FastAPI server in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) and calling its REST endpoints to control crawlers remotely.**

MediaCrawler is a modular Python library designed for scraping content from Chinese social media platforms (Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, Zhihu). According to the NanmiCoder/MediaCrawler source code, the architecture supports two primary integration patterns: embedding as a Python library for programmatic control, or deploying as a standalone HTTP service with FastAPI endpoints.

## Direct Python Integration

For applications that require tight coupling with the crawler logic, import the same factory classes used by the CLI. This method provides full control over configuration and execution flow.

### Importing Core Components

The entry point for programmatic integration is [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py), which exposes `CrawlerFactory` and the global `config` object. The factory maps platform identifiers to concrete implementations located in `media_platform/*`.

```python

# my_app.py

from main import CrawlerFactory, config   # source: main.py

from var import crawler_type_var          # source: var.py

# Configure the crawler at runtime

config.PLATFORM = "dy"                   # e.g., Douyin (dy)

config.SAVE_DATA_OPTION = "json"          # storage format: json, csv, db

config.ENABLE_GET_COMMENTS = True         # enable comment extraction

```

### Launching the Crawler Programmatically

After configuration, instantiate the crawler via `CrawlerFactory.create_crawler()` and execute it using the async `start()` method. This is the same API used internally by the command-line interface.

```python

# Create platform-specific crawler instance

crawler = CrawlerFactory.create_crawler(platform=config.PLATFORM)

# Run the async crawler

import asyncio
asyncio.run(crawler.start())

```

**Key implementation files referenced:**
- [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) – Contains `CrawlerFactory` and entry-point logic
- [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) – Central configuration object (`config.PLATFORM`, `config.SAVE_DATA_OPTION`)
- [`media_platform/douyin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin.py), [`media_platform/xhs.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs.py) – Concrete crawler implementations inheriting from `AbstractCrawler`

## HTTP API Integration

For microservices architectures or polyglot environments, run MediaCrawler as an HTTP server using the FastAPI application defined in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py).

### Starting the FastAPI Server

Launch the server using Uvicorn with automatic reloading for development:

```bash
uv run uvicorn api.main:app --port 8080 --reload

```

### REST Endpoints Reference

External applications can control crawling operations via the following endpoints implemented in [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py):

- **`POST /api/crawler/start`** – Initiates a crawl job with parameters defined in `CrawlerStartRequest` (schema located in [`api/schemas/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py))
- **`POST /api/crawler/stop`** – Gracefully terminates the current crawler process
- **`GET /api/crawler/status`** – Returns `CrawlerStatusResponse` with running state and PID
- **`GET /api/crawler/logs?limit=50`** – Retrieves recent log entries for monitoring
- **`GET /api/config/platforms`** – Lists supported platforms: `xhs`, `dy`, `ks`, `bili`, `wb`, `tieba`, `zhihu`
- **`GET /api/config/options`** – Returns available login types (`qrcode`, `cookie`) and crawl modes (`search`, `detail`, `creator`)

### Sample HTTP Requests

Start a crawler using standard HTTP tools or libraries:

```bash
curl -X POST http://localhost:8080/api/crawler/start \
  -H "Content-Type: application/json" \
  -d '{"platform":"xhs","login_type":"qrcode","crawler_type":"search","keyword":"AI"}'

```

For Python-based integrations, use the `requests` library:

```python
import requests

payload = {
    "platform": "xhs",
    "login_type": "qrcode",
    "crawler_type": "search",
    "keyword": "机器学习",
    "max_pages": 5
}

response = requests.post(
    "http://localhost:8080/api/crawler/start",
    json=payload
)
print(response.json())

```

## WebUI Integration

MediaCrawler includes a **Vue.js front-end** located in the `webui/` directory that communicates with the FastAPI backend. After building (`npm run build`), the interface is served at the root path (`/`) and `/static`. You can embed this dashboard into existing applications via an `<iframe>` or run it independently during development:

```bash
cd webui
npm install
npm run dev  # Serves on http://localhost:5173

```

## Extending MediaCrawler for Custom Platforms

To integrate a new social platform, extend the `AbstractCrawler` class in [`media_platform/your_platform.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/your_platform.py) and register it in `CrawlerFactory.CRAWLERS` within [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py). Add platform-specific configuration in [`config/your_platform_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/your_platform_config.py) if needed, and update the router's platform list in `/api/config/platforms` to expose it via the HTTP API.

## Summary

- **Library integration**: Import `CrawlerFactory` from [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py), configure `config` options, and execute `await crawler.start()` for embedded Python workflows.
- **API integration**: Deploy [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) via Uvicorn to expose REST endpoints for starting, stopping, and monitoring crawlers remotely.
- **Configuration**: All settings are centralized in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), supporting storage backends including CSV, JSON, Excel, SQLite, MySQL, and MongoDB.
- **Extensibility**: New platforms are added by subclassing `AbstractCrawler` and registering with the factory pattern.

## Frequently Asked Questions

### How do I change storage options when integrating MediaCrawler?

Set the `SAVE_DATA_OPTION` attribute on the `config` object before creating the crawler instance. Supported values include `"csv"`, `"json"`, `"excel"`, and database options like `"db"` or `"mongodb"`, as defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py).

### Can I run MediaCrawler without the QR code login?

Yes. When calling the `/api/crawler/start` endpoint or configuring `config.LOGIN_TYPE`, set the value to `"cookie"` instead of `"qrcode"`. Ensure you have valid authentication cookies configured in your environment or config file before starting the crawl.

### Is the FastAPI server suitable for production use?

The FastAPI application in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) includes health checks and environment validation, but you should deploy it behind a reverse proxy (like Nginx) with proper authentication. Use `uvicorn` with multiple workers (`--workers 4`) rather than `--reload` for production workloads.

### How do I access real-time logs from the crawler?

Use the `GET /api/crawler/logs` endpoint with a `limit` query parameter (e.g., `?limit=100`) to retrieve recent entries. For real-time updates, the repository includes WebSocket support in [`api/routers/websocket.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/websocket.py) that streams log messages as they are generated.