# MediaCrawler Usage Examples: CLI, API, WebSocket, and Python SDK Patterns

> Explore MediaCrawler usage examples including CLI, API, WebSocket, and Python SDK patterns for scraping Xiaohongshu, Douyin, Bilibili, and Weibo data. Get started today.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: getting-started
- Published: 2026-08-12

---

**MediaCrawler supports four primary usage patterns—command-line interface, FastAPI server with Web UI, WebSocket programmatic access, and direct Python SDK integration—to scrape data from platforms including Xiaohongshu, Douyin, Bilibili, and Weibo.**

MediaCrawler is a multi-platform self-media data crawler built on **Playwright** or optional **CDP mode** that fetches posts, comments, and creator profiles. This guide demonstrates practical **MediaCrawler usage examples** drawn directly from the `NanmiCoder/MediaCrawler` source code, covering everything from basic CLI commands to embedding the crawler in your own applications.

## Core Architecture Overview

Understanding the codebase structure helps you choose the right integration approach. The project is organized into these key components:

| Component | Description | Source File |
|---|---|---|
| **BaseCrawler** | Common crawling logic, pagination, and data extraction utilities | [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) |
| **Platform Models** | Platform-specific fields and parsers (e.g., Xiaohongshu, Douyin) | [`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py), [`model/m_douyin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_douyin.py) |
| **Configuration** | Runtime options for login method, proxy pool, and output format | [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) |
| **Cache Layer** | Optional Redis or local cache for cookies and rate-limit data | [`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py) |
| **API Service** | FastAPI-based server exposing crawler functions via HTTP/WebSocket | [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py), [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py) |
| **CLI Entry** | Command-line argument parsing and crawler dispatch | [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) |

The general workflow remains consistent across all usage patterns: configure your platform and options, authenticate via QR code or cookies, execute the crawl, and retrieve data in your chosen format.

## Basic CLI Usage Examples

The simplest way to run MediaCrawler is through the command-line interface using `uv run main.py` with platform-specific flags.

### Platform-Specific Search Commands

```bash

# Xiaohongshu (小红书) - search by keyword

uv run main.py --platform xhs --lt qrcode --type search

# Xiaohongshu - crawl specific post IDs

uv run main.py --platform xhs --lt qrcode --type detail

# Douyin (抖音) - search by keyword

uv run main.py --platform dy --lt qrcode --type search

# Bilibili - search by keyword

uv run main.py --platform bili --lt qrcode --type search

# Weibo - search by keyword

uv run main.py --platform weibo --lt qrcode --type search

# Tieba - search by keyword

uv run main.py --platform tieba --lt qrcode --type search

# Zhihu - search by keyword

uv run main.py --platform zhihu --lt qrcode --type search

```

**Parameter reference:**
- `--platform`: Platform identifier (`xhs`, `dy`, `bili`, `weibo`, `tieba`, `zhihu`, `ks` for Kuaishou)
- `--lt`: Login type (`qrcode`, `cookie`, `phone`, or other platform-specific methods)
- `--type`: Crawl mode (`search`, `detail`, `creator`, or `comment`)

All CLI commands are documented in the repository README at [`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md).

## Excel Export Configuration Example

MediaCrawler supports multiple output formats including **JSONL**, **CSV**, **SQLite**, **MySQL**, and **Excel**. To export crawled data directly to Excel, modify your configuration and specify the output format.

### Step 1: Configure Output Format

Edit [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to set the default save option:

```python

# config/base_config.py

SAVE_DATA_OPTION = "excel"   # Options: "jsonl", "csv", "db", "sqlite", "excel"

```

### Step 2: Run with Excel Output

```bash

# Override via CLI flag

uv run main.py --platform xhs --lt qrcode --type search --save_data_option excel

# Douyin to Excel

uv run main.py --platform dy --lt qrcode --type search --save_data_option excel

# Bilibili to Excel

uv run main.py --platform bili --lt qrcode --type search --save_data_option excel

```

Excel files are written to `data/<platform>/` with auto-generated filenames like `xhs_search_20250128_143025.xlsx`. For detailed export options, see [`docs/excel_export_guide.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/docs/excel_export_guide.md).

## FastAPI Server and Web UI Usage

MediaCrawler exposes a **REST API and WebSocket interface** through a FastAPI server, enabling programmatic access and a browser-based management interface.

### Starting the API Server

```bash

# Terminal 1 - FastAPI backend

uv run uvicorn api.main:app --port 8080 --reload

```

The API entry point is defined in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py), with crawler routes in [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py).

### Launching the Web UI

```bash

# Terminal 2 - Vite frontend

cd webui
npm install
npm run dev

```

Access the interface at `http://localhost:5173/`. The Web UI communicates with the API to trigger crawls, monitor progress, and download results.

## WebSocket Programmatic Client Example

For real-time crawl status and streaming results, connect to the **WebSocket endpoint** at `/ws/crawler`. Here is a minimal Python client implementation:

```python
import asyncio
import json
import websockets


async def crawl_xiaohongshu():
    uri = "ws://localhost:8080/ws/crawler"
    
    async with websockets.connect(uri) as websocket:
        # Send crawl configuration

        await websocket.send(json.dumps({
            "platform": "xhs",
            "login_type": "qrcode",
            "crawl_type": "search",
            "save_data_option": "excel",
            "keyword": "your_search_keyword"  # Optional: specify search term

        }))
        
        # Stream results and status messages

        while True:
            message = await websocket.recv()
            data = json.loads(message)
            print(f"Received: {data}")


if __name__ == "__main__":
    asyncio.run(crawl_xiaohongshu())

```

The WebSocket handler processes these payloads in [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py), dispatching to the appropriate platform crawler and streaming progress updates back to the client.

## Python SDK Integration Example

Embed MediaCrawler directly in your Python applications using the **CrawlerManager** service class.

### Direct SDK Usage

```python
import asyncio
from api.services.crawler_manager import CrawlerManager


async def run_embedded_crawl():
    """Execute crawler programmatically without CLI."""
    manager = CrawlerManager()
    
    # Example: Xiaohongshu keyword search to JSONL

    await manager.run_crawler(
        platform="xhs",
        login_type="qrcode",
        crawl_type="search",
        save_data_option="jsonl",  # Options: "excel", "csv", "db", "sqlite", "jsonl"

        keyword="target_keyword"
    )
    
    # Example: Douyin creator profile crawl to Excel

    await manager.run_crawler(
        platform="dy",
        login_type="qrcode",
        crawl_type="creator",
        save_data_option="excel",
        creator_id="creator_sec_id"
    )


# Execute with uv run

# uv run -c "import asyncio; from your_module import run_embedded_crawl; asyncio.run(run_embedded_crawl())"

if __name__ == "__main__":
    asyncio.run(run_embedded_crawl())

```

The `CrawlerManager` class in [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py) handles platform instantiation, login workflow, and result persistence, making it suitable for custom pipelines, scheduled jobs, or integration with larger data systems.

## Playwright vs. CDP Mode Configuration

MediaCrawler supports two browser automation backends. CDP mode connects to an existing Chrome instance; Playwright launches managed browser contexts.

### Switching Modes

Edit [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to select your backend:

```python

# config/base_config.py

ENABLE_CDP_MODE = True   # Default: use Chrome DevTools Protocol

# ENABLE_CDP_MODE = False  # Use Playwright-managed browsers instead

```

### Playwright Setup

If using Playwright mode, install browser binaries first:

```bash
uv run playwright install chromium

```

CDP mode requires a running Chrome instance with remote debugging enabled:

```bash

# Launch Chrome with remote debugging port

chrome --remote-debugging-port=9222 --user-data-dir=/tmp/chrome_dev_profile

```

For detailed CDP configuration, refer to `docs/CDP模式使用指南.md`.

## Complete Configuration Reference Example

Here is a production-ready [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) excerpt showing key options:

```python

# config/base_config.py

# Platform and crawl settings

PLATFORM = "xhs"           # xhs, dy, ks, bili, weibo, tieba, zhihu

CRAWLER_TYPE = "search"    # search, detail, creator, comment

KEYWORDS = ["python教程", "数据分析"]

# Login configuration

LOGIN_TYPE = "qrcode"      # qrcode, cookie, phone

# Browser settings

ENABLE_CDP_MODE = True
HEADLESS = False           # Set True for production servers

# Data persistence

SAVE_DATA_OPTION = "excel"  # jsonl, csv, sqlite, db, excel

DATA_DIR = "data"

# Rate limiting and caching

ENABLE_PROXY = False
PROXY_POOL = ["http://proxy1:8080", "http://proxy2:8080"]
CACHE_TYPE = "local"       # local, redis

REDIS_URL = "redis://localhost:6379/0"

# Platform-specific limits

MAX_NOTES = 100            # Maximum posts to fetch per search

MAX_COMMENTS = 50          # Maximum comments per post

```

## Summary

- **CLI execution** with `uv run main.py` provides the fastest path to crawling data from supported platforms using QR code authentication.
- **Excel export** requires setting `SAVE_DATA_OPTION = "excel"` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) or passing `--save_data_option excel` on the command line, with files written to `data/<platform>/`.
- **FastAPI server** architecture enables both HTTP API access and WebSocket streaming for real-time crawl monitoring, with source code in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) and [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py).
- **Python SDK integration** via `CrawlerManager` from [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py) allows embedding crawlers in custom applications without subprocess calls.
- **CDP mode** (default) connects to existing Chrome instances for reduced detection; **Playwright mode** offers managed browser contexts requiring `uv run playwright install`.

## Frequently Asked Questions

### How do I run MediaCrawler without using the command line?

Use the **Python SDK** approach. Import `CrawlerManager` from [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py) and call `await manager.run_crawler()` with your platform and crawl parameters. This embeds the crawler directly in your application without subprocess overhead.

### Can I export MediaCrawler results directly to Excel?

Yes. Set `SAVE_DATA_OPTION = "excel"` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) or pass `--save_data_option excel` to any CLI command. The crawler generates `.xlsx` files under `data/<platform>/` with timestamps in filenames. See [`docs/excel_export_guide.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/docs/excel_export_guide.md) for formatting options.

### What is the difference between CDP mode and Playwright mode?

**CDP mode** (default, `ENABLE_CDP_MODE = True`) connects MediaCrawler to an existing Chrome instance via the Chrome DevTools Protocol, reducing bot detection by using your normal browser profile. **Playwright mode** (`ENABLE_CDP_MODE = False`) launches fresh, isolated browser contexts managed by Playwright, requiring `uv run playwright install` but offering more consistent environments for server deployments.