MediaCrawler Usage Examples: CLI, API, WebSocket, and Python SDK Patterns

MediaCrawler supports four primary usage patterns—command-line interface, FastAPI server with Web UI, WebSocket programmatic access, and direct Python SDK integration—to scrape data from platforms including Xiaohongshu, Douyin, Bilibili, and Weibo.

MediaCrawler is a multi-platform self-media data crawler built on Playwright or optional CDP mode that fetches posts, comments, and creator profiles. This guide demonstrates practical MediaCrawler usage examples drawn directly from the NanmiCoder/MediaCrawler source code, covering everything from basic CLI commands to embedding the crawler in your own applications.

Core Architecture Overview

Understanding the codebase structure helps you choose the right integration approach. The project is organized into these key components:

Component Description Source File
BaseCrawler Common crawling logic, pagination, and data extraction utilities base/base_crawler.py
Platform Models Platform-specific fields and parsers (e.g., Xiaohongshu, Douyin) model/m_xiaohongshu.py, model/m_douyin.py
Configuration Runtime options for login method, proxy pool, and output format config/base_config.py
Cache Layer Optional Redis or local cache for cookies and rate-limit data cache/redis_cache.py
API Service FastAPI-based server exposing crawler functions via HTTP/WebSocket api/main.py, api/routers/crawler.py
CLI Entry Command-line argument parsing and crawler dispatch main.py

The general workflow remains consistent across all usage patterns: configure your platform and options, authenticate via QR code or cookies, execute the crawl, and retrieve data in your chosen format.

Basic CLI Usage Examples

The simplest way to run MediaCrawler is through the command-line interface using uv run main.py with platform-specific flags.

Platform-Specific Search Commands


# Xiaohongshu (小红书) - search by keyword

uv run main.py --platform xhs --lt qrcode --type search

# Xiaohongshu - crawl specific post IDs

uv run main.py --platform xhs --lt qrcode --type detail

# Douyin (抖音) - search by keyword

uv run main.py --platform dy --lt qrcode --type search

# Bilibili - search by keyword

uv run main.py --platform bili --lt qrcode --type search

# Weibo - search by keyword

uv run main.py --platform weibo --lt qrcode --type search

# Tieba - search by keyword

uv run main.py --platform tieba --lt qrcode --type search

# Zhihu - search by keyword

uv run main.py --platform zhihu --lt qrcode --type search

Parameter reference:

  • --platform: Platform identifier (xhs, dy, bili, weibo, tieba, zhihu, ks for Kuaishou)
  • --lt: Login type (qrcode, cookie, phone, or other platform-specific methods)
  • --type: Crawl mode (search, detail, creator, or comment)

All CLI commands are documented in the repository README at README.md.

Excel Export Configuration Example

MediaCrawler supports multiple output formats including JSONL, CSV, SQLite, MySQL, and Excel. To export crawled data directly to Excel, modify your configuration and specify the output format.

Step 1: Configure Output Format

Edit config/base_config.py to set the default save option:


# config/base_config.py

SAVE_DATA_OPTION = "excel"   # Options: "jsonl", "csv", "db", "sqlite", "excel"

Step 2: Run with Excel Output


# Override via CLI flag

uv run main.py --platform xhs --lt qrcode --type search --save_data_option excel

# Douyin to Excel

uv run main.py --platform dy --lt qrcode --type search --save_data_option excel

# Bilibili to Excel

uv run main.py --platform bili --lt qrcode --type search --save_data_option excel

Excel files are written to data/<platform>/ with auto-generated filenames like xhs_search_20250128_143025.xlsx. For detailed export options, see docs/excel_export_guide.md.

FastAPI Server and Web UI Usage

MediaCrawler exposes a REST API and WebSocket interface through a FastAPI server, enabling programmatic access and a browser-based management interface.

Starting the API Server


# Terminal 1 - FastAPI backend

uv run uvicorn api.main:app --port 8080 --reload

The API entry point is defined in api/main.py, with crawler routes in api/routers/crawler.py.

Launching the Web UI


# Terminal 2 - Vite frontend

cd webui
npm install
npm run dev

Access the interface at http://localhost:5173/. The Web UI communicates with the API to trigger crawls, monitor progress, and download results.

WebSocket Programmatic Client Example

For real-time crawl status and streaming results, connect to the WebSocket endpoint at /ws/crawler. Here is a minimal Python client implementation:

import asyncio
import json
import websockets


async def crawl_xiaohongshu():
    uri = "ws://localhost:8080/ws/crawler"
    
    async with websockets.connect(uri) as websocket:
        # Send crawl configuration

        await websocket.send(json.dumps({
            "platform": "xhs",
            "login_type": "qrcode",
            "crawl_type": "search",
            "save_data_option": "excel",
            "keyword": "your_search_keyword"  # Optional: specify search term

        }))
        
        # Stream results and status messages

        while True:
            message = await websocket.recv()
            data = json.loads(message)
            print(f"Received: {data}")


if __name__ == "__main__":
    asyncio.run(crawl_xiaohongshu())

The WebSocket handler processes these payloads in api/routers/crawler.py, dispatching to the appropriate platform crawler and streaming progress updates back to the client.

Python SDK Integration Example

Embed MediaCrawler directly in your Python applications using the CrawlerManager service class.

Direct SDK Usage

import asyncio
from api.services.crawler_manager import CrawlerManager


async def run_embedded_crawl():
    """Execute crawler programmatically without CLI."""
    manager = CrawlerManager()
    
    # Example: Xiaohongshu keyword search to JSONL

    await manager.run_crawler(
        platform="xhs",
        login_type="qrcode",
        crawl_type="search",
        save_data_option="jsonl",  # Options: "excel", "csv", "db", "sqlite", "jsonl"

        keyword="target_keyword"
    )
    
    # Example: Douyin creator profile crawl to Excel

    await manager.run_crawler(
        platform="dy",
        login_type="qrcode",
        crawl_type="creator",
        save_data_option="excel",
        creator_id="creator_sec_id"
    )


# Execute with uv run

# uv run -c "import asyncio; from your_module import run_embedded_crawl; asyncio.run(run_embedded_crawl())"

if __name__ == "__main__":
    asyncio.run(run_embedded_crawl())

The CrawlerManager class in api/services/crawler_manager.py handles platform instantiation, login workflow, and result persistence, making it suitable for custom pipelines, scheduled jobs, or integration with larger data systems.

Playwright vs. CDP Mode Configuration

MediaCrawler supports two browser automation backends. CDP mode connects to an existing Chrome instance; Playwright launches managed browser contexts.

Switching Modes

Edit config/base_config.py to select your backend:


# config/base_config.py

ENABLE_CDP_MODE = True   # Default: use Chrome DevTools Protocol

# ENABLE_CDP_MODE = False  # Use Playwright-managed browsers instead

Playwright Setup

If using Playwright mode, install browser binaries first:

uv run playwright install chromium

CDP mode requires a running Chrome instance with remote debugging enabled:


# Launch Chrome with remote debugging port

chrome --remote-debugging-port=9222 --user-data-dir=/tmp/chrome_dev_profile

For detailed CDP configuration, refer to docs/CDP模式使用指南.md.

Complete Configuration Reference Example

Here is a production-ready config/base_config.py excerpt showing key options:


# config/base_config.py

# Platform and crawl settings

PLATFORM = "xhs"           # xhs, dy, ks, bili, weibo, tieba, zhihu

CRAWLER_TYPE = "search"    # search, detail, creator, comment

KEYWORDS = ["python教程", "数据分析"]

# Login configuration

LOGIN_TYPE = "qrcode"      # qrcode, cookie, phone

# Browser settings

ENABLE_CDP_MODE = True
HEADLESS = False           # Set True for production servers

# Data persistence

SAVE_DATA_OPTION = "excel"  # jsonl, csv, sqlite, db, excel

DATA_DIR = "data"

# Rate limiting and caching

ENABLE_PROXY = False
PROXY_POOL = ["http://proxy1:8080", "http://proxy2:8080"]
CACHE_TYPE = "local"       # local, redis

REDIS_URL = "redis://localhost:6379/0"

# Platform-specific limits

MAX_NOTES = 100            # Maximum posts to fetch per search

MAX_COMMENTS = 50          # Maximum comments per post

Summary

  • CLI execution with uv run main.py provides the fastest path to crawling data from supported platforms using QR code authentication.
  • Excel export requires setting SAVE_DATA_OPTION = "excel" in config/base_config.py or passing --save_data_option excel on the command line, with files written to data/<platform>/.
  • FastAPI server architecture enables both HTTP API access and WebSocket streaming for real-time crawl monitoring, with source code in api/main.py and api/routers/crawler.py.
  • Python SDK integration via CrawlerManager from api/services/crawler_manager.py allows embedding crawlers in custom applications without subprocess calls.
  • CDP mode (default) connects to existing Chrome instances for reduced detection; Playwright mode offers managed browser contexts requiring uv run playwright install.

Frequently Asked Questions

How do I run MediaCrawler without using the command line?

Use the Python SDK approach. Import CrawlerManager from api/services/crawler_manager.py and call await manager.run_crawler() with your platform and crawl parameters. This embeds the crawler directly in your application without subprocess overhead.

Can I export MediaCrawler results directly to Excel?

Yes. Set SAVE_DATA_OPTION = "excel" in config/base_config.py or pass --save_data_option excel to any CLI command. The crawler generates .xlsx files under data/<platform>/ with timestamps in filenames. See docs/excel_export_guide.md for formatting options.

What is the difference between CDP mode and Playwright mode?

CDP mode (default, ENABLE_CDP_MODE = True) connects MediaCrawler to an existing Chrome instance via the Chrome DevTools Protocol, reducing bot detection by using your normal browser profile. Playwright mode (ENABLE_CDP_MODE = False) launches fresh, isolated browser contexts managed by Playwright, requiring uv run playwright install but offering more consistent environments for server deployments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →