How to Run MediaCrawler from Source: Complete Setup Guide

Clone the repository, install dependencies with uv sync, enable Chrome remote debugging on port 9222, and execute uv run main.py --platform xhs --lt qrcode --type search to begin crawling immediately.

MediaCrawler is an asynchronous Python framework designed to scrape content from major social media platforms including XiaoHongShu, Douyin, Bilibili, Weibo, Tieba, and Zhihu. Running MediaCrawler from source provides complete control over the crawler factory, browser automation modes, and data persistence layers. This guide covers installation via the uv package manager, CDP (Chrome DevTools Protocol) configuration, and execution through both CLI and WebUI interfaces.

Prerequisites

Before running MediaCrawler from source, ensure your environment meets these requirements:

  • uv (Python package manager) ≥ 0.4 — Installation guide
  • Chrome or Edge ≥ 144 — Required for CDP mode operation
  • Node.js ≥ 16.0.0 — Only needed if using the optional WebUI
  • Playwright — Only required if disabling CDP mode (uv run playwright install)

Installation and Setup

Clone the repository and install Python dependencies using uv:

git clone https://github.com/NanmiCoder/MediaCrawler.git
cd MediaCrawler
uv sync

The uv sync command installs the exact Python version and dependencies specified in the project configuration.

Understanding the Execution Architecture

When you run MediaCrawler from source, the execution follows this path in main.py (lines 99-122):

  1. Argument parsing — cmd_arg.parse_cmd() processes CLI options (line 103)
  2. Database initialization — db.init_db() prepares storage if SQL options are selected (lines 104-107)
  3. Crawler instantiation — CrawlerFactory.create_crawler() returns a platform-specific crawler (lines 113-115)
  4. Async execution — crawler.start() launches the asynchronous crawl (lines 114-115)

Platform-specific implementations reside in media_platform/<platform>.py (e.g., xhs.py), each extending AbstractCrawler defined in base/base_crawler.py.

Running from the Command Line

MediaCrawler defaults to CDP (Chrome DevTools Protocol) mode, which connects to your existing Chrome browser instance to reuse cookies and login states. To enable this:

  1. Open Chrome and navigate to chrome://inspect/#remote-debugging
  2. Enable remote debugging (Chrome listens on 127.0.0.1:9222 by default)
  3. Verify ENABLE_CDP_MODE=True in config/base_config.py (lines 55-59)

Execute Crawl Commands

Run a keyword search crawl on XiaoHongShu using:

uv run main.py --platform xhs --lt qrcode --type search

Fetch specific posts by ID using detail mode:

uv run main.py --platform xhs --lt qrcode --type detail

View all available options:

uv run main.py --help

The cmd_arg/arg.py module defines the CLI interface, supporting platforms (xhs, douyin, bilibili, weibo, tieba, zhihu), login types (qrcode, phone, etc.), and crawl types (search, detail).

Alternative: Playwright Mode

If CDP is unavailable, disable it in config/base_config.py by setting ENABLE_CDP_MODE=False, then install Playwright:

uv run playwright install

This runs a headless browser managed entirely by Playwright, though it may face higher detection rates compared to CDP mode.

Optional WebUI Setup

MediaCrawler includes a Vue.js frontend with FastAPI backend for visual configuration.

Start the FastAPI Backend

uv run uvicorn api.main:app --port 8080 --reload

The API server defined in api/main.py exposes endpoints that the frontend consumes.

Launch the Vue Frontend

cd webui
npm install
npm run dev

Access the interface at http://localhost:5173/ to configure platforms, monitor logs, and export data without using the command line.

Production Deployment

Build the frontend for static serving:

cd webui
npm install
npm run build
uv run uvicorn api.main:app --port 8080

This serves the pre-built UI from api/webui/ alongside the API.

Data Storage Configuration

Control output formats in config/base_config.py via SAVE_DATA_OPTION (default: jsonl at line 89). Supported formats include:

  • File-based: csv, json, jsonl, excel
  • Relational databases: sqlite, mysql, postgres

When database options are selected, main.py automatically initializes the connection (lines 110-112). The store/ directory contains platform-specific writers (e.g., store/excel_store_base.py), with Excel flushing occurring post-crawl via _flush_excel_if_needed() (lines 73-84).

Optional word-cloud generation activates when ENABLE_GET_WORDCLOUD=True (line 22) and data format is JSON/JSONL, triggered by _generate_wordcloud_if_needed() (lines 86-99).

Summary

  • Install: Use git clone followed by uv sync to establish the environment
  • Configure: Enable CDP mode in config/base_config.py (lines 55-59) and start Chrome with remote debugging on port 9222
  • Execute: Run uv run main.py with --platform, --lt, and --type arguments to start the asynchronous crawl via main.py lines 99-122
  • Extend: Access the WebUI by running uvicorn api.main:app and the Vue dev server for browser-based control
  • Store: Modify SAVE_DATA_OPTION in base configuration to switch between JSONL, CSV, Excel, or SQL backends managed by modules in store/

Frequently Asked Questions

What is the difference between CDP mode and Playwright mode?

CDP mode (default) connects MediaCrawler to a locally running Chrome instance via the Chrome DevTools Protocol on port 9222, allowing the crawler to reuse your existing browser state, cookies, and extensions. This significantly reduces detection risk. Playwright mode launches a managed headless browser instance when ENABLE_CDP_MODE=False, requiring uv run playwright install but operating independently of your local Chrome installation.

Which file controls the global crawler configuration?

The config/base_config.py file contains all global defaults including ENABLE_CDP_MODE (lines 55-59), SAVE_DATA_OPTION (line 89), and ENABLE_GET_WORDCLOUD (line 22). Changes to this file affect all crawl operations unless overridden by CLI arguments processed in cmd_arg/arg.py.

How do I add a new social media platform to MediaCrawler?

Create a new crawler class in media_platform/<platform>.py that inherits from AbstractCrawler (defined in base/base_crawler.py). Implement the required async methods, then register the platform in the CrawlerFactory so that CrawlerFactory.create_crawler() can instantiate it when --platform <your_platform> is passed to main.py.

Why does my crawl fail to start with a connection error?

Ensure Chrome is running with remote debugging enabled on 127.0.0.1:9222 (the default CDP_DEBUG_PORT specified in config/base_config.py). Verify that ENABLE_CDP_MODE=True (lines 55-59) and that no firewall is blocking the local connection. If using Playwright mode instead, confirm you have executed uv run playwright install to download the necessary browser binaries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →