How to Run MediaCrawler from Source: Complete Setup Guide
Clone the repository, install dependencies with uv sync, enable Chrome remote debugging on port 9222, and execute uv run main.py --platform xhs --lt qrcode --type search to begin crawling immediately.
MediaCrawler is an asynchronous Python framework designed to scrape content from major social media platforms including XiaoHongShu, Douyin, Bilibili, Weibo, Tieba, and Zhihu. Running MediaCrawler from source provides complete control over the crawler factory, browser automation modes, and data persistence layers. This guide covers installation via the uv package manager, CDP (Chrome DevTools Protocol) configuration, and execution through both CLI and WebUI interfaces.
Prerequisites
Before running MediaCrawler from source, ensure your environment meets these requirements:
- uv (Python package manager) ≥ 0.4 — Installation guide
- Chrome or Edge ≥ 144 — Required for CDP mode operation
- Node.js ≥ 16.0.0 — Only needed if using the optional WebUI
- Playwright — Only required if disabling CDP mode (
uv run playwright install)
Installation and Setup
Clone the repository and install Python dependencies using uv:
git clone https://github.com/NanmiCoder/MediaCrawler.git
cd MediaCrawler
uv sync
The uv sync command installs the exact Python version and dependencies specified in the project configuration.
Understanding the Execution Architecture
When you run MediaCrawler from source, the execution follows this path in main.py (lines 99-122):
- Argument parsing —
cmd_arg.parse_cmd()processes CLI options (line 103) - Database initialization —
db.init_db()prepares storage if SQL options are selected (lines 104-107) - Crawler instantiation —
CrawlerFactory.create_crawler()returns a platform-specific crawler (lines 113-115) - Async execution —
crawler.start()launches the asynchronous crawl (lines 114-115)
Platform-specific implementations reside in media_platform/<platform>.py (e.g., xhs.py), each extending AbstractCrawler defined in base/base_crawler.py.
Running from the Command Line
Configure CDP Mode (Recommended)
MediaCrawler defaults to CDP (Chrome DevTools Protocol) mode, which connects to your existing Chrome browser instance to reuse cookies and login states. To enable this:
- Open Chrome and navigate to
chrome://inspect/#remote-debugging - Enable remote debugging (Chrome listens on
127.0.0.1:9222by default) - Verify
ENABLE_CDP_MODE=Trueinconfig/base_config.py(lines 55-59)
Execute Crawl Commands
Run a keyword search crawl on XiaoHongShu using:
uv run main.py --platform xhs --lt qrcode --type search
Fetch specific posts by ID using detail mode:
uv run main.py --platform xhs --lt qrcode --type detail
View all available options:
uv run main.py --help
The cmd_arg/arg.py module defines the CLI interface, supporting platforms (xhs, douyin, bilibili, weibo, tieba, zhihu), login types (qrcode, phone, etc.), and crawl types (search, detail).
Alternative: Playwright Mode
If CDP is unavailable, disable it in config/base_config.py by setting ENABLE_CDP_MODE=False, then install Playwright:
uv run playwright install
This runs a headless browser managed entirely by Playwright, though it may face higher detection rates compared to CDP mode.
Optional WebUI Setup
MediaCrawler includes a Vue.js frontend with FastAPI backend for visual configuration.
Start the FastAPI Backend
uv run uvicorn api.main:app --port 8080 --reload
The API server defined in api/main.py exposes endpoints that the frontend consumes.
Launch the Vue Frontend
cd webui
npm install
npm run dev
Access the interface at http://localhost:5173/ to configure platforms, monitor logs, and export data without using the command line.
Production Deployment
Build the frontend for static serving:
cd webui
npm install
npm run build
uv run uvicorn api.main:app --port 8080
This serves the pre-built UI from api/webui/ alongside the API.
Data Storage Configuration
Control output formats in config/base_config.py via SAVE_DATA_OPTION (default: jsonl at line 89). Supported formats include:
- File-based:
csv,json,jsonl,excel - Relational databases:
sqlite,mysql,postgres
When database options are selected, main.py automatically initializes the connection (lines 110-112). The store/ directory contains platform-specific writers (e.g., store/excel_store_base.py), with Excel flushing occurring post-crawl via _flush_excel_if_needed() (lines 73-84).
Optional word-cloud generation activates when ENABLE_GET_WORDCLOUD=True (line 22) and data format is JSON/JSONL, triggered by _generate_wordcloud_if_needed() (lines 86-99).
Summary
- Install: Use
git clonefollowed byuv syncto establish the environment - Configure: Enable CDP mode in
config/base_config.py(lines 55-59) and start Chrome with remote debugging on port 9222 - Execute: Run
uv run main.pywith--platform,--lt, and--typearguments to start the asynchronous crawl viamain.pylines 99-122 - Extend: Access the WebUI by running
uvicorn api.main:appand the Vue dev server for browser-based control - Store: Modify
SAVE_DATA_OPTIONin base configuration to switch between JSONL, CSV, Excel, or SQL backends managed by modules instore/
Frequently Asked Questions
What is the difference between CDP mode and Playwright mode?
CDP mode (default) connects MediaCrawler to a locally running Chrome instance via the Chrome DevTools Protocol on port 9222, allowing the crawler to reuse your existing browser state, cookies, and extensions. This significantly reduces detection risk. Playwright mode launches a managed headless browser instance when ENABLE_CDP_MODE=False, requiring uv run playwright install but operating independently of your local Chrome installation.
Which file controls the global crawler configuration?
The config/base_config.py file contains all global defaults including ENABLE_CDP_MODE (lines 55-59), SAVE_DATA_OPTION (line 89), and ENABLE_GET_WORDCLOUD (line 22). Changes to this file affect all crawl operations unless overridden by CLI arguments processed in cmd_arg/arg.py.
How do I add a new social media platform to MediaCrawler?
Create a new crawler class in media_platform/<platform>.py that inherits from AbstractCrawler (defined in base/base_crawler.py). Implement the required async methods, then register the platform in the CrawlerFactory so that CrawlerFactory.create_crawler() can instantiate it when --platform <your_platform> is passed to main.py.
Why does my crawl fail to start with a connection error?
Ensure Chrome is running with remote debugging enabled on 127.0.0.1:9222 (the default CDP_DEBUG_PORT specified in config/base_config.py). Verify that ENABLE_CDP_MODE=True (lines 55-59) and that no firewall is blocking the local connection. If using Playwright mode instead, confirm you have executed uv run playwright install to download the necessary browser binaries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →