# How to Run MediaCrawler from Source: Complete Setup Guide

> Learn how to run MediaCrawler from source with this setup guide. Clone the repo, install dependencies, and start crawling immediately by following simple steps.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-29

---

**Clone the repository, install dependencies with `uv sync`, enable Chrome remote debugging on port 9222, and execute `uv run main.py --platform xhs --lt qrcode --type search` to begin crawling immediately.**

MediaCrawler is an asynchronous Python framework designed to scrape content from major social media platforms including XiaoHongShu, Douyin, Bilibili, Weibo, Tieba, and Zhihu. Running MediaCrawler from source provides complete control over the crawler factory, browser automation modes, and data persistence layers. This guide covers installation via the `uv` package manager, CDP (Chrome DevTools Protocol) configuration, and execution through both CLI and WebUI interfaces.

## Prerequisites

Before running MediaCrawler from source, ensure your environment meets these requirements:

- **uv** (Python package manager) ≥ 0.4 — [Installation guide](https://docs.astral.sh/uv/getting-started/installation)
- **Chrome or Edge** ≥ 144 — Required for CDP mode operation
- **Node.js** ≥ 16.0.0 — Only needed if using the optional WebUI
- **Playwright** — Only required if disabling CDP mode (`uv run playwright install`)

## Installation and Setup

Clone the repository and install Python dependencies using `uv`:

```bash
git clone https://github.com/NanmiCoder/MediaCrawler.git
cd MediaCrawler
uv sync

```

The `uv sync` command installs the exact Python version and dependencies specified in the project configuration.

## Understanding the Execution Architecture

When you run MediaCrawler from source, the execution follows this path in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) (lines 99-122):

1. **Argument parsing** — `cmd_arg.parse_cmd()` processes CLI options (line 103)
2. **Database initialization** — `db.init_db()` prepares storage if SQL options are selected (lines 104-107)
3. **Crawler instantiation** — `CrawlerFactory.create_crawler()` returns a platform-specific crawler (lines 113-115)
4. **Async execution** — `crawler.start()` launches the asynchronous crawl (lines 114-115)

Platform-specific implementations reside in `media_platform/<platform>.py` (e.g., [`xhs.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/xhs.py)), each extending `AbstractCrawler` defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py).

## Running from the Command Line

### Configure CDP Mode (Recommended)

MediaCrawler defaults to **CDP (Chrome DevTools Protocol)** mode, which connects to your existing Chrome browser instance to reuse cookies and login states. To enable this:

1. Open Chrome and navigate to `chrome://inspect/#remote-debugging`
2. Enable remote debugging (Chrome listens on `127.0.0.1:9222` by default)
3. Verify `ENABLE_CDP_MODE=True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) (lines 55-59)

### Execute Crawl Commands

Run a keyword search crawl on XiaoHongShu using:

```bash
uv run main.py --platform xhs --lt qrcode --type search

```

Fetch specific posts by ID using detail mode:

```bash
uv run main.py --platform xhs --lt qrcode --type detail

```

View all available options:

```bash
uv run main.py --help

```

The [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) module defines the CLI interface, supporting platforms (`xhs`, `douyin`, `bilibili`, `weibo`, `tieba`, `zhihu`), login types (`qrcode`, `phone`, etc.), and crawl types (`search`, `detail`).

### Alternative: Playwright Mode

If CDP is unavailable, disable it in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) by setting `ENABLE_CDP_MODE=False`, then install Playwright:

```bash
uv run playwright install

```

This runs a headless browser managed entirely by Playwright, though it may face higher detection rates compared to CDP mode.

## Optional WebUI Setup

MediaCrawler includes a Vue.js frontend with FastAPI backend for visual configuration.

### Start the FastAPI Backend

```bash
uv run uvicorn api.main:app --port 8080 --reload

```

The API server defined in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) exposes endpoints that the frontend consumes.

### Launch the Vue Frontend

```bash
cd webui
npm install
npm run dev

```

Access the interface at `http://localhost:5173/` to configure platforms, monitor logs, and export data without using the command line.

### Production Deployment

Build the frontend for static serving:

```bash
cd webui
npm install
npm run build
uv run uvicorn api.main:app --port 8080

```

This serves the pre-built UI from `api/webui/` alongside the API.

## Data Storage Configuration

Control output formats in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) via `SAVE_DATA_OPTION` (default: `jsonl` at line 89). Supported formats include:

- **File-based**: `csv`, `json`, `jsonl`, `excel`
- **Relational databases**: `sqlite`, `mysql`, `postgres`

When database options are selected, [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) automatically initializes the connection (lines 110-112). The `store/` directory contains platform-specific writers (e.g., [`store/excel_store_base.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/excel_store_base.py)), with Excel flushing occurring post-crawl via `_flush_excel_if_needed()` (lines 73-84).

Optional word-cloud generation activates when `ENABLE_GET_WORDCLOUD=True` (line 22) and data format is JSON/JSONL, triggered by `_generate_wordcloud_if_needed()` (lines 86-99).

## Summary

- **Install**: Use `git clone` followed by `uv sync` to establish the environment
- **Configure**: Enable CDP mode in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) (lines 55-59) and start Chrome with remote debugging on port 9222
- **Execute**: Run `uv run main.py` with `--platform`, `--lt`, and `--type` arguments to start the asynchronous crawl via [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) lines 99-122
- **Extend**: Access the WebUI by running `uvicorn api.main:app` and the Vue dev server for browser-based control
- **Store**: Modify `SAVE_DATA_OPTION` in base configuration to switch between JSONL, CSV, Excel, or SQL backends managed by modules in `store/`

## Frequently Asked Questions

### What is the difference between CDP mode and Playwright mode?

**CDP mode** (default) connects MediaCrawler to a locally running Chrome instance via the Chrome DevTools Protocol on port 9222, allowing the crawler to reuse your existing browser state, cookies, and extensions. This significantly reduces detection risk. **Playwright mode** launches a managed headless browser instance when `ENABLE_CDP_MODE=False`, requiring `uv run playwright install` but operating independently of your local Chrome installation.

### Which file controls the global crawler configuration?

The [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) file contains all global defaults including `ENABLE_CDP_MODE` (lines 55-59), `SAVE_DATA_OPTION` (line 89), and `ENABLE_GET_WORDCLOUD` (line 22). Changes to this file affect all crawl operations unless overridden by CLI arguments processed in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py).

### How do I add a new social media platform to MediaCrawler?

Create a new crawler class in `media_platform/<platform>.py` that inherits from `AbstractCrawler` (defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)). Implement the required async methods, then register the platform in the `CrawlerFactory` so that `CrawlerFactory.create_crawler()` can instantiate it when `--platform <your_platform>` is passed to [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py).

### Why does my crawl fail to start with a connection error?

Ensure Chrome is running with remote debugging enabled on `127.0.0.1:9222` (the default `CDP_DEBUG_PORT` specified in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)). Verify that `ENABLE_CDP_MODE=True` (lines 55-59) and that no firewall is blocking the local connection. If using Playwright mode instead, confirm you have executed `uv run playwright install` to download the necessary browser binaries.