# How to Set Up MediaCrawler on a Server: Complete Deployment Guide

> Learn how to set up MediaCrawler on your server with this complete deployment guide. Follow simple steps to clone the repo, install dependencies, configure settings, and launch the application for efficient media crawling.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-29

---

**TLDR:** Clone the repository, install the `uv` package manager and Google Chrome, enable `ENABLE_CDP_MODE = True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), launch Chrome with `--remote-debugging-port=9222`, then execute `uv run main.py` for command-line crawling or `uv run uvicorn api.main:app` to start the FastAPI service.

MediaCrawler is a modular Python crawler that extracts data from Chinese self-media platforms including Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. Deploying it on a server requires the **Chrome DevTools Protocol (CDP)** mode to avoid the overhead of running a full browser GUI while maintaining stealth against anti-bot detection. This guide walks through the production deployment using the actual source structure from `NanmiCoder/MediaCrawler`.

## Architecture Components

Understanding the core files helps troubleshoot server deployments.

- **[`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)** – CLI entry point that parses arguments and loads platform-specific crawler classes.
- **[`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)** – Abstract `BaseCrawler` class handling browser context, login hooks, and pagination logic.
- **[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)** – Centralized settings for platform selection (`PLATFORM`), CDP toggle (`ENABLE_CDP_MODE`), proxy configuration, and data storage format (`SAVE_DATA_OPTION`).
- **[`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py)** – FastAPI application exposing crawler functionality via HTTP endpoints and WebSocket.
- **[`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py)** – Factory selecting between `LocalCache` and `RedisCache` for login state persistence.
- **[`proxy/base_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/base_proxy.py)** – IP rotation logic supporting static lists, Wandou, and Kuaidaili providers.

## Prerequisites for Server Deployment

### System Requirements

- Linux server (Ubuntu 20.04+ or Debian 11+ recommended)
- Python 3.8+ (managed via `uv`)
- Node.js >= 16 (only if using the WebUI)

### Browser Environment Configuration

MediaCrawler requires a real browser environment. On servers without a display, **CDP mode is mandatory**. Install Google Chrome or Chromium, then launch it with remote debugging enabled:

```bash
google-chrome --remote-debugging-port=9222 --user-data-dir=/tmp/chrome-profile --no-sandbox --headless=new &

```

The `--user-data-dir` persists cookies, while `--remote-debugging-port=9222` allows the crawler to attach via CDP.

## Installation Steps

Execute these commands as a non-root user with sudo privileges:

```bash

# 1. Clone the repository

git clone https://github.com/NanmiCoder/MediaCrawler.git
cd MediaCrawler

# 2. Install uv (fast Python package manager)

curl -LsSf https://astral.sh/uv/install.sh | bash

# 3. Resolve dependencies exactly as specified in pyproject.toml

uv sync

# 4. (Optional) Install Playwright browsers as fallback

# Only needed if you disable CDP mode later

uv run playwright install chromium

```

## Server Configuration

Edit [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to optimize for headless server operation:

```python

# Enable CDP to connect to existing Chrome instance

ENABLE_CDP_MODE = True

# Select target platform: xhs, dy, ks, bili, wb, tieba, zhihu

PLATFORM = "xhs"

# QR code login recommended for servers (cookies persist via cache layer)

LOGIN_TYPE = "qrcode"

# Data storage backend: csv, json, jsonl, sqlite, mysql, etc.

SAVE_DATA_OPTION = "json"

# Persist login cookies between restarts

SAVE_LOGIN_STATE = True

```

## Running MediaCrawler on a Server

### Method 1: Command Line Execution

Ensure Chrome is running with remote debugging (see Prerequisites), then:

```bash
uv run main.py --platform xhs --lt qrcode --type search

```

Available `CRAWLER_TYPE` values include `search`, `detail`, and `creator`, defined in the platform implementations that extend `BaseCrawler`.

### Method 2: FastAPI Service Mode

Expose the crawler via HTTP for integration with other services:

```bash
uv run uvicorn api.main:app --host 0.0.0.0 --port 8080 --reload

```

The API in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) provides endpoints for triggering crawls, checking health status, and streaming results. It also serves WebSocket connections for real-time progress updates.

### Method 3: WebUI Deployment

For browser-based management without CLI access:

```bash
cd webui
npm install
npm run dev

```

The WebUI proxies API requests to the FastAPI backend. The Vite configuration in [`webui/vite.config.ts`](https://github.com/NanmiCoder/MediaCrawler/blob/main/webui/vite.config.ts) defines the development server routing.

## Production Reverse Proxy with Nginx

Place the FastAPI service behind Nginx for SSL termination and load balancing:

```nginx
server {
    listen 80;
    server_name media.example.com;

    location / {
        proxy_pass http://127.0.0.1:8080;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

```

## Summary

- **CDP mode is essential** for server deployments to avoid GUI dependencies; configure via `ENABLE_CDP_MODE = True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py).
- **Chrome must run separately** with `--remote-debugging-port=9222` before starting the crawler, allowing [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) to attach via the Chrome DevTools Protocol.
- **Three execution modes** are available: direct CLI via [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py), HTTP API via [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py), or visual management through the `webui/` Vue frontend.
- **State persistence** relies on the cache layer ([`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py)) to store login cookies between sessions, preventing repeated QR code scans.
- **Reverse proxy** configuration with Nginx enables secure public access to the FastAPI endpoints.

## Frequently Asked Questions

### What is the difference between CDP and Playwright mode for server deployment?

**CDP (Chrome DevTools Protocol)** attaches to an existing Chrome instance that you manually launch with remote debugging flags, making it ideal for servers because it reduces memory overhead and detection risk. **Playwright mode** launches a fresh browser instance programmatically, which requires more resources and is harder to run on headless servers without display drivers. According to the [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) implementation, CDP mode reuses cookies from the Chrome user data directory, maintaining session persistence across restarts.

### How do I handle login on a server without a graphical interface?

Use `LOGIN_TYPE = "qrcode"` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and scan the QR code output in the terminal logs during the first run. The `cache/` layer (`LocalCache` or `RedisCache`) persists the resulting cookies if `SAVE_LOGIN_STATE = True` is set. For subsequent runs, the crawler automatically loads these cookies from the cache implementation defined in [`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py), bypassing the need for manual login.

### Can I use Redis instead of local file caching for distributed deployments?

Yes. Modify [`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py) to return `RedisCache` instead of `LocalCache`. The factory pattern in this file instantiates the cache backend used by `BaseCrawler` for login state and rate-limit data. Redis is recommended for production server clusters where multiple crawler instances need shared session state.

### How do I configure IP rotation for anti-detection on a server?

Set up a proxy provider in [`proxy/base_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/base_proxy.py). The repository supports static proxy lists, Wandou, and Kuaidaili pools. Configure your chosen method in [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py), and the mixin classes will automatically inject proxy settings into all HTTP requests made by the platform crawlers. For high-volume crawling, combine this with the `asyncio` concurrency settings also defined in the base configuration.