How to Set Up MediaCrawler on a Server: Complete Deployment Guide
TLDR: Clone the repository, install the uv package manager and Google Chrome, enable ENABLE_CDP_MODE = True in config/base_config.py, launch Chrome with --remote-debugging-port=9222, then execute uv run main.py for command-line crawling or uv run uvicorn api.main:app to start the FastAPI service.
MediaCrawler is a modular Python crawler that extracts data from Chinese self-media platforms including Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. Deploying it on a server requires the Chrome DevTools Protocol (CDP) mode to avoid the overhead of running a full browser GUI while maintaining stealth against anti-bot detection. This guide walks through the production deployment using the actual source structure from NanmiCoder/MediaCrawler.
Architecture Components
Understanding the core files helps troubleshoot server deployments.
main.py– CLI entry point that parses arguments and loads platform-specific crawler classes.base/base_crawler.py– AbstractBaseCrawlerclass handling browser context, login hooks, and pagination logic.config/base_config.py– Centralized settings for platform selection (PLATFORM), CDP toggle (ENABLE_CDP_MODE), proxy configuration, and data storage format (SAVE_DATA_OPTION).api/main.py– FastAPI application exposing crawler functionality via HTTP endpoints and WebSocket.cache/cache_factory.py– Factory selecting betweenLocalCacheandRedisCachefor login state persistence.proxy/base_proxy.py– IP rotation logic supporting static lists, Wandou, and Kuaidaili providers.
Prerequisites for Server Deployment
System Requirements
- Linux server (Ubuntu 20.04+ or Debian 11+ recommended)
- Python 3.8+ (managed via
uv) - Node.js >= 16 (only if using the WebUI)
Browser Environment Configuration
MediaCrawler requires a real browser environment. On servers without a display, CDP mode is mandatory. Install Google Chrome or Chromium, then launch it with remote debugging enabled:
google-chrome --remote-debugging-port=9222 --user-data-dir=/tmp/chrome-profile --no-sandbox --headless=new &
The --user-data-dir persists cookies, while --remote-debugging-port=9222 allows the crawler to attach via CDP.
Installation Steps
Execute these commands as a non-root user with sudo privileges:
# 1. Clone the repository
git clone https://github.com/NanmiCoder/MediaCrawler.git
cd MediaCrawler
# 2. Install uv (fast Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | bash
# 3. Resolve dependencies exactly as specified in pyproject.toml
uv sync
# 4. (Optional) Install Playwright browsers as fallback
# Only needed if you disable CDP mode later
uv run playwright install chromium
Server Configuration
Edit config/base_config.py to optimize for headless server operation:
# Enable CDP to connect to existing Chrome instance
ENABLE_CDP_MODE = True
# Select target platform: xhs, dy, ks, bili, wb, tieba, zhihu
PLATFORM = "xhs"
# QR code login recommended for servers (cookies persist via cache layer)
LOGIN_TYPE = "qrcode"
# Data storage backend: csv, json, jsonl, sqlite, mysql, etc.
SAVE_DATA_OPTION = "json"
# Persist login cookies between restarts
SAVE_LOGIN_STATE = True
Running MediaCrawler on a Server
Method 1: Command Line Execution
Ensure Chrome is running with remote debugging (see Prerequisites), then:
uv run main.py --platform xhs --lt qrcode --type search
Available CRAWLER_TYPE values include search, detail, and creator, defined in the platform implementations that extend BaseCrawler.
Method 2: FastAPI Service Mode
Expose the crawler via HTTP for integration with other services:
uv run uvicorn api.main:app --host 0.0.0.0 --port 8080 --reload
The API in api/main.py provides endpoints for triggering crawls, checking health status, and streaming results. It also serves WebSocket connections for real-time progress updates.
Method 3: WebUI Deployment
For browser-based management without CLI access:
cd webui
npm install
npm run dev
The WebUI proxies API requests to the FastAPI backend. The Vite configuration in webui/vite.config.ts defines the development server routing.
Production Reverse Proxy with Nginx
Place the FastAPI service behind Nginx for SSL termination and load balancing:
server {
listen 80;
server_name media.example.com;
location / {
proxy_pass http://127.0.0.1:8080;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
Summary
- CDP mode is essential for server deployments to avoid GUI dependencies; configure via
ENABLE_CDP_MODE = Trueinconfig/base_config.py. - Chrome must run separately with
--remote-debugging-port=9222before starting the crawler, allowingbase/base_crawler.pyto attach via the Chrome DevTools Protocol. - Three execution modes are available: direct CLI via
main.py, HTTP API viaapi/main.py, or visual management through thewebui/Vue frontend. - State persistence relies on the cache layer (
cache/cache_factory.py) to store login cookies between sessions, preventing repeated QR code scans. - Reverse proxy configuration with Nginx enables secure public access to the FastAPI endpoints.
Frequently Asked Questions
What is the difference between CDP and Playwright mode for server deployment?
CDP (Chrome DevTools Protocol) attaches to an existing Chrome instance that you manually launch with remote debugging flags, making it ideal for servers because it reduces memory overhead and detection risk. Playwright mode launches a fresh browser instance programmatically, which requires more resources and is harder to run on headless servers without display drivers. According to the base/base_crawler.py implementation, CDP mode reuses cookies from the Chrome user data directory, maintaining session persistence across restarts.
How do I handle login on a server without a graphical interface?
Use LOGIN_TYPE = "qrcode" in config/base_config.py and scan the QR code output in the terminal logs during the first run. The cache/ layer (LocalCache or RedisCache) persists the resulting cookies if SAVE_LOGIN_STATE = True is set. For subsequent runs, the crawler automatically loads these cookies from the cache implementation defined in cache/cache_factory.py, bypassing the need for manual login.
Can I use Redis instead of local file caching for distributed deployments?
Yes. Modify cache/cache_factory.py to return RedisCache instead of LocalCache. The factory pattern in this file instantiates the cache backend used by BaseCrawler for login state and rate-limit data. Redis is recommended for production server clusters where multiple crawler instances need shared session state.
How do I configure IP rotation for anti-detection on a server?
Set up a proxy provider in proxy/base_proxy.py. The repository supports static proxy lists, Wandou, and Kuaidaili pools. Configure your chosen method in base_config.py, and the mixin classes will automatically inject proxy settings into all HTTP requests made by the platform crawlers. For high-volume crawling, combine this with the asyncio concurrency settings also defined in the base configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →