How to Install MediaCrawler and Its Dependencies: Complete Setup Guide

To install MediaCrawler, clone the repository, install uv and Node.js prerequisites, run uv sync to resolve Python dependencies, and optionally install Playwright browsers with uv run playwright install unless using CDP mode.

MediaCrawler is a multi-platform media content scraper that supports crawling data from Xiaohongshu, Douyin, Zhihu, and other platforms. To install MediaCrawler correctly, you need to configure the Python environment using uv, install Node.js for platform-specific signing scripts, and set up browser automation dependencies. This guide walks through the complete installation process based on the repository's source code structure and official documentation.

Prerequisites: uv, Node.js, and Chrome

Before syncing the Python environment, install the required system tools.

Install uv (Python Package Manager)

uv is the recommended Python package manager for MediaCrawler. It provides fast dependency resolution and deterministic environments according to the pyproject.toml configuration.

Install uv by following the official documentation at https://docs.astral.sh/uv/getting-started/installation.

Install Node.js (≥ 16)

Node.js is required for platform-specific signing scripts, particularly for Douyin and Zhihu crawling operations that rely on cryptographic utilities.

Download Node.js from https://nodejs.org/en/download/.

Install Google Chrome (Optional for CDP Mode)

If you plan to use CDP mode (Chrome DevTools Protocol), install Google Chrome version 144 or higher. In this mode, MediaCrawler connects to an existing Chrome instance via tools/cdp_browser.py rather than launching its own browser.

Download Chrome from https://www.google.com/chrome/.

Clone the Repository and Sync Dependencies

Once prerequisites are installed, clone the repository and synchronize the Python environment.

git clone https://github.com/NanmiCoder/MediaCrawler.git
cd MediaCrawler

Verify uv is available:

uv --version

Sync the Python environment using uv sync. This command reads pyproject.toml and creates a virtual environment under .venv/ while installing all required packages, including internal modules like cache/, tools/, and platform models such as model/m_xiaohongshu.py referenced in main.py.

uv sync

Install Playwright Browser Binaries (Optional)

Playwright browser installation depends on your configuration in config/base_config.py.

If ENABLE_CDP_MODE = True (the default), MediaCrawler uses tools/cdp_browser.py to connect to your existing Chrome instance. You can skip Playwright installation.

If you disable CDP mode by setting ENABLE_CDP_MODE = False in config/base_config.py, MediaCrawler falls back to tools/browser_launcher.py to launch its own browser instances. Install the required browsers:

uv run playwright install

This downloads compatible Chromium, Firefox, and WebKit binaries that Playwright needs for browser automation.

Verify the Installation

Confirm the CLI is properly wired by displaying the help text:

uv run main.py --help

You should see usage information confirming that the entry point in main.py is functioning correctly and can dispatch to platform-specific crawlers.

Alternative: Using pip and venv

If you cannot use uv, MediaCrawler supports traditional virtual environments using requirements.txt. Create a venv and install dependencies:

python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

pip install -r requirements.txt
playwright install
python main.py --help

Note that requirements.txt serves as the legacy dependency list, while pyproject.toml remains the primary source for uv users.

Summary

  • Install prerequisites: uv, Node.js ≥ 16, and optionally Chrome ≥ 144 for CDP mode.
  • Sync environment: Run uv sync after cloning to install dependencies from pyproject.toml into a .venv/ directory.
  • Configure browser mode: Set ENABLE_CDP_MODE in config/base_config.py to determine whether to use an existing Chrome instance via tools/cdp_browser.py or install Playwright browsers for tools/browser_launcher.py.
  • Verify setup: Test with uv run main.py --help before executing crawl commands.
  • Alternative available: Use pip install -r requirements.txt if uv is unavailable in your environment.

Frequently Asked Questions

What is the difference between CDP mode and standard Playwright mode?

CDP mode (Chrome DevTools Protocol) connects MediaCrawler to an existing Chrome instance using tools/cdp_browser.py, requiring no additional browser downloads. Standard Playwright mode launches headless browsers via tools/browser_launcher.py and requires running uv run playwright install to download Chromium, Firefox, and WebKit binaries. CDP mode is enabled by default in config/base_config.py with ENABLE_CDP_MODE = True.

Can I install MediaCrawler without using uv?

Yes. While uv is the recommended package manager for deterministic environments, you can use traditional Python virtual environments. Create a venv with python -m venv .venv, activate it, and run pip install -r requirements.txt followed by playwright install. This mirrors the uv workflow for environments where uv cannot be installed.

Why is Node.js required for a Python project?

Node.js is required because MediaCrawler uses platform-specific signing scripts for certain sites like Douyin and Zhihu. These scripts handle cryptographic signatures and authentication flows that must execute in a Node.js environment, as implemented in the repository's auxiliary tools.

Where are the platform-specific configurations stored?

Platform configurations reside in config/<platform>_config.py files (e.g., config/xhs_config.py for Xiaohongshu). These files define keyword lists, credentials, and crawling parameters. Global settings like ENABLE_CDP_MODE are controlled in config/base_config.py, which determines whether the crawler uses CDP connections or launches standalone browsers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →