MediaCrawler Command Line: A Complete CLI Guide for NanmiCoder/MediaCrawler

Run MediaCrawler directly from the terminal using python main.py with flags like --platform, --keyword, and --output, or use python -m MediaCrawler.main as a module.

MediaCrawler is an open-source Python framework for extracting data from Chinese social media platforms. This guide covers exactly how to execute crawls from the command line, including all available flags, configuration patterns, and the internal execution flow that makes the CLI work.

How MediaCrawler's CLI Architecture Works

The command-line interface follows a three-layer design that separates argument parsing, configuration loading, and platform-specific execution.

cmd_arg/arg.py defines the complete argparse schema. This file contains flag definitions for --platform, --keyword, --output, --proxy, and authentication parameters. When you type a command, this module validates your input and builds a namespace object.

main.py serves as the orchestration layer. It imports the parsed arguments, dynamically loads platform-specific configurations from files like config/zhihu_config.py or config/weibo_config.py, and instantiates the correct crawler class (e.g., ZhihuCrawler, WeiboCrawler).

base/base_crawler.py provides the abstract foundation. All platform crawlers inherit from this class, which implements common workflows for initialization, HTTP request handling via tools/httpx_util.py, and data persistence to MongoDB, Redis, or local files.

This modular structure means you can extend the CLI with new platforms or custom arguments without modifying the core entry point.

Installing and Running MediaCrawler from the Command Line

Prerequisites and Installation

Before running any crawl commands, install the Python dependencies:

pip install -r requirements.txt

Basic Command Structure

All MediaCrawler executions follow this pattern:

python main.py --platform <PLATFORM> [OPTIONS]

Or run as a module from the repository root:

python -m MediaCrawler.main --platform <PLATFORM> [OPTIONS]

Essential CLI Flags

Flag Required Description Example
--platform Yes Target platform identifier zhihu, weibo, douyin, bilibili, tieba, kuaishou, xiaohongshu
--keyword Context-dependent Search term for keyword-based platforms "人工智能"
--user-id Context-dependent Numeric user ID for profile crawls 123456789
--output No Output file path and format results.xlsx, data.json
--proxy No HTTP/HTTPS proxy URL http://127.0.0.1:7890
--cache No Cache backend selection redis or local

Practical MediaCrawler Command Examples

Search Zhihu for a topic and export to Excel:

python -m MediaCrawler.main \
    --platform zhihu \
    --keyword "人工智能" \
    --output results.xlsx

Example 2: Weibo User Timeline with Proxy and Redis

Crawl a specific Weibo user's posts using a proxy and distributed caching:

python -m MediaCrawler.main \
    --platform weibo \
    --user-id 123456789 \
    --proxy http://127.0.0.1:7890 \
    --cache redis \
    --output weibo.json

Example 3: Discover All Available Options

View the complete argument reference:

python -m MediaCrawler.main --help

Platform-Specific Command Patterns

These platforms support the --keyword flag:

python main.py --platform tieba --keyword "Python教程" --output tieba.xlsx
python main.py --platform bilibili --keyword "机器学习" --output videos.json

User-Centric Platforms (Weibo, Douyin, Xiaohongshu)

These require --user-id instead of --keyword:

python main.py --platform douyin --user-id 987654321 --output douyin_videos.xlsx
python main.py --platform xiaohongshu --user-id 555666777 --proxy http://proxy:8080

Kuaishou and Multi-Mode Platforms

Some platforms support both search and user modes depending on which flags you provide. Check config/<platform>_config.py for platform-specific capabilities.

Customizing MediaCrawler CLI Behavior

Adding Proxy Support

Pass --proxy with a valid HTTP or HTTPS URL to route all requests through an intermediary. This helps bypass rate limits and geo-restrictions:

python main.py --platform weibo --keyword "新闻" --proxy socks5://127.0.0.1:1080

Configuring Cache Backends

The --cache flag accepts two values:

  • local – File-based caching (default, no external dependencies)
  • redis – Distributed caching using the connection parameters from your config

Redis caching is implemented in cache/redis_cache.py, while local caching uses cache/local_cache.py.

Output Format Handling

MediaCrawler automatically detects format from file extension:

Extension Handler Use Case
.xlsx Excel writer Analysis, sharing, Excel workflows
.csv CSV writer Universal compatibility, large datasets
.json JSON writer API integration, nested data structures

Troubleshooting Common CLI Issues

"ModuleNotFoundError: No module named 'MediaCrawler'"

Run from the repository root directory, or use the direct file approach:

python main.py --platform zhihu --keyword "test"

Platform Config Not Found

Ensure config/<platform>_config.py exists. The platform name in --platform must match the config filename prefix exactly.

Authentication Failures

Some platforms require login credentials or cookies. Store these in the appropriate config/<platform>_config.py file or use environment variables as referenced in base/base_crawler.py.

Summary

  • Entry point: main.py orchestrates the entire crawl execution
  • Argument parsing: cmd_arg/arg.py validates all CLI flags including --platform, --keyword, --user-id, --output, --proxy, and --cache
  • Platform selection: Dynamic loading of config/<platform>_config.py and corresponding crawler classes that inherit from base/base_crawler.py
  • HTTP execution: tools/httpx_util.py handles async requests with optional proxy support
  • Data persistence: Built-in exporters for Excel, CSV, and JSON formats
  • Extensibility: Add new platforms by creating config files and crawler classes without modifying main.py

Frequently Asked Questions

What is the minimum command to run MediaCrawler?

python main.py --platform zhihu --keyword "example"

This executes a basic crawl with default settings. The --platform flag triggers dynamic loading of the corresponding config and crawler class from config/zhihu_config.py.

Why does my crawl fail with "platform not found"?

The --platform value must exactly match an available config file in the config/ directory. Valid options include zhihu, weibo, douyin, bilibili, tieba, kuaishou, and xiaohongshu as defined in cmd_arg/arg.py.

How do I use Redis caching from the command line?

Add --cache redis to any command. Ensure your Redis connection parameters are configured in the appropriate config/<platform>_config.py file. The cache/redis_cache.py module handles the actual Redis operations.

Can I run MediaCrawler without installing dependencies?

No. The requirements.txt file lists mandatory dependencies including httpx for HTTP requests and platform-specific packages. Install with pip install -r requirements.txt before first use.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →