MediaCrawler Command Line: A Complete CLI Guide for NanmiCoder/MediaCrawler
Run MediaCrawler directly from the terminal using python main.py with flags like --platform, --keyword, and --output, or use python -m MediaCrawler.main as a module.
MediaCrawler is an open-source Python framework for extracting data from Chinese social media platforms. This guide covers exactly how to execute crawls from the command line, including all available flags, configuration patterns, and the internal execution flow that makes the CLI work.
How MediaCrawler's CLI Architecture Works
The command-line interface follows a three-layer design that separates argument parsing, configuration loading, and platform-specific execution.
cmd_arg/arg.py defines the complete argparse schema. This file contains flag definitions for --platform, --keyword, --output, --proxy, and authentication parameters. When you type a command, this module validates your input and builds a namespace object.
main.py serves as the orchestration layer. It imports the parsed arguments, dynamically loads platform-specific configurations from files like config/zhihu_config.py or config/weibo_config.py, and instantiates the correct crawler class (e.g., ZhihuCrawler, WeiboCrawler).
base/base_crawler.py provides the abstract foundation. All platform crawlers inherit from this class, which implements common workflows for initialization, HTTP request handling via tools/httpx_util.py, and data persistence to MongoDB, Redis, or local files.
This modular structure means you can extend the CLI with new platforms or custom arguments without modifying the core entry point.
Installing and Running MediaCrawler from the Command Line
Prerequisites and Installation
Before running any crawl commands, install the Python dependencies:
pip install -r requirements.txt
Basic Command Structure
All MediaCrawler executions follow this pattern:
python main.py --platform <PLATFORM> [OPTIONS]
Or run as a module from the repository root:
python -m MediaCrawler.main --platform <PLATFORM> [OPTIONS]
Essential CLI Flags
| Flag | Required | Description | Example |
|---|---|---|---|
--platform |
Yes | Target platform identifier | zhihu, weibo, douyin, bilibili, tieba, kuaishou, xiaohongshu |
--keyword |
Context-dependent | Search term for keyword-based platforms | "人工智能" |
--user-id |
Context-dependent | Numeric user ID for profile crawls | 123456789 |
--output |
No | Output file path and format | results.xlsx, data.json |
--proxy |
No | HTTP/HTTPS proxy URL | http://127.0.0.1:7890 |
--cache |
No | Cache backend selection | redis or local |
Practical MediaCrawler Command Examples
Example 1: Zhihu Keyword Search
Search Zhihu for a topic and export to Excel:
python -m MediaCrawler.main \
--platform zhihu \
--keyword "人工智能" \
--output results.xlsx
Example 2: Weibo User Timeline with Proxy and Redis
Crawl a specific Weibo user's posts using a proxy and distributed caching:
python -m MediaCrawler.main \
--platform weibo \
--user-id 123456789 \
--proxy http://127.0.0.1:7890 \
--cache redis \
--output weibo.json
Example 3: Discover All Available Options
View the complete argument reference:
python -m MediaCrawler.main --help
Platform-Specific Command Patterns
Keyword-Based Platforms (Zhihu, Tieba, Bilibili Search)
These platforms support the --keyword flag:
python main.py --platform tieba --keyword "Python教程" --output tieba.xlsx
python main.py --platform bilibili --keyword "机器学习" --output videos.json
User-Centric Platforms (Weibo, Douyin, Xiaohongshu)
These require --user-id instead of --keyword:
python main.py --platform douyin --user-id 987654321 --output douyin_videos.xlsx
python main.py --platform xiaohongshu --user-id 555666777 --proxy http://proxy:8080
Kuaishou and Multi-Mode Platforms
Some platforms support both search and user modes depending on which flags you provide. Check config/<platform>_config.py for platform-specific capabilities.
Customizing MediaCrawler CLI Behavior
Adding Proxy Support
Pass --proxy with a valid HTTP or HTTPS URL to route all requests through an intermediary. This helps bypass rate limits and geo-restrictions:
python main.py --platform weibo --keyword "新闻" --proxy socks5://127.0.0.1:1080
Configuring Cache Backends
The --cache flag accepts two values:
local– File-based caching (default, no external dependencies)redis– Distributed caching using the connection parameters from your config
Redis caching is implemented in cache/redis_cache.py, while local caching uses cache/local_cache.py.
Output Format Handling
MediaCrawler automatically detects format from file extension:
| Extension | Handler | Use Case |
|---|---|---|
.xlsx |
Excel writer | Analysis, sharing, Excel workflows |
.csv |
CSV writer | Universal compatibility, large datasets |
.json |
JSON writer | API integration, nested data structures |
Troubleshooting Common CLI Issues
"ModuleNotFoundError: No module named 'MediaCrawler'"
Run from the repository root directory, or use the direct file approach:
python main.py --platform zhihu --keyword "test"
Platform Config Not Found
Ensure config/<platform>_config.py exists. The platform name in --platform must match the config filename prefix exactly.
Authentication Failures
Some platforms require login credentials or cookies. Store these in the appropriate config/<platform>_config.py file or use environment variables as referenced in base/base_crawler.py.
Summary
- Entry point:
main.pyorchestrates the entire crawl execution - Argument parsing:
cmd_arg/arg.pyvalidates all CLI flags including--platform,--keyword,--user-id,--output,--proxy, and--cache - Platform selection: Dynamic loading of
config/<platform>_config.pyand corresponding crawler classes that inherit frombase/base_crawler.py - HTTP execution:
tools/httpx_util.pyhandles async requests with optional proxy support - Data persistence: Built-in exporters for Excel, CSV, and JSON formats
- Extensibility: Add new platforms by creating config files and crawler classes without modifying
main.py
Frequently Asked Questions
What is the minimum command to run MediaCrawler?
python main.py --platform zhihu --keyword "example"
This executes a basic crawl with default settings. The --platform flag triggers dynamic loading of the corresponding config and crawler class from config/zhihu_config.py.
Why does my crawl fail with "platform not found"?
The --platform value must exactly match an available config file in the config/ directory. Valid options include zhihu, weibo, douyin, bilibili, tieba, kuaishou, and xiaohongshu as defined in cmd_arg/arg.py.
How do I use Redis caching from the command line?
Add --cache redis to any command. Ensure your Redis connection parameters are configured in the appropriate config/<platform>_config.py file. The cache/redis_cache.py module handles the actual Redis operations.
Can I run MediaCrawler without installing dependencies?
No. The requirements.txt file lists mandatory dependencies including httpx for HTTP requests and platform-specific packages. Install with pip install -r requirements.txt before first use.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →