What Crawler Types Does MediaCrawler's CLI Support?

MediaCrawler's command-line interface currently supports only one crawler type: search, which performs keyword-based content discovery across supported social media platforms.

MediaCrawler is an open-source scraping framework designed for extracting data from Chinese social media platforms like Weibo, Tieba, and others. When invoking the tool via CLI, you control the crawling behavior through the --crawler-type argument. According to the source code in the NanmiCoder/MediaCrawler repository, this parameter maps to the CrawlerTypeEnum enumeration, which currently defines search as the sole available operational mode.

The CrawlerTypeEnum Definition

The supported crawler types are defined in the CrawlerTypeEnum class located in [cmd_arg/arg.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py).

At line 63 of this file, the enumeration contains a single entry:

class CrawlerTypeEnum(Enum):
    search = "search"

This enumeration serves as the source of truth for all valid values passed to the --crawler-type (or -t) CLI option. The search value represents the standard keyword-based crawling mode that queries platforms for content matching specified search terms.

How the Search Crawler Type Works

The search crawler type operates as the default operational mode for MediaCrawler. When activated, it executes keyword-based searches across the configured platform, retrieving posts, comments, and metadata that match the keywords provided via the --keyword argument.

If you omit the --crawler-type parameter entirely, the CLI automatically falls back to search mode. This default assignment occurs at line 184 in cmd_arg/arg.py, where the argument parser initializes the crawler type to "search" when no explicit value is provided.

Specifying Crawler Types in CLI Commands

You can explicitly declare the crawler type using either the short or long flag format. While omitting the parameter triggers the default search behavior, explicit declaration ensures clarity in scripts and automation workflows.


# Use the short flag (-t)

media_crawler -p weibo -t search -k "open source"

# Use the long flag (--crawler-type)

media_crawler --platform weibo --crawler-type search --keyword "open source"

# Omitting the flag defaults to search automatically

media_crawler -p weibo -k "open source"

All three commands above execute identical search operations against the Weibo platform.

API Schema Validation

The crawler type definition extends beyond the CLI into the API layer. The file [api/schemas/crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/crawler.py) mirrors the CrawlerTypeEnum for payload validation in API requests.

This schema ensures that HTTP requests to the MediaCrawler API also default to search mode when no type is specified, maintaining consistency between command-line and programmatic interfaces.

Testing the Crawler Type Configuration

The project includes unit tests that validate the crawler type initialization. In [tests/test_cmd_arg_tieba.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_cmd_arg_tieba.py), test cases instantiate crawlers using the default search type, verifying that the argument parser correctly handles the enumeration and applies the default value when processing Tieba-specific crawl configurations.

Summary

  • MediaCrawler's CLI supports only the search crawler type as defined in cmd_arg/arg.py at line 63.
  • search is the default mode when --crawler-type is omitted, initialized at line 184 in the same file.
  • Use -t or --crawler-type followed by search to explicitly declare the operational mode.
  • The API schema in api/schemas/crawler.py mirrors this enumeration for HTTP request validation.
  • Unit tests in tests/test_cmd_arg_tieba.py verify the default search type behavior.

Frequently Asked Questions

What is the default crawler type in MediaCrawler?

The default crawler type is search. If you run the CLI without specifying the --crawler-type or -t flag, MediaCrawler automatically defaults to search mode according to the argument parsing logic at line 184 in cmd_arg/arg.py.

How do I specify the crawler type when running MediaCrawler?

Use the --crawler-type flag (or -t for short) followed by the desired type. Currently, only search is valid: media_crawler -p weibo -t search -k "keyword". This flag connects to the CrawlerTypeEnum defined in the CLI argument parser.

Are there other crawler types besides search in MediaCrawler?

No. As of the current implementation in the main branch, the CrawlerTypeEnum in cmd_arg/arg.py contains only the search entry. The maintainers may expand this enumeration in future releases to support additional crawling strategies like user-profile crawling or trending-topic discovery.

Where is the crawler type validation defined in the source code?

The primary definition resides in cmd_arg/arg.py within the CrawlerTypeEnum class. Additionally, the API layer validates crawler types through api/schemas/crawler.py, which duplicates the enumeration to ensure HTTP payloads contain valid operational modes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →