What Programming Language Is MediaCrawler Written In?
MediaCrawler is implemented entirely in Python 3.11+, leveraging modern async libraries like Playwright and FastAPI for browser automation and data extraction.
The NanmiCoder/MediaCrawler repository is a comprehensive social media scraping framework designed for platforms like XiaoHongShu, Weibo, and Bilibili. Understanding what programming language MediaCrawler is written in reveals why it offers such robust asynchronous crawling capabilities and seamless integration with data science tooling.
Python as the Core Implementation Language
MediaCrawler contains no compiled language source files (e.g., .c, .go, or .js for core logic). The entire codebase consists of Python modules organized under standard package structures. This Python-centric design enables rapid development of platform-specific crawlers while maintaining clean abstraction layers through object-oriented patterns.
Version Requirements and Configuration
The project's pyproject.toml explicitly declares Python ≥3.11 as the minimum version in lines 7-8:
[project]
requires-python = ">=3.11"
This requirement ensures access to modern Python features like improved asyncio exception groups and enhanced typing support, which the crawler relies on for concurrent browser automation tasks.
Python Ecosystem and Key Dependencies
According to the dependencies listed in pyproject.toml, MediaCrawler utilizes a modern Python tooling stack:
- Playwright (async API) — Handles browser automation for social media login flows and dynamic content extraction
- FastAPI — Powers the optional HTTP API server for exposing crawler functionality via REST endpoints
- SQLAlchemy — Provides the ORM layer for persisting scraped data to databases
- Pandas / OpenPyXL — Enables Excel and CSV export utilities for data analysis workflows
- uv — Serves as the fast dependency resolver and package manager (referenced in README installation instructions)
- Typer — Drives CLI argument parsing in the
cmd_argmodule
All third-party integrations are Python packages installable via standard pip or uv workflows, confirming the project's commitment to the Python ecosystem.
Architecture and Source File Organization
The repository structure demonstrates idiomatic Python package organization with clear separation of concerns:
main.py— Contains the CLI entry point and crawler orchestration logicbase/base_crawler.py(lines 20-27) — Defines abstract base classes establishing the crawler contract that all platform implementations must followmedia_platform/— Houses platform-specific implementations such asxhs.py(XiaoHongShu),weibo.py, andbilibilicrawlersconfig/— Stores runtime configuration files controlling crawl parameterstools/— Provides helper utilities including async file writers and Chrome DevTools Protocol (CDP) browser wrappers
Command Line Interface Usage
You interact with MediaCrawler through Python execution commands using uv:
# Install dependencies via uv
uv sync
# Execute search on XiaoHongShu platform
uv run main.py --platform xhs --lt qrcode --type search
# Crawl specific post by ID in detail mode
uv run main.py --platform xhs --lt qrcode --type detail
Programmatic Python API
For custom automation scripts, import the crawler classes directly:
from media_platform.xhs import XiaoHongShuCrawler
import asyncio
async def run():
crawler = XiaoHongShuCrawler()
await crawler.start() # Starts the complete crawl workflow
asyncio.run(run())
Factory Pattern Implementation
The CrawlerFactory in main.py enables dynamic crawler instantiation:
from main import CrawlerFactory
# Create Bilibili crawler instance
crawler = CrawlerFactory.create_crawler("bili")
await crawler.start()
Summary
- MediaCrawler is written exclusively in Python, requiring version 3.11 or higher as specified in
pyproject.toml. - The project leverages async/await patterns throughout, utilizing libraries like Playwright for non-blocking browser automation.
- All entry points—including the CLI (
main.py), abstract base classes (base/base_crawler.py), and platform modules (media_platform/*)—are pure Python implementations. - The absence of compiled extensions means the entire codebase is inspectable and modifiable using standard Python development workflows.
Frequently Asked Questions
Is MediaCrawler written in JavaScript or TypeScript?
No. While MediaCrawler uses Playwright to control browser instances (which involves JavaScript execution within pages), all core logic, orchestration, and data processing are implemented in Python. The browser automation is handled through Playwright's Python async API, not Node.js.
What is the minimum Python version required for MediaCrawler?
MediaCrawler requires Python 3.11 or higher. This is enforced in the requires-python field of pyproject.toml (lines 7-8) to ensure compatibility with modern asyncio features and type hinting syntax used throughout the codebase.
Does MediaCrawler use compiled extensions for performance?
No. MediaCrawler relies on Python's native async/await concurrency rather than C extensions or compiled modules for its core functionality. Performance for I/O-bound operations comes from asyncio and Playwright's efficient browser management, not from compiled Python extensions.
Can I extend MediaCrawler with my own Python modules?
Yes. You can create custom crawlers by implementing the abstract base classes defined in base/base_crawler.py (specifically lines 20-27 where the crawler contract is established). New platform modules placed in the media_platform/ directory and registered through the CrawlerFactory in main.py integrate seamlessly with the existing Python architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →