How to Use MediaCrawler for Scraping Data from Bilibili: A Complete Technical Guide
MediaCrawler provides an asynchronous, Playwright-based framework for extracting Bilibili videos, comments, and creator profiles with support for multiple storage backends including CSV, MongoDB, SQLite, and Excel.
MediaCrawler is an open-source Python scraping framework designed for Chinese social media platforms. When configured for Bilibili, it orchestrates browser automation to handle authentication, rate limiting, and data extraction while offering both command-line and programmatic APIs for flexible deployment.
Understanding the Bilibili Crawling Architecture
The repository follows a modular pipeline architecture that separates concerns between configuration, HTTP client operations, and data persistence.
Core Components
- Command-line parser (
cmd_arg/arg.py): Theparse_cmdfunction converts CLI flags like--platform biliand--type searchinto a configuration object that updates global settings. - Configuration module (
config/bilibili_config.py): Contains Bilibili-specific defaults includingBILI_SPECIFIED_ID_LISTfor target videos, search mode parameters, and video quality settings. - Crawler core (
media_platform/bilibili/core.py): TheBilibiliCrawlerclass orchestrates the browser lifecycle throughBilibiliCrawler.start(), manages login state, and dispatches to specific crawling modes. - HTTP client (
media_platform/bilibili/client.py): TheBilibiliClientclass provides low-level API methods includingsearch_video_by_keyword,get_video_info, andget_video_all_comments. - Storage implementations (
store/bilibili/_store_impl.py): Concrete classes likeBiliCsvStoreImplement,BiliDbStoreImplement, andBiliMongoStoreImplementhandle persistence logic for different backends.
Execution Flow
- Initialization: CLI arguments are parsed by
parse_cmdincmd_arg/arg.pyto populate the globalconfigobject. - Browser Launch:
BilibiliCrawler.start()initializes a Playwright browser instance (Chromium by default) and handles authentication via QR-code, phone, or cookie login. - Mode Dispatch: Based on
config.CRAWLER_TYPE, the crawler executes one of three paths:- Search mode: Calls
search()→search_by_keywords()to crawl search result pages. - Detail mode: Calls
get_specified_videos()to fetch metadata for specific BV IDs or URLs. - Creator mode: Calls
get_all_creator_details()to extract fan counts, followings, and dynamic posts.
- Search mode: Calls
- Data Extraction: For each video, the crawler optionally downloads media files via
BilibiliVideo.store_videoinstore/bilibili/bilibilli_store_media.pyand extracts comments throughBilibiliClient.get_video_all_comments. - Persistence: Results are routed to the storage backend selected via
--save_data_option, with implementations instore/bilibili/_store_impl.pyhandling async writes.
Installation and Prerequisites
Install the required dependencies and browser binaries before running Bilibili scraping tasks:
pip install -r requirements.txt
playwright install chromium
The requirements.txt includes essential packages such as playwright, aiofiles, sqlalchemy, and pymongo that support the async architecture and multiple storage backends.
Command-Line Usage Patterns
MediaCrawler supports three primary crawling modes for Bilibili, controlled by the --type parameter.
Search Videos by Keyword
Use search mode to discover videos based on keywords and time ranges:
python -m MediaCrawler \
--platform bili \
--type search \
--keywords "gaming,anime" \
--save_data_option csv \
--headless true
This command respects the BILI_SEARCH_MODE setting in config/bilibili_config.py (defaulting to normal search) and saves video metadata, comments, and creator information to CSV files under data/bili/.
Crawl Specific Videos by ID or URL
Use detail mode when you have specific video identifiers:
python -m MediaCrawler \
--platform bili \
--type detail \
--specified_id "BV1dwuKzmE26,BV14Q4y1n7jz" \
--save_data_option mongodb
The crawler parses each entry through parse_video_info_from_url to extract the numeric bvid, then calls BilibiliClient.get_video_info to retrieve full metadata and comments.
Harvest Creator Profiles and Fan Data
Use creator mode to extract comprehensive channel statistics:
python -m MediaCrawler \
--platform bili \
--type creator \
--creator_id "20813884,https://space.bilibili.com/434377496" \
--save_data_option jsonl
This invokes get_all_creator_details() which iterates through creator pages and extracts fan/following counts via BilibiliClient.get_creator_all_fans, respecting the CREATOR_MODE configuration in config/bilibili_config.py.
Configuring Storage Backends
Select your persistence layer using the --save_data_option flag with one of the following values: csv, db, json, jsonl, sqlite, mongodb, or excel.
Each option maps to a specific implementation in store/bilibili/_store_impl.py:
- CSV:
BiliCsvStoreImplementwrites to flat files indata/bili/. - MongoDB:
BiliMongoStoreImplementstores documents in collections prefixed withbilibili_. - SQLite:
BiliDbStoreImplementuses SQLAlchemy for relational storage. - Excel: Generates
.xlsxfiles for business analysis workflows.
All storage classes inherit from AbstractStore defined in base/base_crawler.py and implement async methods: store_content(), store_comment(), and store_creator().
Programmatic API Integration
Beyond CLI usage, you can embed MediaCrawler directly in Python applications.
Running the Crawler from Python
Import the core components and simulate CLI arguments programmatically:
import asyncio
from cmd_arg.arg import parse_cmd
from media_platform.bilibili.core import BilibiliCrawler
async def scrape_bilibili():
# Simulate command-line arguments
args = await parse_cmd([
"--platform", "bili",
"--type", "search",
"--keywords", "anime,music",
"--save_data_option", "jsonl",
"--headless", "true"
])
crawler = BilibiliCrawler()
await crawler.start() # Handles browser init and login
await crawler.close()
asyncio.run(scrape_bilibili())
The parse_cmd function returns a SimpleNamespace that updates global configuration values, allowing BilibiliCrawler.start() to execute with the specified parameters.
Implementing Custom Storage Backends
To store data in a custom system like PostgreSQL, implement the AbstractStore interface:
from base.base_crawler import AbstractStore
class PostgresStoreImplement(AbstractStore):
async def store_content(self, content_item: dict):
# Custom INSERT logic for video metadata
pass
async def store_comment(self, comment_item: dict):
# Custom INSERT logic for comments
pass
async def store_creator(self, creator: dict):
# Custom INSERT logic for creator profiles
pass
Register your implementation by mapping the storage key in the platform initialization code, then invoke with --save_data_option postgres.
Key Configuration Files
| File | Purpose | Key Elements |
|---|---|---|
cmd_arg/arg.py |
CLI parsing and config injection | parse_cmd function |
config/bilibili_config.py |
Platform-specific constants | BILI_SPECIFIED_ID_LIST, BILI_SEARCH_MODE |
media_platform/bilibili/core.py |
Main orchestration logic | BilibiliCrawler class, start() method |
media_platform/bilibili/client.py |
API wrapper | BilibiliClient with search_video_by_keyword, get_video_info |
media_platform/bilibili/login.py |
Authentication flows | QR-code and cookie login handlers |
store/bilibili/_store_impl.py |
Storage backends | BiliCsvStoreImplement, BiliMongoStoreImplement |
Summary
- MediaCrawler uses Playwright to automate Chromium for Bilibili scraping, handling JavaScript rendering and authentication automatically.
- Three operational modes—search, detail, and creator—support keyword discovery, specific video extraction, and channel statistics harvesting respectively.
- The architecture separates concerns between
cmd_arg/arg.pyfor CLI parsing,media_platform/bilibili/core.pyfor orchestration, andstore/bilibili/_store_impl.pyfor persistence. - Storage is pluggable via
--save_data_option, supporting everything from CSV files to MongoDB with async batch writes. - Concurrency is controlled by
config.MAX_CONCURRENCY_NUM(default 10) through asyncio semaphores to prevent rate limiting.
Frequently Asked Questions
How does MediaCrawler handle Bilibili authentication?
MediaCrawler supports three authentication methods implemented in media_platform/bilibili/login.py: QR-code scanning, phone number login, and cookie-based sessions. When running BilibiliCrawler.start(), the system checks for existing cookies in config.COOKIES and falls back to interactive QR-code login if no valid session exists.
Can I scrape specific Bilibili videos by URL instead of BV ID?
Yes. The --specified_id parameter accepts both raw BV IDs (e.g., BV1dwuKzmE26) and full URLs (e.g., https://www.bilibili.com/video/BV1dwuKzmE26/). The parse_video_info_from_url function in the crawler core extracts the numeric identifier automatically before calling BilibiliClient.get_video_info.
What is the difference between JSON and JSONL storage options?
The JSON option writes all records to a single JSON array file, while JSONL (JSON Lines) writes one JSON object per line. JSONL is preferred for large scraping jobs because it allows streaming writes and prevents memory issues when processing millions of comments or video records.
How do I limit concurrent requests to avoid IP bans?
Control concurrency using the --max_concurrency_num flag or by setting config.MAX_CONCURRENCY_NUM in config/bilibili_config.py. The default value is 10 simultaneous requests, managed by asyncio semaphores in BilibiliCrawler. Reducing this to 3-5 for creator scraping mode is recommended when operating without proxy rotation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →