How to Use MediaCrawler for Scraping Data from Bilibili: A Complete Technical Guide

MediaCrawler provides an asynchronous, Playwright-based framework for extracting Bilibili videos, comments, and creator profiles with support for multiple storage backends including CSV, MongoDB, SQLite, and Excel.

MediaCrawler is an open-source Python scraping framework designed for Chinese social media platforms. When configured for Bilibili, it orchestrates browser automation to handle authentication, rate limiting, and data extraction while offering both command-line and programmatic APIs for flexible deployment.

Understanding the Bilibili Crawling Architecture

The repository follows a modular pipeline architecture that separates concerns between configuration, HTTP client operations, and data persistence.

Core Components

  • Command-line parser (cmd_arg/arg.py): The parse_cmd function converts CLI flags like --platform bili and --type search into a configuration object that updates global settings.
  • Configuration module (config/bilibili_config.py): Contains Bilibili-specific defaults including BILI_SPECIFIED_ID_LIST for target videos, search mode parameters, and video quality settings.
  • Crawler core (media_platform/bilibili/core.py): The BilibiliCrawler class orchestrates the browser lifecycle through BilibiliCrawler.start(), manages login state, and dispatches to specific crawling modes.
  • HTTP client (media_platform/bilibili/client.py): The BilibiliClient class provides low-level API methods including search_video_by_keyword, get_video_info, and get_video_all_comments.
  • Storage implementations (store/bilibili/_store_impl.py): Concrete classes like BiliCsvStoreImplement, BiliDbStoreImplement, and BiliMongoStoreImplement handle persistence logic for different backends.

Execution Flow

  1. Initialization: CLI arguments are parsed by parse_cmd in cmd_arg/arg.py to populate the global config object.
  2. Browser Launch: BilibiliCrawler.start() initializes a Playwright browser instance (Chromium by default) and handles authentication via QR-code, phone, or cookie login.
  3. Mode Dispatch: Based on config.CRAWLER_TYPE, the crawler executes one of three paths:
    • Search mode: Calls search() → search_by_keywords() to crawl search result pages.
    • Detail mode: Calls get_specified_videos() to fetch metadata for specific BV IDs or URLs.
    • Creator mode: Calls get_all_creator_details() to extract fan counts, followings, and dynamic posts.
  4. Data Extraction: For each video, the crawler optionally downloads media files via BilibiliVideo.store_video in store/bilibili/bilibilli_store_media.py and extracts comments through BilibiliClient.get_video_all_comments.
  5. Persistence: Results are routed to the storage backend selected via --save_data_option, with implementations in store/bilibili/_store_impl.py handling async writes.

Installation and Prerequisites

Install the required dependencies and browser binaries before running Bilibili scraping tasks:

pip install -r requirements.txt
playwright install chromium

The requirements.txt includes essential packages such as playwright, aiofiles, sqlalchemy, and pymongo that support the async architecture and multiple storage backends.

Command-Line Usage Patterns

MediaCrawler supports three primary crawling modes for Bilibili, controlled by the --type parameter.

Search Videos by Keyword

Use search mode to discover videos based on keywords and time ranges:

python -m MediaCrawler \
  --platform bili \
  --type search \
  --keywords "gaming,anime" \
  --save_data_option csv \
  --headless true

This command respects the BILI_SEARCH_MODE setting in config/bilibili_config.py (defaulting to normal search) and saves video metadata, comments, and creator information to CSV files under data/bili/.

Crawl Specific Videos by ID or URL

Use detail mode when you have specific video identifiers:

python -m MediaCrawler \
  --platform bili \
  --type detail \
  --specified_id "BV1dwuKzmE26,BV14Q4y1n7jz" \
  --save_data_option mongodb

The crawler parses each entry through parse_video_info_from_url to extract the numeric bvid, then calls BilibiliClient.get_video_info to retrieve full metadata and comments.

Harvest Creator Profiles and Fan Data

Use creator mode to extract comprehensive channel statistics:

python -m MediaCrawler \
  --platform bili \
  --type creator \
  --creator_id "20813884,https://space.bilibili.com/434377496" \
  --save_data_option jsonl

This invokes get_all_creator_details() which iterates through creator pages and extracts fan/following counts via BilibiliClient.get_creator_all_fans, respecting the CREATOR_MODE configuration in config/bilibili_config.py.

Configuring Storage Backends

Select your persistence layer using the --save_data_option flag with one of the following values: csv, db, json, jsonl, sqlite, mongodb, or excel.

Each option maps to a specific implementation in store/bilibili/_store_impl.py:

  • CSV: BiliCsvStoreImplement writes to flat files in data/bili/.
  • MongoDB: BiliMongoStoreImplement stores documents in collections prefixed with bilibili_.
  • SQLite: BiliDbStoreImplement uses SQLAlchemy for relational storage.
  • Excel: Generates .xlsx files for business analysis workflows.

All storage classes inherit from AbstractStore defined in base/base_crawler.py and implement async methods: store_content(), store_comment(), and store_creator().

Programmatic API Integration

Beyond CLI usage, you can embed MediaCrawler directly in Python applications.

Running the Crawler from Python

Import the core components and simulate CLI arguments programmatically:

import asyncio
from cmd_arg.arg import parse_cmd
from media_platform.bilibili.core import BilibiliCrawler

async def scrape_bilibili():
    # Simulate command-line arguments

    args = await parse_cmd([
        "--platform", "bili",
        "--type", "search",
        "--keywords", "anime,music",
        "--save_data_option", "jsonl",
        "--headless", "true"
    ])
    
    crawler = BilibiliCrawler()
    await crawler.start()   # Handles browser init and login

    await crawler.close()

asyncio.run(scrape_bilibili())

The parse_cmd function returns a SimpleNamespace that updates global configuration values, allowing BilibiliCrawler.start() to execute with the specified parameters.

Implementing Custom Storage Backends

To store data in a custom system like PostgreSQL, implement the AbstractStore interface:

from base.base_crawler import AbstractStore

class PostgresStoreImplement(AbstractStore):
    async def store_content(self, content_item: dict):
        # Custom INSERT logic for video metadata

        pass
    
    async def store_comment(self, comment_item: dict):
        # Custom INSERT logic for comments

        pass
    
    async def store_creator(self, creator: dict):
        # Custom INSERT logic for creator profiles

        pass

Register your implementation by mapping the storage key in the platform initialization code, then invoke with --save_data_option postgres.

Key Configuration Files

File Purpose Key Elements
cmd_arg/arg.py CLI parsing and config injection parse_cmd function
config/bilibili_config.py Platform-specific constants BILI_SPECIFIED_ID_LIST, BILI_SEARCH_MODE
media_platform/bilibili/core.py Main orchestration logic BilibiliCrawler class, start() method
media_platform/bilibili/client.py API wrapper BilibiliClient with search_video_by_keyword, get_video_info
media_platform/bilibili/login.py Authentication flows QR-code and cookie login handlers
store/bilibili/_store_impl.py Storage backends BiliCsvStoreImplement, BiliMongoStoreImplement

Summary

  • MediaCrawler uses Playwright to automate Chromium for Bilibili scraping, handling JavaScript rendering and authentication automatically.
  • Three operational modes—search, detail, and creator—support keyword discovery, specific video extraction, and channel statistics harvesting respectively.
  • The architecture separates concerns between cmd_arg/arg.py for CLI parsing, media_platform/bilibili/core.py for orchestration, and store/bilibili/_store_impl.py for persistence.
  • Storage is pluggable via --save_data_option, supporting everything from CSV files to MongoDB with async batch writes.
  • Concurrency is controlled by config.MAX_CONCURRENCY_NUM (default 10) through asyncio semaphores to prevent rate limiting.

Frequently Asked Questions

How does MediaCrawler handle Bilibili authentication?

MediaCrawler supports three authentication methods implemented in media_platform/bilibili/login.py: QR-code scanning, phone number login, and cookie-based sessions. When running BilibiliCrawler.start(), the system checks for existing cookies in config.COOKIES and falls back to interactive QR-code login if no valid session exists.

Can I scrape specific Bilibili videos by URL instead of BV ID?

Yes. The --specified_id parameter accepts both raw BV IDs (e.g., BV1dwuKzmE26) and full URLs (e.g., https://www.bilibili.com/video/BV1dwuKzmE26/). The parse_video_info_from_url function in the crawler core extracts the numeric identifier automatically before calling BilibiliClient.get_video_info.

What is the difference between JSON and JSONL storage options?

The JSON option writes all records to a single JSON array file, while JSONL (JSON Lines) writes one JSON object per line. JSONL is preferred for large scraping jobs because it allows streaming writes and prevents memory issues when processing millions of comments or video records.

How do I limit concurrent requests to avoid IP bans?

Control concurrency using the --max_concurrency_num flag or by setting config.MAX_CONCURRENCY_NUM in config/bilibili_config.py. The default value is 10 simultaneous requests, managed by asyncio semaphores in BilibiliCrawler. Reducing this to 3-5 for creator scraping mode is recommended when operating without proxy rotation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →