# How to Use MediaCrawler for Scraping Data from Bilibili: A Complete Technical Guide

> Learn how to scrape Bilibili data with NanmiCoder MediaCrawler. Extract videos, comments, and profiles using this async, Playwright-based framework. Supports CSV, MongoDB, SQLite, and Excel.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-01

---

**MediaCrawler provides an asynchronous, Playwright-based framework for extracting Bilibili videos, comments, and creator profiles with support for multiple storage backends including CSV, MongoDB, SQLite, and Excel.**

MediaCrawler is an open-source Python scraping framework designed for Chinese social media platforms. When configured for Bilibili, it orchestrates browser automation to handle authentication, rate limiting, and data extraction while offering both command-line and programmatic APIs for flexible deployment.

## Understanding the Bilibili Crawling Architecture

The repository follows a modular pipeline architecture that separates concerns between configuration, HTTP client operations, and data persistence.

### Core Components

- **Command-line parser** ([`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py)): The `parse_cmd` function converts CLI flags like `--platform bili` and `--type search` into a configuration object that updates global settings.
- **Configuration module** ([`config/bilibili_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/bilibili_config.py)): Contains Bilibili-specific defaults including `BILI_SPECIFIED_ID_LIST` for target videos, search mode parameters, and video quality settings.
- **Crawler core** ([`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py)): The `BilibiliCrawler` class orchestrates the browser lifecycle through `BilibiliCrawler.start()`, manages login state, and dispatches to specific crawling modes.
- **HTTP client** ([`media_platform/bilibili/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/client.py)): The `BilibiliClient` class provides low-level API methods including `search_video_by_keyword`, `get_video_info`, and `get_video_all_comments`.
- **Storage implementations** ([`store/bilibili/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/bilibili/_store_impl.py)): Concrete classes like `BiliCsvStoreImplement`, `BiliDbStoreImplement`, and `BiliMongoStoreImplement` handle persistence logic for different backends.

### Execution Flow

1. **Initialization**: CLI arguments are parsed by `parse_cmd` in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) to populate the global `config` object.
2. **Browser Launch**: `BilibiliCrawler.start()` initializes a Playwright browser instance (Chromium by default) and handles authentication via QR-code, phone, or cookie login.
3. **Mode Dispatch**: Based on `config.CRAWLER_TYPE`, the crawler executes one of three paths:
   - **Search mode**: Calls `search()` → `search_by_keywords()` to crawl search result pages.
   - **Detail mode**: Calls `get_specified_videos()` to fetch metadata for specific BV IDs or URLs.
   - **Creator mode**: Calls `get_all_creator_details()` to extract fan counts, followings, and dynamic posts.
4. **Data Extraction**: For each video, the crawler optionally downloads media files via `BilibiliVideo.store_video` in [`store/bilibili/bilibilli_store_media.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/bilibili/bilibilli_store_media.py) and extracts comments through `BilibiliClient.get_video_all_comments`.
5. **Persistence**: Results are routed to the storage backend selected via `--save_data_option`, with implementations in [`store/bilibili/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/bilibili/_store_impl.py) handling async writes.

## Installation and Prerequisites

Install the required dependencies and browser binaries before running Bilibili scraping tasks:

```bash
pip install -r requirements.txt
playwright install chromium

```

The [`requirements.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/requirements.txt) includes essential packages such as `playwright`, `aiofiles`, `sqlalchemy`, and `pymongo` that support the async architecture and multiple storage backends.

## Command-Line Usage Patterns

MediaCrawler supports three primary crawling modes for Bilibili, controlled by the `--type` parameter.

### Search Videos by Keyword

Use **search mode** to discover videos based on keywords and time ranges:

```bash
python -m MediaCrawler \
  --platform bili \
  --type search \
  --keywords "gaming,anime" \
  --save_data_option csv \
  --headless true

```

This command respects the `BILI_SEARCH_MODE` setting in [`config/bilibili_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/bilibili_config.py) (defaulting to *normal* search) and saves video metadata, comments, and creator information to CSV files under `data/bili/`.

### Crawl Specific Videos by ID or URL

Use **detail mode** when you have specific video identifiers:

```bash
python -m MediaCrawler \
  --platform bili \
  --type detail \
  --specified_id "BV1dwuKzmE26,BV14Q4y1n7jz" \
  --save_data_option mongodb

```

The crawler parses each entry through `parse_video_info_from_url` to extract the numeric `bvid`, then calls `BilibiliClient.get_video_info` to retrieve full metadata and comments.

### Harvest Creator Profiles and Fan Data

Use **creator mode** to extract comprehensive channel statistics:

```bash
python -m MediaCrawler \
  --platform bili \
  --type creator \
  --creator_id "20813884,https://space.bilibili.com/434377496" \
  --save_data_option jsonl

```

This invokes `get_all_creator_details()` which iterates through creator pages and extracts fan/following counts via `BilibiliClient.get_creator_all_fans`, respecting the `CREATOR_MODE` configuration in [`config/bilibili_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/bilibili_config.py).

## Configuring Storage Backends

Select your persistence layer using the `--save_data_option` flag with one of the following values: `csv`, `db`, `json`, `jsonl`, `sqlite`, `mongodb`, or `excel`.

Each option maps to a specific implementation in [`store/bilibili/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/bilibili/_store_impl.py):

- **CSV**: `BiliCsvStoreImplement` writes to flat files in `data/bili/`.
- **MongoDB**: `BiliMongoStoreImplement` stores documents in collections prefixed with `bilibili_`.
- **SQLite**: `BiliDbStoreImplement` uses SQLAlchemy for relational storage.
- **Excel**: Generates `.xlsx` files for business analysis workflows.

All storage classes inherit from `AbstractStore` defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) and implement async methods: `store_content()`, `store_comment()`, and `store_creator()`.

## Programmatic API Integration

Beyond CLI usage, you can embed MediaCrawler directly in Python applications.

### Running the Crawler from Python

Import the core components and simulate CLI arguments programmatically:

```python
import asyncio
from cmd_arg.arg import parse_cmd
from media_platform.bilibili.core import BilibiliCrawler

async def scrape_bilibili():
    # Simulate command-line arguments

    args = await parse_cmd([
        "--platform", "bili",
        "--type", "search",
        "--keywords", "anime,music",
        "--save_data_option", "jsonl",
        "--headless", "true"
    ])
    
    crawler = BilibiliCrawler()
    await crawler.start()   # Handles browser init and login

    await crawler.close()

asyncio.run(scrape_bilibili())

```

The `parse_cmd` function returns a `SimpleNamespace` that updates global configuration values, allowing `BilibiliCrawler.start()` to execute with the specified parameters.

### Implementing Custom Storage Backends

To store data in a custom system like PostgreSQL, implement the `AbstractStore` interface:

```python
from base.base_crawler import AbstractStore

class PostgresStoreImplement(AbstractStore):
    async def store_content(self, content_item: dict):
        # Custom INSERT logic for video metadata

        pass
    
    async def store_comment(self, comment_item: dict):
        # Custom INSERT logic for comments

        pass
    
    async def store_creator(self, creator: dict):
        # Custom INSERT logic for creator profiles

        pass

```

Register your implementation by mapping the storage key in the platform initialization code, then invoke with `--save_data_option postgres`.

## Key Configuration Files

| File | Purpose | Key Elements |
|------|---------|--------------|
| [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) | CLI parsing and config injection | `parse_cmd` function |
| [`config/bilibili_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/bilibili_config.py) | Platform-specific constants | `BILI_SPECIFIED_ID_LIST`, `BILI_SEARCH_MODE` |
| [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) | Main orchestration logic | `BilibiliCrawler` class, `start()` method |
| [`media_platform/bilibili/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/client.py) | API wrapper | `BilibiliClient` with `search_video_by_keyword`, `get_video_info` |
| [`media_platform/bilibili/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/login.py) | Authentication flows | QR-code and cookie login handlers |
| [`store/bilibili/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/bilibili/_store_impl.py) | Storage backends | `BiliCsvStoreImplement`, `BiliMongoStoreImplement` |

## Summary

- **MediaCrawler** uses Playwright to automate Chromium for Bilibili scraping, handling JavaScript rendering and authentication automatically.
- Three operational modes—**search**, **detail**, and **creator**—support keyword discovery, specific video extraction, and channel statistics harvesting respectively.
- The architecture separates concerns between [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) for CLI parsing, [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) for orchestration, and [`store/bilibili/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/bilibili/_store_impl.py) for persistence.
- Storage is pluggable via `--save_data_option`, supporting everything from CSV files to MongoDB with async batch writes.
- Concurrency is controlled by `config.MAX_CONCURRENCY_NUM` (default 10) through asyncio semaphores to prevent rate limiting.

## Frequently Asked Questions

### How does MediaCrawler handle Bilibili authentication?

MediaCrawler supports three authentication methods implemented in [`media_platform/bilibili/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/login.py): QR-code scanning, phone number login, and cookie-based sessions. When running `BilibiliCrawler.start()`, the system checks for existing cookies in `config.COOKIES` and falls back to interactive QR-code login if no valid session exists.

### Can I scrape specific Bilibili videos by URL instead of BV ID?

Yes. The `--specified_id` parameter accepts both raw BV IDs (e.g., `BV1dwuKzmE26`) and full URLs (e.g., `https://www.bilibili.com/video/BV1dwuKzmE26/`). The `parse_video_info_from_url` function in the crawler core extracts the numeric identifier automatically before calling `BilibiliClient.get_video_info`.

### What is the difference between JSON and JSONL storage options?

The **JSON** option writes all records to a single JSON array file, while **JSONL** (JSON Lines) writes one JSON object per line. JSONL is preferred for large scraping jobs because it allows streaming writes and prevents memory issues when processing millions of comments or video records.

### How do I limit concurrent requests to avoid IP bans?

Control concurrency using the `--max_concurrency_num` flag or by setting `config.MAX_CONCURRENCY_NUM` in [`config/bilibili_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/bilibili_config.py). The default value is 10 simultaneous requests, managed by asyncio semaphores in `BilibiliCrawler`. Reducing this to 3-5 for creator scraping mode is recommended when operating without proxy rotation.