# Directory Structure for Platform-Specific Code in NanmiCoder/MediaCrawler

> Explore the NanmiCoder/MediaCrawler directory structure for platform-specific code. Learn how to isolate each social media platform with consistent internal files and mirrored storage logic.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: architecture
- Published: 2026-07-03

---

**MediaCrawler isolates each social-media platform in its own sub-package under `media_platform/`, with consistent internal files ([`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py), [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py), etc.) and mirrored storage logic in `store/`.**

MediaCrawler organizes platform-specific crawling logic into modular sub-packages under the `media_platform/` directory. This architecture keeps authentication flows, API clients, and data models cleanly separated while allowing shared utilities to handle caching and storage. According to the NanmiCoder/MediaCrawler source code, this pattern supports seven major Chinese social platforms.

## The media_platform Directory Layout

The repository groups all platform-specific code under `media_platform/` at the project root. Each platform occupies its own sub-package containing login workflows, HTTP clients, and field definitions.

Supported platforms include:

- **Zhihu** (`media_platform/zhihu/`) - Handles login, client requests, and core crawling workflows
- **Xiaohongshu (XHS)** (`media_platform/xhs/`) - Playwright-based authentication and API client
- **Weibo** (`media_platform/weibo/`) - Login implementation and data extraction helpers
- **Baidu Tieba** (`media_platform/tieba/`) - GraphQL request handling and helper functions
- **Kuaishou** (`media_platform/kuaishou/`) - GraphQL query definitions and core logic
- **Douyin** (`media_platform/douyin/`) - Authentication and crawling utilities
- **Bilibili** (`media_platform/bilibili/`) - Login and data extraction helpers

## Consistent Internal File Structure

Each platform package follows an identical internal layout in NanmiCoder/MediaCrawler. This standardization makes it easy to navigate between different platforms once you understand one package.

The standard files include:

- [`__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/__init__.py) - Package entry point and public symbol exports
- [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py) - Platform-specific authentication (cookies, Playwright, etc.)
- [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py) - Low-level HTTP or GraphQL client
- [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py) - High-level crawler orchestration
- [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py) - Data extraction helper functions
- [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py) - Typed data structures for platform items
- [`exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/exception.py) - Custom exceptions for platform-specific errors

## Data Persistence Layer

The `store/` directory mirrors the `media_platform/` structure for data persistence. Each platform has its own implementation module (e.g., [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py), [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py)).

These storage modules utilize the generic `AsyncFileWriter` utility from [`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) to write platform-segregated files under `data/<platform>/`.

## Practical Usage Examples

Here are concrete examples of interacting with the platform-specific code structure.

### Importing a Platform Client

```python
from media_platform.zhihu.client import ZhihuClient

zhihu = ZhihuClient()
await zhihu.login()
profile = await zhihu.fetch_user_profile(user_id="123456")

```

### Running via the CLI Runner

```python

# Example: Starting a Tieba crawler

from media_platform.tieba.client import BaiduTieBaClient
from tools.app_runner import AppRunner

runner = AppRunner(platform="tieba", crawler_type="search")
await runner.run()

```

### Storing Platform Data

```python
from store.weibo._store_impl import WeiboStoreImpl

store = WeiboStoreImpl()
await store.save_media(media_item)

```

## Key Source Files

Critical implementation files in the directory structure include:

- [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) - Zhihu HTTP client and login flow
- [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) - XHS Playwright-based authentication
- [`media_platform/weibo/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/client.py) - Weibo request handling
- [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py) - High-level Tieba crawling logic
- [`media_platform/kuaishou/graphql.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/graphql.py) - Kuaishou GraphQL query definitions
- [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py) - Zhihu-specific persistence implementation
- [`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) - Async writer for `data/<platform>/` files
- [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) - CLI entry point that selects platforms based on arguments

## Summary

- MediaCrawler organizes platform-specific code under `media_platform/` with sub-packages for Zhihu, XHS, Weibo, Tieba, Kuaishou, Douyin, and Bilibili
- Each platform package contains standardized files: [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py), [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py), [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py), [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py), and [`exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/exception.py)
- The `store/` directory mirrors platform structure for data persistence using `AsyncFileWriter`
- Platform data outputs to segregated `data/<platform>/` directories
- Entry point [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) routes to appropriate platform packages based on CLI arguments

## Frequently Asked Questions

### How do I add a new platform to MediaCrawler?

Create a new sub-package under `media_platform/` following the existing template. Include [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py), [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py), [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py), [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py), and [`exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/exception.py). Add corresponding storage implementation in `store/<platform>/_store_impl.py` and update [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to recognize the new platform key.

### Where does MediaCrawler store downloaded data?

Data persists to `data/<platform>/` directories via the `AsyncFileWriter` utility in [`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py). Each platform's storage implementation in `store/<platform>/_store_impl.py` manages platform-specific file formatting and organization. The writer creates segregated output directories automatically based on the platform parameter.

### What is the difference between client.py and core.py in platform packages?

[`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py) handles low-level HTTP/GraphQL communication and authentication state, while [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py) implements high-level crawling orchestration logic that coordinates the client, data extraction, and storage operations. The client manages session cookies and rate limiting, whereas core manages the crawling workflow and business logic.

### How does the CLI runner select which platform to use?

The `AppRunner` class in [`tools/app_runner.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/app_runner.py) receives a `platform` parameter (e.g., `"tieba"`, `"zhihu"`) and dynamically imports the corresponding package from `media_platform/` to execute the appropriate crawling workflow. This allows [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to route commands to the correct platform implementation without hardcoding platform-specific logic in the entry point.