Directory Structure for Platform-Specific Code in NanmiCoder/MediaCrawler

MediaCrawler isolates each social-media platform in its own sub-package under media_platform/, with consistent internal files (login.py, client.py, core.py, etc.) and mirrored storage logic in store/.

MediaCrawler organizes platform-specific crawling logic into modular sub-packages under the media_platform/ directory. This architecture keeps authentication flows, API clients, and data models cleanly separated while allowing shared utilities to handle caching and storage. According to the NanmiCoder/MediaCrawler source code, this pattern supports seven major Chinese social platforms.

The media_platform Directory Layout

The repository groups all platform-specific code under media_platform/ at the project root. Each platform occupies its own sub-package containing login workflows, HTTP clients, and field definitions.

Supported platforms include:

  • Zhihu (media_platform/zhihu/) - Handles login, client requests, and core crawling workflows
  • Xiaohongshu (XHS) (media_platform/xhs/) - Playwright-based authentication and API client
  • Weibo (media_platform/weibo/) - Login implementation and data extraction helpers
  • Baidu Tieba (media_platform/tieba/) - GraphQL request handling and helper functions
  • Kuaishou (media_platform/kuaishou/) - GraphQL query definitions and core logic
  • Douyin (media_platform/douyin/) - Authentication and crawling utilities
  • Bilibili (media_platform/bilibili/) - Login and data extraction helpers

Consistent Internal File Structure

Each platform package follows an identical internal layout in NanmiCoder/MediaCrawler. This standardization makes it easy to navigate between different platforms once you understand one package.

The standard files include:

  • __init__.py - Package entry point and public symbol exports
  • login.py - Platform-specific authentication (cookies, Playwright, etc.)
  • client.py - Low-level HTTP or GraphQL client
  • core.py - High-level crawler orchestration
  • help.py - Data extraction helper functions
  • field.py - Typed data structures for platform items
  • exception.py - Custom exceptions for platform-specific errors

Data Persistence Layer

The store/ directory mirrors the media_platform/ structure for data persistence. Each platform has its own implementation module (e.g., store/zhihu/_store_impl.py, store/xhs/_store_impl.py).

These storage modules utilize the generic AsyncFileWriter utility from tools/async_file_writer.py to write platform-segregated files under data/<platform>/.

Practical Usage Examples

Here are concrete examples of interacting with the platform-specific code structure.

Importing a Platform Client

from media_platform.zhihu.client import ZhihuClient

zhihu = ZhihuClient()
await zhihu.login()
profile = await zhihu.fetch_user_profile(user_id="123456")

Running via the CLI Runner


# Example: Starting a Tieba crawler

from media_platform.tieba.client import BaiduTieBaClient
from tools.app_runner import AppRunner

runner = AppRunner(platform="tieba", crawler_type="search")
await runner.run()

Storing Platform Data

from store.weibo._store_impl import WeiboStoreImpl

store = WeiboStoreImpl()
await store.save_media(media_item)

Key Source Files

Critical implementation files in the directory structure include:

Summary

  • MediaCrawler organizes platform-specific code under media_platform/ with sub-packages for Zhihu, XHS, Weibo, Tieba, Kuaishou, Douyin, and Bilibili
  • Each platform package contains standardized files: login.py, client.py, core.py, help.py, field.py, and exception.py
  • The store/ directory mirrors platform structure for data persistence using AsyncFileWriter
  • Platform data outputs to segregated data/<platform>/ directories
  • Entry point main.py routes to appropriate platform packages based on CLI arguments

Frequently Asked Questions

How do I add a new platform to MediaCrawler?

Create a new sub-package under media_platform/ following the existing template. Include login.py, client.py, core.py, help.py, field.py, and exception.py. Add corresponding storage implementation in store/<platform>/_store_impl.py and update main.py to recognize the new platform key.

Where does MediaCrawler store downloaded data?

Data persists to data/<platform>/ directories via the AsyncFileWriter utility in tools/async_file_writer.py. Each platform's storage implementation in store/<platform>/_store_impl.py manages platform-specific file formatting and organization. The writer creates segregated output directories automatically based on the platform parameter.

What is the difference between client.py and core.py in platform packages?

client.py handles low-level HTTP/GraphQL communication and authentication state, while core.py implements high-level crawling orchestration logic that coordinates the client, data extraction, and storage operations. The client manages session cookies and rate limiting, whereas core manages the crawling workflow and business logic.

How does the CLI runner select which platform to use?

The AppRunner class in tools/app_runner.py receives a platform parameter (e.g., "tieba", "zhihu") and dynamically imports the corresponding package from media_platform/ to execute the appropriate crawling workflow. This allows main.py to route commands to the correct platform implementation without hardcoding platform-specific logic in the entry point.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →