Directory Structure for Platform-Specific Code in NanmiCoder/MediaCrawler
MediaCrawler isolates each social-media platform in its own sub-package under media_platform/, with consistent internal files (login.py, client.py, core.py, etc.) and mirrored storage logic in store/.
MediaCrawler organizes platform-specific crawling logic into modular sub-packages under the media_platform/ directory. This architecture keeps authentication flows, API clients, and data models cleanly separated while allowing shared utilities to handle caching and storage. According to the NanmiCoder/MediaCrawler source code, this pattern supports seven major Chinese social platforms.
The media_platform Directory Layout
The repository groups all platform-specific code under media_platform/ at the project root. Each platform occupies its own sub-package containing login workflows, HTTP clients, and field definitions.
Supported platforms include:
- Zhihu (
media_platform/zhihu/) - Handles login, client requests, and core crawling workflows - Xiaohongshu (XHS) (
media_platform/xhs/) - Playwright-based authentication and API client - Weibo (
media_platform/weibo/) - Login implementation and data extraction helpers - Baidu Tieba (
media_platform/tieba/) - GraphQL request handling and helper functions - Kuaishou (
media_platform/kuaishou/) - GraphQL query definitions and core logic - Douyin (
media_platform/douyin/) - Authentication and crawling utilities - Bilibili (
media_platform/bilibili/) - Login and data extraction helpers
Consistent Internal File Structure
Each platform package follows an identical internal layout in NanmiCoder/MediaCrawler. This standardization makes it easy to navigate between different platforms once you understand one package.
The standard files include:
__init__.py- Package entry point and public symbol exportslogin.py- Platform-specific authentication (cookies, Playwright, etc.)client.py- Low-level HTTP or GraphQL clientcore.py- High-level crawler orchestrationhelp.py- Data extraction helper functionsfield.py- Typed data structures for platform itemsexception.py- Custom exceptions for platform-specific errors
Data Persistence Layer
The store/ directory mirrors the media_platform/ structure for data persistence. Each platform has its own implementation module (e.g., store/zhihu/_store_impl.py, store/xhs/_store_impl.py).
These storage modules utilize the generic AsyncFileWriter utility from tools/async_file_writer.py to write platform-segregated files under data/<platform>/.
Practical Usage Examples
Here are concrete examples of interacting with the platform-specific code structure.
Importing a Platform Client
from media_platform.zhihu.client import ZhihuClient
zhihu = ZhihuClient()
await zhihu.login()
profile = await zhihu.fetch_user_profile(user_id="123456")
Running via the CLI Runner
# Example: Starting a Tieba crawler
from media_platform.tieba.client import BaiduTieBaClient
from tools.app_runner import AppRunner
runner = AppRunner(platform="tieba", crawler_type="search")
await runner.run()
Storing Platform Data
from store.weibo._store_impl import WeiboStoreImpl
store = WeiboStoreImpl()
await store.save_media(media_item)
Key Source Files
Critical implementation files in the directory structure include:
media_platform/zhihu/client.py- Zhihu HTTP client and login flowmedia_platform/xhs/login.py- XHS Playwright-based authenticationmedia_platform/weibo/client.py- Weibo request handlingmedia_platform/tieba/core.py- High-level Tieba crawling logicmedia_platform/kuaishou/graphql.py- Kuaishou GraphQL query definitionsstore/zhihu/_store_impl.py- Zhihu-specific persistence implementationtools/async_file_writer.py- Async writer fordata/<platform>/filesmain.py- CLI entry point that selects platforms based on arguments
Summary
- MediaCrawler organizes platform-specific code under
media_platform/with sub-packages for Zhihu, XHS, Weibo, Tieba, Kuaishou, Douyin, and Bilibili - Each platform package contains standardized files:
login.py,client.py,core.py,help.py,field.py, andexception.py - The
store/directory mirrors platform structure for data persistence usingAsyncFileWriter - Platform data outputs to segregated
data/<platform>/directories - Entry point
main.pyroutes to appropriate platform packages based on CLI arguments
Frequently Asked Questions
How do I add a new platform to MediaCrawler?
Create a new sub-package under media_platform/ following the existing template. Include login.py, client.py, core.py, help.py, field.py, and exception.py. Add corresponding storage implementation in store/<platform>/_store_impl.py and update main.py to recognize the new platform key.
Where does MediaCrawler store downloaded data?
Data persists to data/<platform>/ directories via the AsyncFileWriter utility in tools/async_file_writer.py. Each platform's storage implementation in store/<platform>/_store_impl.py manages platform-specific file formatting and organization. The writer creates segregated output directories automatically based on the platform parameter.
What is the difference between client.py and core.py in platform packages?
client.py handles low-level HTTP/GraphQL communication and authentication state, while core.py implements high-level crawling orchestration logic that coordinates the client, data extraction, and storage operations. The client manages session cookies and rate limiting, whereas core manages the crawling workflow and business logic.
How does the CLI runner select which platform to use?
The AppRunner class in tools/app_runner.py receives a platform parameter (e.g., "tieba", "zhihu") and dynamically imports the corresponding package from media_platform/ to execute the appropriate crawling workflow. This allows main.py to route commands to the correct platform implementation without hardcoding platform-specific logic in the entry point.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →