How Platform-Specific field.py Modules Define Data Schemas in MediaCrawler
Each platform-specific field.py module in MediaCrawler encapsulates typed constants and data structures that define valid request parameters and API response schemas for that specific media platform.
The NanmiCoder/MediaCrawler repository organizes data schemas through isolated field.py files located in each platform's subdirectory under media_platform/. These modules serve as the single source of truth for platform-specific values, ensuring that crawling logic remains clean and maintainable while preventing typos in API parameter strings.
Schema Design Pattern in MediaCrawler field.py Files
Every field.py module follows a consistent architectural pattern that separates platform knowledge from crawling logic. According to the source code, these files contain no network logic—they purely provide typed constants that higher-level code consumes when building query dictionaries or parsing JSON responses.
Enum Classes for Request Parameters
The primary building block is the Enum class, which represents a closed set of allowed values for API parameters. These enums map human-readable names to the exact integer or string values each platform expects.
Typical implementations include enums for search types, sort orders, time ranges, and channel filters.
NamedTuple for Response Objects
Some platforms, notably Xiaohongshu, extend the pattern to include NamedTuple classes. These define lightweight, immutable data structures that mirror the fields returned by the platform's API, providing dot-notation access to response attributes like note.title or note.video_url.
Constants Import Pattern
Certain modules import platform-specific constant definitions from a shared constant package. For example, media_platform/zhihu/field.py imports from constant import zhihu as zhihu_constant to map enum values to the exact string literals the Zhihu API requires.
Platform-Specific Schema Implementations
Each media platform tailors its field.py to expose the specific parameters and response structures relevant to that API.
Douyin (media_platform/douyin/field.py)
The Douyin implementation defines three core enums that control search behavior:
- SearchChannelType: Enumerates content channels (
aweme_general,aweme_video_web) - SearchSortType: Maps sorting preferences to integer codes (
GENERAL=0,MOST_LIKE=1,LATEST=2) - PublishTimeType: Defines time filter ranges (
UNLIMITED=0,ONE_DAY=1,ONE_WEEK=2)
from media_platform.douyin.field import SearchChannelType, SearchSortType, PublishTimeType
params = {
"channel": SearchChannelType.VIDEO.value,
"sort_type": SearchSortType.MOST_LIKE.value,
"publish_time": PublishTimeType.ONE_WEEK.value,
}
Zhihu (media_platform/zhihu/field.py)
Zhihu's schema focuses on search configuration through three enums that reference external constants:
- SearchTime: Time range filters (
ONE_DAY,ONE_WEEK,ONE_MONTH) - SearchType: Content type selectors (
ANSWER,VIDEO,ARTICLE) mapped viazhihu_constant - SearchSort: Ordering methods (
UPVOTED_COUNT,NEWEST)
from media_platform.zhihu.field import SearchTime, SearchType, SearchSort
query = {
"time": SearchTime.ONE_WEEK.value,
"type": SearchType.VIDEO.value,
"sort": SearchSort.UPVOTED_COUNT.value,
}
Xiaohongshu (media_platform/xhs/field.py)
Xiaohongshu's field.py provides the most comprehensive schema definition, combining request enums with a response data structure:
Enums:
- FeedType: Content category feeds (
FOOD,TRAVEL,FASHION) - NoteType: Content format classifications
- SearchSortType and SearchNoteType: Query modifiers
NamedTuple:
- Note: A complete data structure covering IDs, title, media URLs, statistics, and timestamps
from media_platform.xhs.field import FeedType, Note
# Build a feed request
feed_params = {"type": FeedType.FOOD.value}
# Parse a note response
note = Note(*json_item) # json_item is a list/tuple matching the fields
print(note.title, note.video_url)
Bilibili (media_platform/bilibili/field.py)
Bilibili organizes schemas around content ordering:
- SearchOrderType: Result sorting (
MOST_CLICK,LAST_PUBLISH,MOST_STAR) - CommentOrderType: Comment thread ordering (
MIXED,TIME,LIKE_COUNT)
from media_platform.bilibili.field import SearchOrderType
params = {"order": SearchOrderType.MOST_CLICK.value}
Tieba (media_platform/tieba/field.py)
The Tieba implementation handles forum-specific search parameters:
- SearchSortType: Time and relevance ordering (
TIME_DESC,TIME_ASC,RELEVANCE) - SearchNoteType: Thread type filtering (
MAIN_THREAD,FIXED_THREAD)
from media_platform.tieba.field import SearchSortType, SearchNoteType
search_opts = {
"sort": SearchSortType.TIME_DESC.value,
"type": SearchNoteType.MAIN_THREAD.value,
}
Kuaishou (media_platform/kuaishou/field.py)
Currently, the Kuaishou module contains only a file header and serves as a placeholder. It follows the same extension pattern as other platforms, ready to accept Enum and NamedTuple definitions as the implementation matures.
Why MediaCrawler Uses field.py for Schema Definition
This architecture provides three critical advantages for the crawling framework:
- Isolation of Platform Knowledge: All platform-specific constants reside in one location, making updates trivial when APIs change their parameter values or add new options.
- Type Safety: Using
EnumandNamedTupleprovides IDE autocomplete hints and prevents runtime errors from string typos in parameter dictionaries. - Reusability: Higher-level crawlers import these enums and data classes without embedding literal string values throughout the codebase.
Unified Usage Example Across Platforms
The consistent interface allows developers to construct multi-platform requests using the same import pattern:
from media_platform.douyin.field import SearchChannelType, SearchSortType
from media_platform.zhihu.field import SearchTime, SearchSort
from media_platform.xhs.field import FeedType, Note
# Build a mixed-platform request map
platform_requests = {
"douyin": {
"channel": SearchChannelType.VIDEO.value,
"sort": SearchSortType.LATEST.value,
},
"zhihu": {
"time": SearchTime.ONE_DAY.value,
"sort": SearchSort.UPVOTED_COUNT.value,
},
"xhs": {
"feed": FeedType.TRAVEL.value,
},
}
Summary
- Platform-specific
field.pymodules in MediaCrawler define data schemas usingEnumclasses for request parameters and optionalNamedTupleclasses for response structures. - No network logic exists in these files; they serve purely as typed constant repositories consumed by higher-level crawling code.
- Key files include
media_platform/douyin/field.py,media_platform/zhihu/field.py,media_platform/xhs/field.py,media_platform/bilibili/field.py, andmedia_platform/tieba/field.py. - Xiaohongshu uniquely implements a
NoteNamedTuple for structured response parsing, while other platforms focus primarily on request parameter enums. - Isolation and type safety are the primary architectural benefits, preventing API string typos and centralizing platform-specific knowledge.
Frequently Asked Questions
What is the purpose of field.py in MediaCrawler?
The field.py modules serve as schema definition files that encapsulate platform-specific constants, enums, and data structures. They provide the exact parameter values and response field mappings required to communicate with each media platform's API, ensuring type safety and preventing hardcoded string literals throughout the crawling logic.
How does Xiaohongshu's field.py differ from other platforms?
While most platforms define only Enum classes for request parameters, media_platform/xhs/field.py additionally includes a Note NamedTuple that maps the complete structure of API responses. This allows developers to parse JSON items into immutable objects with dot-notation access to fields like note.title, note.video_url, and note.like_count.
Does field.py contain network logic?
No. According to the source code analysis, these modules strictly contain typed constants and data structure definitions. They do not implement HTTP requests, response handling, or any network operations. The crawling logic in other modules imports these constants to build query dictionaries and parse responses.
How do I add a new platform to MediaCrawler?
Create a new directory under media_platform/ for your target platform, then implement a field.py file following the established pattern: define Enum classes for all API parameters (search types, sort orders, time filters), import platform-specific constants if needed, and optionally define NamedTuple classes for complex response objects. The Kuaishou implementation demonstrates the minimal placeholder structure acceptable for new platforms.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →