# Enabling Second-Level Comment Crawling for All Platforms in MediaCrawler

> Learn how to enable second-level comment crawling in MediaCrawler for all platforms. Use the ENABLE_GET_SUB_COMMENTS flag for hierarchical comment extraction via CLI or config.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: new-feature-announcement
- Published: 2026-07-31

---

**MediaCrawler supports hierarchical comment extraction through a single global flag `ENABLE_GET_SUB_COMMENTS` that activates sub-comment retrieval logic across all platform-specific clients when enabled via CLI or configuration.**

The MediaCrawler repository provides a unified framework for scraping content from Chinese social media platforms. While first-level (root) comments are collected by default, second-level (reply) comments require explicit activation through a centralized configuration mechanism that propagates to every supported platform client.

## How the Configuration Flag Controls Sub-Comment Crawling

The feature centers on the **`ENABLE_GET_SUB_COMMENTS`** boolean defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) at lines 16-18. By default, this value is set to `False`, ensuring that sub-comment crawling remains opt-in to conserve bandwidth and storage resources.

When you launch the crawler via command line, the argument parser in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) (lines 13-18) exposes the `--get_sub_comment` option. This CLI flag accepts truthy values (`yes`, `true`, `1`, `y`, `t`) and writes the parsed boolean directly to `config.ENABLE_GET_SUB_COMMENTS`.

Each platform client implements a guard clause that checks this flag before executing sub-comment retrieval logic. For example, in [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) (lines 51-53), the method returns an empty list immediately if the flag is disabled:

```python
if not config.ENABLE_GET_SUB_COMMENTS:
    return []

```

When the flag is `True`, clients enter a pagination loop that checks the root comment's `sub_comment_count`, repeatedly calls platform-specific APIs (such as `get_child_comments` or `get_note_sub_comments`), extracts comment objects via `_extractor.extract_comments`, and accumulates results into `all_sub_comments` for merging with first-level data.

## Enabling Second-Level Comments via CLI

The simplest activation method is passing the `--get_sub_comment` flag when launching the crawler from your terminal:

```bash
python media_crawler.py --platform zhihu --get_sub_comment true

```

The argument parser accepts multiple truthy representations, allowing flexibility in shell scripts and automation pipelines. Valid positive values include `yes`, `true`, `1`, `y`, and `t`.

## Enabling Second-Level Comments Programmatically

For custom scripts or Jupyter notebooks that instantiate clients directly, import the configuration object and set the attribute before creating any client instances:

```python
from config.base_config import config

# Enable sub-comment crawling globally

config.ENABLE_GET_SUB_COMMENTS = True

# Subsequent client instances will automatically fetch replies

from media_platform.zhihu.client import ZhihuClient
client = ZhihuClient()

```

This approach ensures that all subsequent platform client instantiations respect the updated configuration without requiring command-line arguments.

## Platform-Specific Implementation Details

Each platform client implements the `get_comments_all_sub_comments` method with logic tailored to that platform's API structure and pagination mechanisms.

### Zhihu

In [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) (lines 33-84), the client loops while `parment_comment.sub_comment_count > 0`, invoking `get_child_comments` for each pagination batch and extracting structured data via `_extractor.extract_comments`.

### XiaoHongShu (XHS)

The XHS client in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) (lines 474-535) monitors `sub_comment_has_more` and `sub_comment_cursor` fields to manage pagination through the `get_note_sub_comments` endpoint.

### Weibo

Located in [`media_platform/weibo/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/client.py) (lines 24-27 and 246), the implementation reads the nested `comments` field within comment items and forwards these replies to the callback function if present.

### Tieba

The Tieba client at [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) (lines 502-595 and 532) calculates the total number of sub-comment pages using integer division (`sub_comment_count // 10 + 1`) and generates sequential URLs to fetch each page.

### Kuaishou

In [`media_platform/kuaishou/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/client.py) (lines 277-305), the logic checks the `hasSubComments` boolean flag and utilizes cursor-based pagination through the `pcursorV2` parameter to retrieve nested replies.

### Douyin and Bilibili

These platforms handle sub-comments through the `is_fetch_sub_comments` parameter within their core classes rather than the global flag alone. See [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) (line 259) and [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) (line 358) for the specific implementation details that integrate with the broader configuration system.

## Data Storage and Schema Compatibility

When sub-comments are collected, they are treated as standard comment objects and persisted according to your selected storage backend—whether MongoDB, SQLite, JSONL, or others. The data models already include a **`sub_comment_count`** field (visible in [`model/m_zhihu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_zhihu.py) and analogous platform models) that records the number of replies, ensuring backward compatibility with existing database schemas and downstream analytics pipelines.

## Summary

- **Global Flag**: Set `ENABLE_GET_SUB_COMMENTS = True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) or use the `--get_sub_comment` CLI argument to activate hierarchical comment crawling.
- **Platform Coverage**: The flag controls sub-comment retrieval in Zhihu, XHS, Weibo, Tieba, and Kuaishou clients through guard clauses in their respective `get_comments_all_sub_comments` methods.
- **CLI Integration**: The [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) module parses `--get_sub_comment` and accepts multiple truthy value formats for shell scripting convenience.
- **Pagination Logic**: Each platform implements custom pagination handling (cursor-based, page-based, or count-based) to traverse the full reply hierarchy.
- **Storage Ready**: Sub-comments integrate seamlessly with existing storage backends and preserve the `sub_comment_count` metadata field.

## Frequently Asked Questions

### Does enabling sub-comment crawling significantly increase API rate limit usage?

Yes. Each second-level comment requires additional API calls proportional to the reply count. Platforms like Tieba calculate pages as `sub_comment_count // 10 + 1`, while cursor-based platforms like Kuaishou iterate until `hasSubComments` returns false. Monitor your request quotas when enabling this feature on high-engagement posts.

### Which platforms support second-level comment extraction in MediaCrawler?

All major platforms in the repository support the feature: **Zhihu**, **XiaoHongShu (XHS)**, **Weibo**, **Tieba**, **Kuaishou**, **Douyin**, and **Bilibili**. Each implements the logic in their respective client files (e.g., [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py)) or core classes ([`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py)).

### How are sub-comments structured in the output data?

Sub-comments are extracted using the same `_extractor.extract_comments` method as first-level comments and merged into the final result set. They maintain identical schema fields including `sub_comment_count`, allowing recursive analysis of conversation depth without separate processing pipelines.

### Can I disable sub-comment crawling mid-execution?

While the `ENABLE_GET_SUB_COMMENTS` flag is checked at the start of each `get_comments_all_sub_comments` invocation, changing the configuration value during runtime only affects subsequent comment retrieval calls. Already fetched sub-comments remain in memory, but new parent comments will skip sub-comment fetching if the flag is set to `False` before their processing begins.