# How MediaCrawler Implements Second-Level Comment Crawling with ENABLE_GET_SUB_COMMENTS

> Learn how MediaCrawler enables second-level comment crawling by setting ENABLE_GET_SUB_COMMENTS to True. This activates the get_comments_all_sub_comments method for nested replies on Zhihu, Douyin, and Bilibili.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: deep-dive
- Published: 2026-06-29

---

**Set `ENABLE_GET_SUB_COMMENTS = True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to enable automatic fetching of nested replies, which triggers the `get_comments_all_sub_comments` method to paginate through child comments across supported platforms like Zhihu, Douyin, and Bilibili.**

MediaCrawler is an open-source multi-platform content crawling framework that supports optional second-level comment extraction via a centralized configuration flag. When enabled, the framework automatically detects parent comments containing nested replies and initiates a dedicated pagination workflow to retrieve all sub-comments while respecting platform rate limits. This implementation maintains API consistency across platforms including Zhihu, Douyin, Bilibili, Weibo, Tieba, and Kuaishou.

## Configuration Flag in base_config.py

The feature is controlled by the global boolean `ENABLE_GET_SUB_COMMENTS` defined in the central configuration file. By default, sub-comment crawling is disabled to minimize API calls and processing time.

```python

# config/base_config.py (lines 118-119)

ENABLE_GET_SUB_COMMENTS = False      # ← toggle to enable sub‑comment crawling

```

You can activate this feature programmatically by setting `config.ENABLE_GET_SUB_COMMENTS = True` before initializing the client, or via the command line using `--enable-sub-comments true`, which is parsed in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) and mapped to the same configuration variable.

## Entry Point for Root and Sub-Comment Retrieval

For platforms like Zhihu, the primary entry point is `get_note_all_comments` in the platform-specific client. After retrieving batches of root comments, this method conditionally invokes the sub-comment crawler:

```python

# media_platform/zhihu/client.py (lines 28-30)

await self.get_comments_all_sub_comments(content, comments,
                                         crawl_interval=crawl_interval,
                                         callback=callback)

```

This pattern is replicated across all supported platforms in their respective `media_platform/{platform}/client.py` files, ensuring uniform behavior whether you are crawling Zhihu, Douyin, or Bilibili.

## The Sub-Comment Crawling Workflow

The method `get_comments_all_sub_comments` in [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) implements a seven-step workflow to retrieve nested replies. Each step includes specific guard clauses and pagination logic to handle platform API limitations.

### Step 1: Configuration Guard Check

The method immediately returns an empty list if the feature is disabled, preventing unnecessary API calls:

```python

# media_platform/zhihu/client.py (lines 51-53)

if not config.ENABLE_GET_SUB_COMMENTS:
    return []

```

### Step 2: Iterate Over Parent Comments

The crawler loops through each root comment, filtering for those with a non-zero `sub_comment_count` field:

```python

# media_platform/zhihu/client.py (lines 55-57)

for parment_comment in comments:
    if parment_comment.sub_comment_count == 0:
        continue

```

### Step 3: Pagination Loop

For each valid parent comment, the method enters a `while` loop that continues until the API reports `is_end` or the pagination offset stops changing:

```python

# media_platform/zhihu/client.py (lines 63-70)

while not is_end:
    # ... API call logic ...

    if offset == child_comment_res.get("offset", offset):
        break
    is_end = child_comment_res.get("is_end", True)

```

### Step 4: Child Comment API Request

Inside the pagination loop, the method calls `get_child_comments`, which maps to the platform's child comment endpoint (e.g., Zhihu's `/child_comment`):

```python

# media_platform/zhihu/client.py (lines 65-66)

child_comment_res = await self.get_child_comments(parment_comment.comment_id, offset, limit)

```

### Step 5: Data Extraction and Aggregation

Raw JSON responses are transformed into structured `ZhihuComment` objects using the platform's extractor class:

```python

# media_platform/zhihu/client.py (lines 71-72)

sub_comments = self._extractor.extract_comments(content, child_comment_res.get("data"))
all_sub_comments.extend(sub_comments)

```

### Step 6: Optional Callback Execution

If a callback function was provided (useful for streaming results to databases or real-time processing), it is invoked after each batch:

```python

# media_platform/zhihu/client.py (lines 78-79)

if callback:
    await callback(sub_comments)

```

### Step 7: Rate Limiting

The method respects the `crawl_interval` parameter to avoid triggering platform anti-bot measures:

```python

# media_platform/zhihu/client.py (line 82)

await asyncio.sleep(crawl_interval)

```

## Practical Implementation Examples

### Enable via Command Line

When running MediaCrawler from the terminal, pass the flag to activate sub-comment crawling:

```bash
python main.py --enable-sub-comments true --platform zhihu

```

The argument parser in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) automatically maps this to `config.ENABLE_GET_SUB_COMMENTS`.

### Programmatic Usage with Callback

For custom pipelines, enable the flag in Python and provide a callback to process sub-comments incrementally:

```python
import asyncio
import config
from media_platform.zhihu.client import ZhiHuClient

async def sub_comment_handler(sub_comments):
    """Process each batch of sub-comments as they arrive."""
    print(f"Received {len(sub_comments)} nested replies")
    # Insert into database or stream to analytics here

async def main():
    # Enable second-level comment crawling

    config.ENABLE_GET_SUB_COMMENTS = True
    
    client = ZhiHuClient(
        timeout=15,
        proxy=None,
        headers={"User-Agent": "MediaCrawler"},
        playwright_page=None,
        cookie_dict={"d_c0": "..."}
    )
    
    # Fetch content and all comments (root + sub)

    content = await client.get_note_by_keyword("Python", page=1)
    comments = await client.get_note_all_comments(
        content[0],
        crawl_interval=0.5,
        callback=sub_comment_handler
    )

asyncio.run(main())

```

In this example, `get_note_all_comments` automatically invokes `get_comments_all_sub_comments` because the configuration flag is enabled, returning both root and nested comments in a unified list.

## Summary

- **Configuration**: Set `ENABLE_GET_SUB_COMMENTS = True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to activate the feature globally.
- **Entry Point**: Platform clients use `get_note_all_comments`, which internally calls `get_comments_all_sub_comments` after fetching root comments.
- **Workflow**: The sub-comment crawler checks the configuration flag, iterates parent comments with replies, paginates through child API endpoints, extracts data using platform-specific extractors, and supports optional callbacks for streaming.
- **Rate Limiting**: Built-in `crawl_interval` sleeping prevents API abuse across all platforms.
- **Consistency**: The same `ENABLE_GET_SUB_COMMENTS` flag and method patterns work identically across Zhihu, Douyin, Bilibili, Weibo, Tieba, and Kuaishou implementations.

## Frequently Asked Questions

### What happens if ENABLE_GET_SUB_COMMENTS is False?

When `ENABLE_GET_SUB_COMMENTS` remains `False` (the default), the `get_comments_all_sub_comments` method returns an empty list immediately without making any API calls to fetch child comments. Only root-level comments are retrieved and processed, significantly reducing API quota usage and execution time.

### Which platforms support second-level comment crawling?

The implementation in [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) serves as the reference pattern, but the same configuration flag and workflow are implemented across all supported platforms including Douyin, Bilibili, Weibo, Tieba, and Kuaishou. Each platform provides its own `get_child_comments` method and extractor, but the control logic remains consistent.

### How does the callback function work in sub-comment crawling?

The optional `callback` parameter in `get_comments_all_sub_comments` accepts an async function that receives a list of `ZhihuComment` objects (or platform-specific equivalents) after each pagination batch. This enables real-time processing, such as writing to databases or updating progress bars, without waiting for the entire sub-comment tree to load.

### Can I adjust the delay between sub-comment API requests?

Yes, the `crawl_interval` parameter passed to `get_note_all_comments` or `get_comments_all_sub_comments` controls the delay between pagination requests. The method executes `await asyncio.sleep(crawl_interval)` after each batch, allowing you to respect platform rate limits by setting values typically between 0.5 and 2.0 seconds.