# Error Handling Strategies for Individual Platform Scrapers in MediaCrawler

> Discover MediaCrawler's robust error handling strategies for platform scrapers. Learn how custom exceptions, automatic retries, and graceful degradation ensure reliable data extraction.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: best-practices
- Published: 2026-07-03

---

**MediaCrawler implements a layered error-handling architecture that combines custom exception hierarchies, automatic retries with Tenacity, and graceful degradation to ensure resilient data extraction across Zhihu, XiaoHongShu, Weibo, Douyin, Kuaishou, and BiliBili.**

The NanmiCoder/MediaCrawler repository isolates platform-specific failures through targeted error handling strategies that prevent transient network issues from terminating entire crawl sessions. Each platform scraper implements a consistent pattern of custom exceptions, defensive try/except blocks, and automatic retry mechanisms that maintain operational continuity during large-scale data extraction.

## Custom Exception Hierarchy for Platform-Specific Failures

Each platform defines lightweight exception classes inheriting from `httpx.RequestError` to categorize distinct failure modes. The three core exceptions used across all platforms include **DataFetchError** for generic fetch failures, **IPBlockError** for rate-limiting scenarios, and **ForbiddenError** for HTTP 403 responses.

In [`media_platform/zhihu/exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/exception.py), the hierarchy is defined as:

```python
from httpx import RequestError

class DataFetchError(RequestError):
    """something error when fetch"""

class IPBlockError(RequestError):
    """fetch so fast that the server block us ip"""

class ForbiddenError(RequestError):
    """Forbidden"""

```

This pattern repeats across [`media_platform/xhs/exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/exception.py), [`media_platform/weibo/exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/exception.py), and other platform directories, enabling consistent error classification while allowing platform-specific extensions.

## Targeted Try/Except Blocks in Core Crawlers

The core crawler classes wrap high-level operations in targeted exception handling to prevent single-request failures from aborting batch processes. In [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py), the `search` method demonstrates this pattern:

```python
try:
    content_list = await self.zhihu_client.get_note_by_keyword(...)
    # Process content...

except DataFetchError:
    utils.logger.error("[ZhihuCrawler.search] Search content error")
    return

```

Similarly, in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py), the pagination loop handles failures gracefully:

```python
try:
    notes_res = await self.xhs_client.get_note_by_keyword(...)
    if not notes_res or not notes_res.get("has_more", False):
        break
except DataFetchError:
    utils.logger.error("[XiaoHongShuCrawler.search] Get note detail error")
    break

```

This approach ensures that **DataFetchError** terminates only the current operation while allowing the crawler to proceed with subsequent items or pages.

## Automatic Retry with Tenacity

The HTTP client wrappers implement transient failure recovery using the **Tenacity** library. Each platform's client class (e.g., `ZhiHuClient`, `XiaoHongShuClient`, `WeiboClient`) decorates its request method with a retry policy that attempts failed requests up to three times with a one-second delay:

```python
from tenacity import retry, stop_after_attempt, wait_fixed

# media_platform/weibo/client.py

@retry(stop=stop_after_attempt(3), wait=wait_fixed(1))
async def request(self, method, url, **kwargs):
    async with make_async_client(proxy=self.proxy) as client:
        response = await client.request(method, url, timeout=self.timeout, **kwargs)
    return response

```

This decorator captures temporary network glitches and server-side errors without requiring manual retry loops in the business logic.

## Proxy Refresh and Browser Fallbacks

MediaCrawler implements **ProxyRefreshMixin** across all platform clients to handle proxy expiration before requests execute. When the preferred Chrome DevTools Protocol (CDP) mode fails during browser initialization, the crawlers fall back to standard Playwright launch:

```python

# media_platform/zhihu/core.py

try:
    # Attempt CDP mode launch

    browser = await launch_browser_with_cdp(...)
except Exception as e:
    utils.logger.error(f"[ZhihuCrawler] CDP mode launch failed, falling back: {e}")
    # Fallback to standard mode

    browser = await self.browser_launcher.launch()

```

This dual-path initialization ensures crawlers remain operational even when advanced debugging features are unavailable in the target environment.

## Concurrency Limits and Rate Control

To prevent **IPBlockError** triggers, each platform scraper implements **asyncio.Semaphore** to cap parallel requests and configurable sleep intervals between operations:

```python
semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)

# Within fetch operations

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)

```

The `MAX_CONCURRENCY_NUM` limits simultaneous connections, while `CRAWLER_MAX_SLEEP_SEC` introduces mandatory delays after each page or detail fetch, reducing the likelihood of rate-limiting responses from target servers.

## Centralized HTTP Status Code Interpretation

The low-level `request` methods in client files centralize HTTP status mapping to custom exceptions. In [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py), this logic distinguishes between permanent failures and recoverable states:

```python
if response.status_code != 200:
    if response.status_code == 403:
        raise ForbiddenError(response.text)
    elif response.status_code == 404:
        return {}  # Zhihu returns 404 for notes without comments

    raise DataFetchError(response.text)

```

This centralized handling ensures that HTTP 403 responses trigger **ForbiddenError**, while 404 responses (common when comments are disabled) return empty dictionaries rather than raising exceptions, allowing the crawl to continue without interruption.

## Summary

- **Custom exceptions** (`DataFetchError`, `IPBlockError`, `ForbiddenError`) provide platform-specific error categorization in files like [`media_platform/zhihu/exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/exception.py) and [`media_platform/xhs/exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/exception.py).
- **Tenacity decorators** automatically retry failed requests up to three times with fixed one-second intervals in all platform client classes.
- **Targeted try/except blocks** in core crawler methods prevent single failures from terminating entire batch operations across Zhihu, XiaoHongShu, Weibo, and other platforms.
- **ProxyRefreshMixin** and CDP fallback mechanisms ensure connectivity resilience across different browser launch modes.
- **Concurrency controls** using `asyncio.Semaphore` and configurable `CRAWLER_MAX_SLEEP_SEC` intervals mitigate rate-limiting risks.

## Frequently Asked Questions

### What exceptions does MediaCrawler use for rate limiting?

MediaCrawler defines **IPBlockError** in each platform's [`exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/exception.py) file to specifically identify when scraping speed triggers server-side IP blocks. This exception inherits from `httpx.RequestError` and is raised when the platform detects rapid request patterns that result in blocking responses, allowing the crawler to log the incident and potentially adjust its rate.

### How does MediaCrawler handle HTTP 403 errors?

When client classes encounter HTTP 403 status codes, they raise **ForbiddenError** (defined in `media_platform/{platform}/exception.py`). This allows core crawlers in [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py) files to distinguish between authentication failures and transient network errors, typically logging the error and either retrying or aborting the specific operation while preserving the overall crawl session.

### Can MediaCrawler recover from failed browser launches?

Yes. The crawler implementations wrap browser initialization in broad exception handlers that catch any launch failure and automatically fall back to standard Playwright launch mode when CDP (Chrome DevTools Protocol) mode fails. This ensures continuous operation across different environment configurations where CDP might be unavailable or restricted.

### How many times does MediaCrawler retry failed requests?

According to the Tenacity configuration in [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py) files across all platforms (Zhihu, XiaoHongShu, Weibo, Douyin, Kuaishou, and BiliBili), MediaCrawler retries failed requests **three times** with a **one-second fixed delay** between attempts, as specified by `@retry(stop=stop_after_attempt(3), wait=wait_fixed(1))`.