Error Handling Strategies for Individual Platform Scrapers in MediaCrawler

MediaCrawler implements a layered error-handling architecture that combines custom exception hierarchies, automatic retries with Tenacity, and graceful degradation to ensure resilient data extraction across Zhihu, XiaoHongShu, Weibo, Douyin, Kuaishou, and BiliBili.

The NanmiCoder/MediaCrawler repository isolates platform-specific failures through targeted error handling strategies that prevent transient network issues from terminating entire crawl sessions. Each platform scraper implements a consistent pattern of custom exceptions, defensive try/except blocks, and automatic retry mechanisms that maintain operational continuity during large-scale data extraction.

Custom Exception Hierarchy for Platform-Specific Failures

Each platform defines lightweight exception classes inheriting from httpx.RequestError to categorize distinct failure modes. The three core exceptions used across all platforms include DataFetchError for generic fetch failures, IPBlockError for rate-limiting scenarios, and ForbiddenError for HTTP 403 responses.

In media_platform/zhihu/exception.py, the hierarchy is defined as:

from httpx import RequestError

class DataFetchError(RequestError):
    """something error when fetch"""

class IPBlockError(RequestError):
    """fetch so fast that the server block us ip"""

class ForbiddenError(RequestError):
    """Forbidden"""

This pattern repeats across media_platform/xhs/exception.py, media_platform/weibo/exception.py, and other platform directories, enabling consistent error classification while allowing platform-specific extensions.

Targeted Try/Except Blocks in Core Crawlers

The core crawler classes wrap high-level operations in targeted exception handling to prevent single-request failures from aborting batch processes. In media_platform/zhihu/core.py, the search method demonstrates this pattern:

try:
    content_list = await self.zhihu_client.get_note_by_keyword(...)
    # Process content...

except DataFetchError:
    utils.logger.error("[ZhihuCrawler.search] Search content error")
    return

Similarly, in media_platform/xhs/core.py, the pagination loop handles failures gracefully:

try:
    notes_res = await self.xhs_client.get_note_by_keyword(...)
    if not notes_res or not notes_res.get("has_more", False):
        break
except DataFetchError:
    utils.logger.error("[XiaoHongShuCrawler.search] Get note detail error")
    break

This approach ensures that DataFetchError terminates only the current operation while allowing the crawler to proceed with subsequent items or pages.

Automatic Retry with Tenacity

The HTTP client wrappers implement transient failure recovery using the Tenacity library. Each platform's client class (e.g., ZhiHuClient, XiaoHongShuClient, WeiboClient) decorates its request method with a retry policy that attempts failed requests up to three times with a one-second delay:

from tenacity import retry, stop_after_attempt, wait_fixed

# media_platform/weibo/client.py

@retry(stop=stop_after_attempt(3), wait=wait_fixed(1))
async def request(self, method, url, **kwargs):
    async with make_async_client(proxy=self.proxy) as client:
        response = await client.request(method, url, timeout=self.timeout, **kwargs)
    return response

This decorator captures temporary network glitches and server-side errors without requiring manual retry loops in the business logic.

Proxy Refresh and Browser Fallbacks

MediaCrawler implements ProxyRefreshMixin across all platform clients to handle proxy expiration before requests execute. When the preferred Chrome DevTools Protocol (CDP) mode fails during browser initialization, the crawlers fall back to standard Playwright launch:


# media_platform/zhihu/core.py

try:
    # Attempt CDP mode launch

    browser = await launch_browser_with_cdp(...)
except Exception as e:
    utils.logger.error(f"[ZhihuCrawler] CDP mode launch failed, falling back: {e}")
    # Fallback to standard mode

    browser = await self.browser_launcher.launch()

This dual-path initialization ensures crawlers remain operational even when advanced debugging features are unavailable in the target environment.

Concurrency Limits and Rate Control

To prevent IPBlockError triggers, each platform scraper implements asyncio.Semaphore to cap parallel requests and configurable sleep intervals between operations:

semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)

# Within fetch operations

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)

The MAX_CONCURRENCY_NUM limits simultaneous connections, while CRAWLER_MAX_SLEEP_SEC introduces mandatory delays after each page or detail fetch, reducing the likelihood of rate-limiting responses from target servers.

Centralized HTTP Status Code Interpretation

The low-level request methods in client files centralize HTTP status mapping to custom exceptions. In media_platform/zhihu/client.py, this logic distinguishes between permanent failures and recoverable states:

if response.status_code != 200:
    if response.status_code == 403:
        raise ForbiddenError(response.text)
    elif response.status_code == 404:
        return {}  # Zhihu returns 404 for notes without comments

    raise DataFetchError(response.text)

This centralized handling ensures that HTTP 403 responses trigger ForbiddenError, while 404 responses (common when comments are disabled) return empty dictionaries rather than raising exceptions, allowing the crawl to continue without interruption.

Summary

  • Custom exceptions (DataFetchError, IPBlockError, ForbiddenError) provide platform-specific error categorization in files like media_platform/zhihu/exception.py and media_platform/xhs/exception.py.
  • Tenacity decorators automatically retry failed requests up to three times with fixed one-second intervals in all platform client classes.
  • Targeted try/except blocks in core crawler methods prevent single failures from terminating entire batch operations across Zhihu, XiaoHongShu, Weibo, and other platforms.
  • ProxyRefreshMixin and CDP fallback mechanisms ensure connectivity resilience across different browser launch modes.
  • Concurrency controls using asyncio.Semaphore and configurable CRAWLER_MAX_SLEEP_SEC intervals mitigate rate-limiting risks.

Frequently Asked Questions

What exceptions does MediaCrawler use for rate limiting?

MediaCrawler defines IPBlockError in each platform's exception.py file to specifically identify when scraping speed triggers server-side IP blocks. This exception inherits from httpx.RequestError and is raised when the platform detects rapid request patterns that result in blocking responses, allowing the crawler to log the incident and potentially adjust its rate.

How does MediaCrawler handle HTTP 403 errors?

When client classes encounter HTTP 403 status codes, they raise ForbiddenError (defined in media_platform/{platform}/exception.py). This allows core crawlers in core.py files to distinguish between authentication failures and transient network errors, typically logging the error and either retrying or aborting the specific operation while preserving the overall crawl session.

Can MediaCrawler recover from failed browser launches?

Yes. The crawler implementations wrap browser initialization in broad exception handlers that catch any launch failure and automatically fall back to standard Playwright launch mode when CDP (Chrome DevTools Protocol) mode fails. This ensures continuous operation across different environment configurations where CDP might be unavailable or restricted.

How many times does MediaCrawler retry failed requests?

According to the Tenacity configuration in client.py files across all platforms (Zhihu, XiaoHongShu, Weibo, Douyin, Kuaishou, and BiliBili), MediaCrawler retries failed requests three times with a one-second fixed delay between attempts, as specified by @retry(stop=stop_after_attempt(3), wait=wait_fixed(1)).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →