Handling Anti-Crawling and Sliding CAPTCHA Verification in MediaCrawler: A Complete Guide

MediaCrawler implements a two-layer defense system that masks browser automation fingerprints and solves sliding CAPTCHAs through computer vision and human-like mouse simulation.

MediaCrawler is an open-source scraping framework designed for Chinese social media platforms like Douyin and Tieba. Handling anti-crawling and sliding CAPTCHA verification in MediaCrawler requires understanding both browser-level stealth techniques and precise mechanical puzzle solving.

Browser-Level Anti-Detection Techniques

MediaCrawler masks automation signatures by injecting JavaScript that runs on every page load. When a new browser context is created via Playwright, the crawler uses the add_init_script API to execute a lightweight snippet that hides typical headless browser indicators.

In media_platform/tieba/core.py at line 91, the injection routine loads code that overrides navigator.webdriver, removes Chrome-driver residues, and synthesizes realistic plugin and language lists. This makes the headless browser appear as a standard Chrome instance to platform detection algorithms.

Solving Sliding CAPTCHA Challenges

For platforms presenting "drag-the-puzzle-piece" verifications, MediaCrawler executes a four-phase pipeline implemented across media_platform/douyin/login.py and tools/slider_util.py.

Image Capture and Gap Detection

The process begins in media_platform/douyin/login.py at line 23, where the move_slider coroutine fetches both the background image via back_selector and the puzzle piece via gap_selector. These images are passed to the Slide class defined in tools/slider_util.py at line 34.

The Slide.discern() method downloads the images and preprocesses them using clear_white to remove surrounding whitespace. It then employs OpenCV's template_match function to locate the exact x-coordinate where the puzzle piece fits into the background gap.

Human-Like Movement Generation

Once the offset distance is calculated, MediaCrawler generates a realistic mouse trajectory. The get_tracks helper in tools/slider_util.py at line 78 accepts a slider_level argument that selects between two algorithms:

  • get_track_simple: A basic acceleration/deceleration model for straightforward movements
  • easing.get_tracks: Advanced easing functions from tools/easing.py that create more human-like curves with variable velocity

Executing the Drag Operation

Using Playwright's mouse API, the solver moves the cursor step-by-step along the generated track array. The check_page_display_slider method in the Douyin login flow orchestrates this by accepting parameters like move_step=12 and slider_level="hard" to fine-tune the behavior. Final offset corrections ensure the total dragged distance exactly matches the detected gap coordinate.

Implementation Examples

The following patterns demonstrate how to implement these anti-crawling measures in your own MediaCrawler extensions.

To solve a Douyin sliding CAPTCHA with hard difficulty:

from media_platform.douyin.login import DouYinLogin
from tools import utils

async def login_with_slider(page):
    douyin = DouYinLogin(page)
    await douyin.check_page_display_slider(move_step=12, slider_level="hard")
    # Executes: image capture → gap detection → track building → drag simulation

To inject anti-detection scripts for Tieba:

from media_platform.tieba.core import TieBaCrawler

async def start_crawler():
    crawler = TieBaCrawler()
    await crawler._inject_anti_detection_scripts()
    # Masks webdriver, plugins, languages, and other automation fingerprints

To generate movement tracks independently:

from tools.slider_util import get_tracks

distance = 120  # pixels detected by OpenCV

track = get_tracks(distance, slider_level="easy")  # Returns: [15, 23, 30, …]

Key files referenced in this implementation include:

Summary

  • MediaCrawler uses Playwright's add_init_script API to inject JavaScript that masks navigator.webdriver and other automation fingerprints on every page load.
  • The sliding CAPTCHA solver in tools/slider_util.py combines OpenCV template matching with configurable movement tracks to simulate human drag behavior.
  • Two difficulty levels ("easy" and "hard") control whether the crawler uses simple acceleration or advanced easing functions for mouse movement.
  • Implementation spans media_platform/douyin/login.py for execution and media_platform/tieba/core.py for browser stealth initialization.

Frequently Asked Questions

How does MediaCrawler avoid browser automation detection?

MediaCrawler injects a JavaScript snippet via Playwright's add_init_script that overrides properties like navigator.webdriver, removes Chrome-driver residues, and fabricates realistic plugin and language lists. This injection occurs in media_platform/tieba/core.py and runs on every page load to present a standard Chrome signature to anti-bot systems.

What algorithm does MediaCrawler use to solve sliding CAPTCHAs?

The solver uses OpenCV's template matching algorithm via the Slide class in tools/slider_util.py. It compares the puzzle piece image against the background to find the exact x-coordinate offset, then generates a human-like mouse trajectory using either simple acceleration or easing functions depending on the slider_level parameter.

Can the slider solver be used for platforms other than Douyin?

Yes. While the implementation example shows Douyin in media_platform/douyin/login.py, the core utilities in tools/slider_util.py and tools/easing.py are platform-agnostic. You can import the Slide class and get_tracks function to implement similar solvers for any platform presenting comparable gap-based verification challenges.

Where is the anti-detection script injected in the codebase?

The injection routine is located in media_platform/tieba/core.py at line 91. This method initializes the Playwright context with stealth scripts before navigating to target URLs, ensuring that automation fingerprints are masked from the initial connection handshake.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →