How MediaCrawler Handles Sliding Captcha Verification: OpenCV and Track Simulation

MediaCrawler solves sliding captcha verification by downloading gap and background images, using OpenCV template matching to calculate the target offset, and generating human-like mouse tracks to simulate natural slider movement before executing the slide via browser automation.

Sliding captcha verification presents a significant hurdle for automated data collection, requiring precise visual analysis and human-like interaction patterns to avoid detection. The MediaCrawler repository implements a deterministic computer vision pipeline that transforms visual puzzle-solving into programmatic slider manipulation. According to the source code in the NanmiCoder/MediaCrawler project, the system combines edge detection algorithms, template matching, and physics-based motion simulation to bypass verification gates on platforms like Douyin and Weibo.

The Sliding Captcha Workflow

The complete sliding captcha verification pipeline in MediaCrawler consists of four distinct phases: image acquisition, preprocessing and analysis, track generation, and browser execution. All core logic resides in tools/slider_util.py, which provides the Slide class for image handling and the discern() method for coordinate calculation. Higher-level login modules in media_platform/douyin/login.py and similar files consume these utilities through the utils.get_tracks() interface.

Image Acquisition and Preprocessing

Downloading Gap and Background Images

MediaCrawler initiates the verification process by fetching the two essential visual components: the slider gap (puzzle piece) and the background image containing the missing section. The Slide.check_is_img_path method handles these downloads using httpx, automatically storing retrieved images in the temp_image/ directory for processing.

Cleaning and Greyscale Conversion

Before computer vision analysis, the downloaded gap image undergoes preprocessing via Slide.clear_white to crop surrounding whitespace that could interfere with matching algorithms. Both images are then converted to greyscale to standardize the input for OpenCV operations. This preprocessing ensures that template matching focuses on the structural edges of the puzzle piece rather than color variations or background artifacts.

Detecting the Target Offset with Computer Vision

The system locates the correct slider position using a two-stage computer vision approach implemented in the Slide.discern() method. First, cv2.Canny applies edge detection to highlight the轮廓 of the puzzle piece within the background noise. Then, cv2.matchTemplate performs normalized cross-correlation matching between the preprocessed gap image and the background canvas.

The matching operation returns the x-coordinate representing the horizontal distance the slider must travel to align the puzzle piece correctly. This offset value becomes the input parameter for track generation. For debugging purposes, MediaCrawler writes the matched result visualization to out.jpg in the project root, allowing developers to verify that the detected position aligns with the actual gap location.

Generating Human-Like Mouse Tracks

Raw distance values cannot be sent directly to browser automation APIs, as instantaneous movement triggers anti-bot detection. MediaCrawler solves this by converting pixel distances into acceleration curves that mimic human hand tremor and momentum.

Simple Acceleration with get_track_simple

For the default easy difficulty level, the system employs get_track_simple to create a piecewise-accelerated track. This algorithm generates a list of incremental pixel movements that start slowly, accelerate through the middle of the motion, and decelerate as the slider approaches the target.

Advanced Easing Curves

When sophisticated motion patterns are required, MediaCrawler imports tools/easing.py to access physics-based interpolation functions. The easing.get_tracks method produces "ease-out-expo" curves that simulate the natural deceleration of human muscle control, creating tracks with variable timing between steps that statistical bot detection models classify as organic.

Executing the Slide Action

The generated track list feeds into the browser automation layer, which supports Playwright, Selenium, or Chrome DevTools Protocol (CDP) drivers. The implementation iterates through each offset in the track, dispatching individual mouse move events to create the appearance of continuous motion:

from tools.slider_util import Slide
from tools.utils import get_tracks

# 1. Load the captcha images and compute the distance to slide

slide = Slide(gap_url, bg_url)          # gap and background URLs from the login page

distance = slide.discern()              # → x-coordinate of the correct position

# 2. Generate a human-like movement track

track = get_tracks(distance, level="easy")   # list of pixel moves

# 3. Feed the track to the browser (Playwright example)

for dx in track:
    page.mouse.move(current_x + dx, slider_y, steps=1)
    current_x += dx
page.mouse.up()

Each incremental movement updates the slider position by a small delta, with steps=1 ensuring that the browser registers distinct motion events rather than teleporting the cursor to the final destination.

Integration in Login Modules

Platform-specific login implementations abstract the verification complexity through the utils.get_tracks(distance, level) helper function. Files such as media_platform/douyin/login.py and media_platform/weibo/login.py extract captcha URLs from the login page DOM, instantiate the Slide class, and orchestrate the slide action without containing low-level OpenCV or easing logic. This modular architecture allows the MediaCrawler project to maintain consistent sliding captcha verification behavior across multiple social media platforms while centralizing image processing utilities in the tools/ directory.

Summary

  • Image Processing: MediaCrawler downloads captcha components via httpx and preprocesses them using Slide.clear_white and greyscale conversion in tools/slider_util.py.
  • Offset Detection: The Slide.discern() method employs cv2.Canny edge detection and cv2.matchTemplate to calculate the exact horizontal distance required to solve the puzzle.
  • Motion Simulation: Distance values transform into human-like tracks using either get_track_simple for basic acceleration or easing.get_tracks for advanced "ease-out-expo" curves defined in tools/easing.py.
  • Browser Execution: Incremental mouse movements execute through Playwright or Selenium, with tracks generated by utils.get_tracks() consumed by login modules like media_platform/douyin/login.py.

Frequently Asked Questions

How does MediaCrawler calculate the exact distance for a sliding captcha?

MediaCrawler calculates the exact distance using OpenCV template matching in the Slide.discern() method. The system first applies cv2.Canny edge detection to both the gap image and background, then uses cv2.matchTemplate to find the correlation peak that indicates where the puzzle piece fits, returning the x-coordinate as the slide distance.

What makes the mouse tracks generated by MediaCrawler appear human-like?

The tracks simulate human motor control by using acceleration curves rather than linear movement. The get_track_simple function creates piecewise acceleration with variable step distances, while tools/easing.py provides "ease-out-expo" functions that mimic natural deceleration patterns, preventing instantaneous velocity changes that trigger anti-bot systems.

Which file contains the core logic for sliding captcha verification in MediaCrawler?

The core logic resides in tools/slider_util.py, which contains the Slide class for image handling and the discern() method for offset calculation. Supporting utilities for track generation exist in tools/easing.py, while tools/utils.py re-exports the get_tracks function for convenient consumption by login modules.

How does MediaCrawler handle the image processing before template matching?

Before template matching, MediaCrawler downloads images to temp_image/ and processes them through Slide.clear_white to remove surrounding whitespace from the gap image. Both images undergo greyscale conversion to standardize input for OpenCV operations, ensuring that the cv2.matchTemplate algorithm focuses on structural edges rather than color data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →