How MediaCrawler Handles Slider Verification in Automated Crawling

MediaCrawler solves slider CAPTCHAs by downloading the gap puzzle and background images, using OpenCV template matching to calculate the exact pixel offset, and generating human-like mouse trajectories that accelerate and decelerate to mimic real user behavior.

MediaCrawler is an open-source crawling framework designed to automate data extraction from media platforms that employ anti-bot measures. When encountering slider verification challenges, the crawler leverages computer vision techniques implemented in tools/slider_util.py to programmatically determine the correct slide distance and simulate realistic user interactions.

Image Acquisition and Preprocessing

The Slide class in tools/slider_util.py initiates the verification bypass by handling image inputs through remote URLs or local file paths.

Downloading CAPTCHA Images

The Slide.check_is_img_path() method detects whether supplied URLs point to remote resources. When remote images are detected, MediaCrawler issues HTTP GET requests via httpx with realistic browser headers to avoid simple bot detection. The downloaded bytes decode into NumPy arrays, optionally resize the image, and cache files locally under a ./temp_image/ directory for processing.

Cleaning the Puzzle Piece

Before analysis, the Slide.clear_white() method scans the gap image to remove surrounding empty pixels. This isolates the actual puzzle piece from background whitespace, ensuring that subsequent edge detection operates only on relevant visual data rather than transparent or white borders.

Computer Vision Detection Pipeline

Once images are preprocessed, MediaCrawler applies OpenCV algorithms to locate the puzzle piece within the background.

Edge Detection with Canny Filtering

The Slide.image_edge_detection() method converts both the cleaned gap image and the background into binary edge maps using the Canny filter. This transformation reduces the images to high-contrast outlines, making template matching more reliable across different color schemes and lighting conditions in the CAPTCHA.

Template Matching for Offset Calculation

The Slide.template_match() method executes OpenCV's matchTemplate function with the TM_CCOEFF_NORMED correlation coefficient method. This compares the edge-detected gap image against the background image to find the highest-probability match location. The X-coordinate of this match represents the exact pixel distance the slider must travel, which the discern() method returns as the final offset value.

Generating Human-Like Mouse Movements

Raw pixel distances trigger bot detection, so MediaCrawler creates non-linear movement tracks that replicate human hand tremors and acceleration patterns.

Simple Trajectory Generation

The get_track_simple(distance) function generates a list of incremental moves that accelerate for the first 80% of the distance and then decelerate. This non-uniform velocity profile defeats naive speed checks that flag linear or instantaneous movements as automated.

Configurable Difficulty Levels

The get_tracks(distance, level="easy") function serves as a configuration wrapper. For the default "easy" level, it delegates to get_track_simple(). When requesting harder levels, it employs sophisticated easing-based generators to create more complex human-like curves suitable for advanced anti-bot systems.

Complete Implementation Example

The following pattern demonstrates the integration of image analysis and track generation in MediaCrawler:

from tools.slider_util import Slide, get_tracks

# Initialize with remote URLs or local file paths

slider = Slide(
    gap="https://example.com/captcha/gap.png",
    bg="https://example.com/captcha/background.png"
)

# Calculate the required slide distance in pixels

distance = slider.discern()
print(f"Slider offset: {distance} pixels")

# Generate realistic mouse movement track

track = get_tracks(distance, level="easy")
print(f"Movement increments: {track}")

# Apply to headless browser (Playwright/Selenium example)

# page.mouse.down()

# for step in track:

#     current_x += step

#     page.mouse.move(current_x, start_y)

# page.mouse.up()

The Slide class constructor automatically handles image download and resizing, while the get_tracks() output provides the granular steps needed for browser automation frameworks to complete the verification without manual intervention.

Summary

  • Image handling: Slide.check_is_img_path() in tools/slider_util.py downloads remote CAPTCHA images via httpx and saves them to ./temp_image/.
  • Preprocessing: Slide.clear_white() removes whitespace from puzzle pieces to isolate the target area.
  • Computer vision: Slide.image_edge_detection() applies Canny filters, and Slide.template_match() uses TM_CCOEFF_NORMED to find the exact X-coordinate offset.
  • Movement simulation: get_track_simple() creates acceleration/deceleration patterns, while get_tracks() provides configurable difficulty levels.
  • Implementation: The discern() method returns the pixel distance required to solve the CAPTCHA, enabling automated slider completion.

Frequently Asked Questions

How does MediaCrawler calculate the exact slider distance?

MediaCrawler uses OpenCV template matching in tools/slider_util.py. The Slide.template_match() method runs cv2.matchTemplate with the TM_CCOEFF_NORMED algorithm on edge-detected images to find the best-match position of the gap piece within the background, returning the precise X-coordinate offset via the discern() method.

What makes the mouse tracks appear human-like?

The get_track_simple() function generates movement increments that accelerate during the first 80% of the distance and decelerate toward the end. This variable velocity pattern mimics natural hand movements, unlike linear or constant-speed tracks that anti-bot systems easily flag as automated.

Can MediaCrawler process both local and remote slider images?

Yes. The Slide class constructor accepts both file system paths and HTTP URLs. When remote URLs are supplied, Slide.check_is_img_path() automatically downloads the images using httpx with browser-like headers, stores them temporarily, and converts them into NumPy arrays for OpenCV processing.

Where is the slider verification logic located in the repository?

The core implementation resides in tools/slider_util.py, which contains the Slide class and track generation functions. This module is re-exported through tools/utils.py for convenient access throughout the codebase, with practical usage examples potentially found in platform-specific login handlers like media_platform/weibo/login.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →