# MediaCrawler | 程序员阿江-Relakkes | Knowledge Base | Instagit

小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 ｜ 评论爬虫、微博帖子 ｜ 评论爬虫、百度贴吧帖子 ｜ 百度贴吧评论回复爬虫  | 知乎问答文章｜评论爬虫

GitHub Stars: 53.3k

Repository: https://github.com/NanmiCoder/MediaCrawler

---

## Articles

### [How to Troubleshoot "No Such Table" Database Initialization Errors in MediaCrawler](/NanmiCoder/MediaCrawler/troubleshoot-no-such-table-database-initialization-errors)

Fix "no such table" database errors in MediaCrawler. Learn how to resolve SQLAlchemy ORM initialization issues by using the --init_db flag or checking SAVE_DATA_OPTION.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Configure a Custom User Agent and Browser Fingerprint in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-custom-user-agent-and-browser-fingerprint)

Learn to configure custom user agent and browser fingerprint in MediaCrawler. Modify viewport, locale, and timezone using CDPBrowserManager and context_options for advanced control.

- Tags: how-to-guide
- Published: 2026-08-14

### [Async Architecture and Concurrency Control (`MAX_CONCURRENCY_NUM`) in MediaCrawler](/NanmiCoder/MediaCrawler/mediacrawler-async-architecture-and-concurrency-control)

Explore MediaCrawler's async architecture and MAX_CONCURRENCY_NUM for safe, serialized crawling. Learn how Python's asyncio and semaphores manage simultaneous operations effectively.

- Tags: internals
- Published: 2026-08-14

### [How to Configure Media File (Image/Video) Download and Storage in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-media-file-download-and-storage)

Learn how to configure media file downloads and storage in MediaCrawler. Easily set save options, enable downloads, and customize storage paths for your images and videos.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Configure MediaCrawler for International Xiaohongshu (rednote.com) vs the Domestic Version](/NanmiCoder/MediaCrawler/configure-mediacrawler-international-vs-domestic-xiaohongshu)

Learn to configure MediaCrawler for international Xiaohongshu rednote.com vs domestic versions. Edit BASE_URL and COOKIE_DOMAIN constants in xhs_config.py to switch between www.xiaohongshu.com and www.rednote.com for targeted c...

- Tags: how-to-guide
- Published: 2026-08-14

### [MediaCrawler Data Deduplication Strategies for Database Storage Mode](/NanmiCoder/MediaCrawler/mediacrawler-data-deduplication-strategies-database-mode)

Discover MediaCrawler's data deduplication strategies for database storage. Learn how upsert and unique MongoDB indexes prevent duplicate records effectively.

- Tags: best-practices
- Published: 2026-08-14

### [How to Implement Breakpoint Resume for Large-Scale Crawling Jobs in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-implement-breakpoint-resume-for-large-scale-crawling)

Learn to implement breakpoint resume for large-scale crawling jobs in MediaCrawler. This guide shows how to use the --start CLI argument and JSON checkpointing for seamless restarts.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Switch Between Headless and Non-Headless Browser Modes in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-switch-between-headless-and-non-headless-browser-modes)

Easily switch MediaCrawler between headless and non-headless browser modes. Learn to configure HEADLESS and CDP_HEADLESS flags via CLI, config file, or code for flexible operation.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Run the MediaCrawler WebUI for Non-Command-Line Usage](/NanmiCoder/MediaCrawler/how-to-run-mediacrawler-webui-interface)

Easily run the MediaCrawler WebUI without the command line. Launch the graphical interface via development mode or build static assets for production deployment.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Debug Login Failures (QR Code Scanning, Phone Verification) in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-debug-login-failures-in-mediacrawler)

Debug MediaCrawler QR code login failures with visible browsers, longer timeouts, and URL logging. Resolve phone verification redirect issues effectively.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Configure Rate Limiting and Crawl Intervals in MediaCrawler to Avoid Platform Bans](/NanmiCoder/MediaCrawler/how-to-configure-rate-limiting-and-crawl-intervals)

Learn to configure rate limiting and crawl intervals in MediaCrawler to prevent platform bans. Set global or custom intervals to manage request pacing effectively and ensure smooth crawling.

- Tags: best-practices
- Published: 2026-08-14

### [How to Manage Browser Contexts and Prevent Memory Leaks in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-manage-browser-contexts-and-prevent-memory-leaks)

Learn how MediaCrawler manages browser contexts and prevents memory leaks using isolated incognito sessions and explicit cleanup for efficient crawling.

- Tags: performance
- Published: 2026-08-14

### [How to Extend MediaCrawler to Support a New Social Media Platform: A Complete Developer Guide](/NanmiCoder/MediaCrawler/how-to-extend-mediacrawler-for-new-platform)

Learn to extend MediaCrawler for new platforms by implementing four abstract classes and registering your platform. A complete developer guide for NanmiCoder/MediaCrawler.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Generate a Word Cloud from Crawled Comments Using MediaCrawler](/NanmiCoder/MediaCrawler/how-to-generate-word-cloud-from-crawled-comments)

Learn to generate a word cloud from crawled comments using MediaCrawler. This guide covers tokenization, stopword filtering, and visualization with Python's wordcloud library.

- Tags: tutorial
- Published: 2026-08-14

### [How to Configure Playwright Proxy Settings with Different Proxy Providers in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-playwright-proxy-settings)

Learn to configure Playwright proxy settings in MediaCrawler. Easily integrate with providers like kuaidaili, wandouhttp, or static IPs by enabling proxy settings and specifying your provider name.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Implement Second-Level Comment Crawling (Nested Replies) in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-implement-second-level-comment-crawling)

Learn how to implement second-level comment crawling in MediaCrawler using the get_comments_all_sub_comments coroutine. Discover how to detect, paginate, and fetch nested replies efficiently.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Handle Login State Persistence and Cookie-Based Authentication in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-handle-login-state-persistence-and-cookie-authentication)

Learn how MediaCrawler manages login state persistence and cookie authentication using persistent user data directories, cookie conversion, and platform-specific verification.

- Tags: how-to-guide
- Published: 2026-08-14

### [How the CrawlerFactory Pattern Creates Platform-Specific Crawler Instances in MediaCrawler](/NanmiCoder/MediaCrawler/how-crawlerfactory-pattern-works-in-mediacrawler)

Learn how the CrawlerFactory pattern in MediaCrawler efficiently creates platform-specific crawler instances using a static registry to decouple instantiation from business logic.

- Tags: internals
- Published: 2026-08-14

### [MediaCrawler Storage Layer Configuration: SQLite, MySQL, MongoDB, JSON/CSV/Excel Options Explained](/NanmiCoder/MediaCrawler/mediacrawler-storage-layer-configuration-options)

Explore MediaCrawler storage options: SQLite, MySQL, MongoDB, and file formats like JSON, CSV, Excel. Configure your data backend easily with NanmiCoder.

- Tags: configuration
- Published: 2026-08-14

### [How to Implement Multi-Account Rotation with IP Proxy Pools in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-implement-multi-account-rotation-with-ip-proxy-pools)

Learn to implement multi-account rotation with IP proxy pools in MediaCrawler. Automate IP and user credential rotation using built-in proxy pooling and account manager.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Configure CDP Mode and Connect to an Existing Chrome Browser in MediaCrawler for Reduced Detection](/NanmiCoder/MediaCrawler/how-to-configure-cdp-mode-and-connect-to-existing-chrome-browser-for-reduced-detection)

Learn how to configure CDP mode and connect to an existing Chrome browser in MediaCrawler for reduced detection. Follow simple steps to enhance your crawling experience.

- Tags: how-to-guide
- Published: 2026-08-14

### [How MediaCrawler Handles Platform-Specific API Signatures and JavaScript Injection](/NanmiCoder/MediaCrawler/how-does-mediacrawler-handle-platform-specific-api-signatures-and-js-injection)

Discover how MediaCrawler tackles platform-specific API signatures and JS injection using dedicated modules and stealth injection to bypass anti-bot measures for popular platforms.

- Tags: deep-dive
- Published: 2026-08-14

### [MediaCrawler Usage Examples: CLI, API, WebSocket, and Python SDK Patterns](/NanmiCoder/MediaCrawler/are-there-any-examples-of-mediacrawler-usage)

Explore MediaCrawler usage examples including CLI, API, WebSocket, and Python SDK patterns for scraping Xiaohongshu, Douyin, Bilibili, and Weibo data. Get started today.

- Tags: getting-started
- Published: 2026-08-12

### [How to Report a Bug in MediaCrawler: The Complete GitHub Issues Guide](/NanmiCoder/MediaCrawler/how-to-report-a-bug-in-mediacrawler)

Learn how to report a bug in MediaCrawler using the complete GitHub Issues guide. Follow the template and checklist for faster fixes and help the developers.

- Tags: how-to-guide
- Published: 2026-08-12

### [What License Is MediaCrawler Distributed Under?](/NanmiCoder/MediaCrawler/what-license-is-mediacrawler-distributed-under)

Discover the terms of MediaCrawler's Non-Commercial Learning License 1.1. Learn how you can use modify and redistribute this software for non-commercial learning and research.

- Tags: faq
- Published: 2026-08-12

### [How to Contribute to NanmiCoder/MediaCrawler: A Complete Guide for Developers](/NanmiCoder/MediaCrawler/how-to-contribute-to-nanmocoder-mediacrawler)

Learn how to contribute to NanmiCoderMediaCrawler. Fork the repo, set up your environment with uv, run tests, implement changes, and submit a pull request to join the development.

- Tags: how-to-guide
- Published: 2026-08-12

### [What is the Underlying Technology Stack of MediaCrawler? A Deep Dive into This Async Python Framework](/NanmiCoder/MediaCrawler/what-is-the-underlying-technology-stack-of-mediacrawler)

Explore the async Python framework MediaCrawler its tech stack including httpx Playwright SQLAlchemy and MongoDB Understand its architecture for scraping Chinese media platforms

- Tags: deep-dive
- Published: 2026-08-12

### [How to Parse Scraped Data with MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/how-to-parse-scraped-data-from-mediacrawler)

Learn how to parse scraped data from MediaCrawler using pure-function helpers. Convert raw URLs, HTML, and API responses into strongly-typed data objects with this complete guide.

- Tags: how-to-guide
- Published: 2026-08-12

### [MediaCrawler Logging: How the Open-Source Crawler Handles Python Logging](/NanmiCoder/MediaCrawler/what-are-the-logging-capabilities-of-mediacrawler)

Explore MediaCrawler's robust Python logging capabilities. Discover its centralized setup, structured output, configurable levels, and noise suppression for seamless debugging.

- Tags: how-to-guide
- Published: 2026-08-12

### [Does MediaCrawler Support Rate Limiting? Yes—Here's How It Works](/NanmiCoder/MediaCrawler/does-mediacrawler-support-rate-limiting)

Learn how MediaCrawler supports rate limiting with configurable sleep intervals, concurrency caps, proxy rotation, and exponential back-off retries. Get started today.

- Tags: how-to-guide
- Published: 2026-08-12

### [How to Integrate MediaCrawler with Other Tools: A Complete Integration Guide](/NanmiCoder/MediaCrawler/how-to-integrate-mediacrawler-with-other-tools)

Learn how to integrate MediaCrawler with other tools using Python imports, REST API calls, or the WebUI for seamless automation and embedding. Explore the complete integration guide.

- Tags: how-to-guide
- Published: 2026-08-12

### [MediaCrawler API: FastAPI HTTP & WebSocket Interface Explained](/NanmiCoder/MediaCrawler/is-there-an-api-for-mediacrawler)

Explore the MediaCrawler API with FastAPI. Control crawls, monitor status, and access data programmatically via its HTTP and WebSocket interfaces.

- Tags: api-reference
- Published: 2026-08-12

### [What Data Formats Does MediaCrawler Output? Complete Guide to File and Database Export Options](/NanmiCoder/MediaCrawler/what-data-formats-does-mediacrawler-output)

Explore MediaCrawler's seven data output formats including CSV, JSON, JSONL, Excel, SQLite, MySQL, and PostgreSQL. Learn to configure export options easily.

- Tags: api-reference
- Published: 2026-08-12

### [How to Extend MediaCrawler with Custom Crawlers: A Step‑by‑Step Guide](/NanmiCoder/MediaCrawler/can-mediacrawler-be-extended-with-custom-crawlers)

Extend MediaCrawler with custom crawlers using its simple three-method interface and factory registration pattern. Build your own crawlers easily with this step-by-step guide for NanmiCoder/MediaCrawler.

- Tags: how-to-guide
- Published: 2026-08-12

### [MediaCrawler Architecture: A Complete Guide to Its Main Modules](/NanmiCoder/MediaCrawler/what-are-the-main-modules-in-mediacrawler-architecture)

Explore the MediaCrawler architecture and its eight main modules: Base Engine, Site-Specific Crawlers, Data Storage, Proxy System, and more. Learn how this modular framework scrapes Chinese social media.

- Tags: architecture
- Published: 2026-08-12

### [MediaCrawler Command Line: A Complete CLI Guide for NanmiCoder/MediaCrawler](/NanmiCoder/MediaCrawler/how-to-run-mediacrawler-from-command-line)

Master the MediaCrawler CLI with NanmiCoder. Learn to run MediaCrawler from the command line using python main.py or as a module for efficient media crawling.

- Tags: how-to-guide
- Published: 2026-08-12

### [How to Configure MediaCrawler Settings: Complete Guide to Python Config Files and CLI Options](/NanmiCoder/MediaCrawler/how-to-configure-mediacrawler-settings)

Learn how to configure MediaCrawler settings easily by editing Python config files or using CLI options. Customize your media crawling experience without altering source code.

- Tags: how-to-guide
- Published: 2026-08-12

### [Setting Up Distributed Crawling with Redis Cache in MediaCrawler](/NanmiCoder/MediaCrawler/setting-up-distributed-crawling-redis-cache-media-crawler)

Set up distributed crawling with Redis cache in MediaCrawler. Coordinate multiple workers, share sessions, throttle requests, and deduplicate tasks globally for efficient web scraping.

- Tags: how-to-guide
- Published: 2026-07-31

### [Integrating Third-Party Proxy Provider APIs with MediaCrawler: A Complete Implementation Guide](/NanmiCoder/MediaCrawler/integrating-third-party-proxy-provider-apis-media-crawler)

Easily integrate third-party proxy provider APIs with MediaCrawler. Our guide shows how to implement a single method for seamless proxy handling in your crawling projects.

- Tags: how-to-guide
- Published: 2026-07-31

### [Handling Session Expiration and Re-Authentication Flow in MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/handling-session-expiration-re-authentication-flow-media-crawler)

Master session expiration and re-authentication in MediaCrawler. Learn how our automatic system ensures uninterrupted crawls across platforms like Zhihu and XiaoHongShu. Get the complete guide now.

- Tags: how-to-guide
- Published: 2026-07-31

### [Enabling Media Downloading During MediaCrawler Crawling: Images, Videos, and Music](/NanmiCoder/MediaCrawler/enabling-media-downloading-images-videos-media-crawler-crawling)

Easily enable media downloading for images, videos, and music with MediaCrawler. Its CDP integration automatically extracts and saves media URLs during crawling on platforms like Douyin and Bilibili. Get started now.

- Tags: how-to-guide
- Published: 2026-07-31

### [How to Specify CUSTOM_BROWSER_PATH for Chrome or Edge in MediaCrawler CDP Mode](/NanmiCoder/MediaCrawler/specifying-custom-browser-path-chrome-edge-media-crawler-cdp-mode)

Learn to specify CUSTOM_BROWSER_PATH for Chrome or Edge in MediaCrawler CDP mode. Configure the browser path in base_config.py to launch specific browser versions via CDP when CDP_CONNECT_EXISTING is False.

- Tags: how-to-guide
- Published: 2026-07-31

### [Switching MediaCrawler Login from QR Code to Phone Number Authentication: A Complete Guide](/NanmiCoder/MediaCrawler/switching-media-crawler-login-qr-code-phone-number-authentication)

Easily switch MediaCrawler login from QR code to phone number authentication. Learn how to update your configuration or use command-line arguments for a seamless transition.

- Tags: how-to-guide
- Published: 2026-07-31

### [Implementing Request Rate Limiting in MediaCrawler: A Complete Configuration Guide](/NanmiCoder/MediaCrawler/implementing-request-rate-limiting-media-crawler)

Configure request rate limiting in MediaCrawler using CRAWLER_MAX_SLEEP_SEC and MAX_CONCURRENCY_NUM. Prevent IP throttling and optimize crawling performance with this essential guide.

- Tags: how-to-guide
- Published: 2026-07-31

### [Using MediaCrawler FastAPI WebUI for Programmatic Crawler Control](/NanmiCoder/MediaCrawler/using-media-crawler-fastapi-webui-programmatic-crawler-control)

Control MediaCrawler programmatically with its FastAPI WebUI. Start, monitor, and stop crawlers via REST endpoints without altering core logic. Integrate seamless automation today.

- Tags: how-to-guide
- Published: 2026-07-31

### [Exporting MediaCrawler Data to Excel with Proper Formatting](/NanmiCoder/MediaCrawler/exporting-media-crawler-data-excel-proper-formatting)

Easily export MediaCrawler data to Excel with professional formatting. MediaCrawler's Excel export layer automatically adds blue headers auto-sized columns and thin borders for clean data presentation.

- Tags: how-to-guide
- Published: 2026-07-31

### [Handling Anti-Crawling and Sliding CAPTCHA Verification in MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/handling-anti-crawling-sliding-captcha-verification-media-crawler)

Learn to bypass anti-crawling and sliding CAPTCHA verification in MediaCrawler. Our guide explains how computer vision and human-like mouse simulation solve these challenges.

- Tags: how-to-guide
- Published: 2026-07-31

### [Tuning MediaCrawler MAX_CONCURRENCY_NUM and CRAWLER_MAX_SLEEP_SEC for Performance](/NanmiCoder/MediaCrawler/tuning-media-crawler-max-concurrency-num-crawler-max-sleep-sec-performance)

Optimize MediaCrawler performance by tuning MAX_CONCURRENCY_NUM and CRAWLER_MAX_SLEEP_SEC. Learn how to balance speed and rate limits for efficient web crawling.

- Tags: performance
- Published: 2026-07-31

### [How to Add a New Social Media Platform to MediaCrawler: A Complete Implementation Guide](/NanmiCoder/MediaCrawler/how-to-add-new-social-media-platform-media-crawler)

Learn how to add a new social media platform to MediaCrawler. This guide covers implementing crawler classes, store modules, API clients, and configuration for seamless integration.

- Tags: how-to-guide
- Published: 2026-07-31

### [MediaCrawler Platform API Signature Handling: XHS and Douyin Implementation Guide](/NanmiCoder/MediaCrawler/media-crawler-platform-api-signature-handling-xhs-sign-douyin-sign)

Master XHS and Douyin API signature handling with MediaCrawler. This guide details pure-Python XHS signing and original JavaScript execution for Douyin's a_bogus parameter.

- Tags: api-reference
- Published: 2026-07-31

### [Adding Custom Words and Chinese Fonts for MediaCrawler Word Clouds: A Complete Guide](/NanmiCoder/MediaCrawler/adding-custom-words-chinese-fonts-media-crawler-word-clouds)

Easily add custom words and Chinese fonts to MediaCrawler word clouds. Configure CUSTOM_WORDS and FONT_PATH in base_config.py, enable word clouds, and run the crawler for stunning visualizations.

- Tags: how-to-guide
- Published: 2026-07-31

### [Enabling Second-Level Comment Crawling for All Platforms in MediaCrawler](/NanmiCoder/MediaCrawler/enabling-second-level-comment-crawling-all-platforms-media-crawler)

Learn how to enable second-level comment crawling in MediaCrawler for all platforms. Use the ENABLE_GET_SUB_COMMENTS flag for hierarchical comment extraction via CLI or config.

- Tags: new-feature-announcement
- Published: 2026-07-31

### [MediaCrawler Login State Caching and Cookie Persistence Management: A Technical Guide](/NanmiCoder/MediaCrawler/media-crawler-login-state-caching-cookie-persistence-management)

Master MediaCrawler login state caching and cookie persistence. Learn how NanmiCoder/MediaCrawler avoids repeated logins by managing user data and Playwright cookies for seamless sessions.

- Tags: how-to-guide
- Published: 2026-07-31

### [Implementing Multi-Account Crawling with IP Rotation in MediaCrawler](/NanmiCoder/MediaCrawler/implementing-multi-account-crawling-ip-rotation-media-crawler)

Learn how to implement multi-account crawling with IP rotation in MediaCrawler. Utilize a shared proxy pool and automatic IP rotation to avoid rate limits and boost your crawling efficiency.

- Tags: how-to-guide
- Published: 2026-07-31

### [Configuring MediaCrawler Data Storage Backends: MySQL, PostgreSQL, SQLite, Excel, and JSONL](/NanmiCoder/MediaCrawler/configuring-media-crawler-data-storage-backends-mysql-postgresql-sqlite-excel-jsonl)

Easily configure MediaCrawler data storage for MySQL, PostgreSQL, SQLite, Excel, and JSONL. Switch backends effortlessly without changing crawler code using config.SAVE_DATA_OPTION.

- Tags: how-to-guide
- Published: 2026-07-31

### [MediaCrawler IP Proxy Configuration: KuaiDaili, WanDouHTTP, and Static Setup](/NanmiCoder/MediaCrawler/media-crawler-ip-proxy-configuration-kuaidaili-wandouhttp-static-setup)

Configure MediaCrawler IP proxies with KuaiDaili, WanDouHTTP, or static setup. Learn dynamic IP rotation and async pool management for efficient crawling.

- Tags: how-to-guide
- Published: 2026-07-31

### [How to Configure CDP Mode in MediaCrawler with Existing Chrome and Login State](/NanmiCoder/MediaCrawler/how-to-configure-cdp-mode-media-crawler-existing-chrome-login-state)

Configure CDP mode in MediaCrawler to attach to an existing Chrome instance preserving login state. Learn how to launch Chrome and set config options for seamless integration.

- Tags: how-to-guide
- Published: 2026-07-31

### [What Kind of Media Can MediaCrawler Crawl? Complete Guide to Supported Platforms and Content Types](/NanmiCoder/MediaCrawler/what-kind-of-media-can-mediacrawler-crawl)

Discover what media MediaCrawler can crawl. Extract text posts, images, videos, and comments from 7 Chinese platforms like Weibo, Douyin, and Bilibili. Get the complete guide to supported content.

- Tags: getting-started
- Published: 2026-07-29

### [How to Set Up MediaCrawler on a Server: Complete Deployment Guide](/NanmiCoder/MediaCrawler/how-to-set-up-mediacrawler-on-server)

Learn how to set up MediaCrawler on your server with this complete deployment guide. Follow simple steps to clone the repo, install dependencies, configure settings, and launch the application for efficient media crawling.

- Tags: how-to-guide
- Published: 2026-07-29

### [MediaCrawler Performance Considerations: Optimizing Throughput and Resource Usage](/NanmiCoder/MediaCrawler/what-are-performance-considerations-for-mediacrawler)

Optimize MediaCrawler performance by balancing concurrency, politeness, browser modes, and I/O for maximum throughput and minimal resource usage. Learn how to enhance your crawling now.

- Tags: performance
- Published: 2026-07-29

### [How to Handle Rate Limiting in MediaCrawler: Configuration and Implementation Guide](/NanmiCoder/MediaCrawler/how-to-handle-rate-limiting-in-mediacrawler)

Learn to handle rate limiting in MediaCrawler. Discover how CRAWLER_MAX_SLEEP_SEC and asyncio.sleep prevent throttling and IP bans for smooth data crawling.

- Tags: how-to-guide
- Published: 2026-07-29

### [How to Use MediaCrawler for Downloading Videos: A Complete Guide](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-for-downloading-videos)

Easily download videos from Douyin, Bilibili, and Kuaishou using MediaCrawler. Our guide shows you how to use its CLI commands or Python async APIs for efficient video downloads.

- Tags: how-to-guide
- Published: 2026-07-29

### [MediaCrawler License: Understanding the Non-Commercial Learning License 1.1](/NanmiCoder/MediaCrawler/what-is-the-license-for-mediacrawler)

Understand MediaCrawler License 1.1. This non-commercial learning license allows you to use, modify, and share MediaCrawler for personal learning and research only. Get the details now.

- Tags: license
- Published: 2026-07-29

### [How to Contribute to the MediaCrawler Project: A Developer's Guide](/NanmiCoder/MediaCrawler/how-to-contribute-to-the-mediacrawler-project)

Learn how to contribute to the MediaCrawler project a modular web-scraping framework built on Playwright. Follow our GitHub pull request guide to get started.

- Tags: how-to-guide
- Published: 2026-07-29

### [How to Troubleshoot MediaCrawler Errors: Complete Diagnostic Guide](/NanmiCoder/MediaCrawler/how-to-troubleshoot-mediacrawler-errors)

Troubleshoot MediaCrawler errors effectively. Diagnose issues from CLI to HTTP handling and apply fixes like cache clearing or proxy overrides for seamless operation.

- Tags: how-to-guide
- Published: 2026-07-29

### [MediaCrawler API: FastAPI HTTP Interface for Orchestrating Crawl Jobs](/NanmiCoder/MediaCrawler/what-is-the-api-for-mediacrawler)

Explore the MediaCrawler API, a FastAPI HTTP interface. Start, stop, and monitor crawl jobs with RESTful endpoints & WebSocket streams. Manage downloaded files efficiently.

- Tags: api-reference
- Published: 2026-07-29

### [How to Integrate MediaCrawler with Other Applications: Python and HTTP API Guide](/NanmiCoder/MediaCrawler/how-to-integrate-mediacrawler-with-other-applications)

Learn to integrate MediaCrawler with other applications using Python imports or its HTTP API. Control crawlers remotely and streamline your media management with this comprehensive guide.

- Tags: how-to-guide
- Published: 2026-07-29

### [How to Customize MediaCrawler's Scraping Logic: A Developer's Guide to Extending the Crawler](/NanmiCoder/MediaCrawler/how-to-customize-mediacrawlers-scraping-logic)

Customize MediaCrawler scraping logic by subclassing AbstractCrawler overriding platform-specific methods and registering your custom class via CrawlerFactory. Extend MediaCrawler easily.

- Tags: how-to-guide
- Published: 2026-07-29

### [Supported Media Source Types in MediaCrawler: The Complete 7-Platform Guide](/NanmiCoder/MediaCrawler/what-are-supported-media-source-types-in-mediacrawler)

Discover the 7 supported media source types in MediaCrawler including XiaoHongShu, Douyin, and Bilibili. Access diverse Chinese social media content efficiently with our guide.

- Tags: getting-started
- Published: 2026-07-29

### [How to Add a New Media Source to MediaCrawler: Complete Implementation Guide](/NanmiCoder/MediaCrawler/how-to-add-new-media-source-to-mediacrawler)

Learn how to add a new media source to MediaCrawler with this complete implementation guide. Follow our step-by-step instructions to integrate custom crawlers.

- Tags: how-to-guide
- Published: 2026-07-29

### [Where Are the Configuration Files Located in MediaCrawler?](/NanmiCoder/MediaCrawler/where-are-the-configuration-files-located-for-mediacrawler)

Find MediaCrawler configuration files in the main/config/ directory. Discover platform-specific modules and the shared base_config.py for all your settings.

- Tags: how-to-guide
- Published: 2026-07-29

### [How to Configure MediaCrawler: A Complete Guide to Base, Platform-Specific, and Database Settings](/NanmiCoder/MediaCrawler/how-to-configure-mediacrawler)

Configure MediaCrawler effectively with this comprehensive guide. Master base, platform, and database settings without touching source code. Customize your media management now.

- Tags: how-to-guide
- Published: 2026-07-29

### [What Programming Language Is MediaCrawler Written In?](/NanmiCoder/MediaCrawler/what-programming-language-is-mediacrawler-written-in)

Discover what programming language MediaCrawler uses. This powerful tool is built with Python 3.11+ and modern async libraries for efficient data extraction.

- Tags: getting-started
- Published: 2026-07-29

### [How to Run MediaCrawler from Source: Complete Setup Guide](/NanmiCoder/MediaCrawler/how-to-run-mediacrawler-from-source)

Learn how to run MediaCrawler from source with this setup guide. Clone the repo, install dependencies, and start crawling immediately by following simple steps.

- Tags: how-to-guide
- Published: 2026-07-29

### [MediaCrawler Ethical Guidelines for Xiaohongshu (XHS): Responsible Web Scraping Practices](/NanmiCoder/MediaCrawler/ethical-guidelines-using-media-crawler-xhs)

Discover MediaCrawler ethical guidelines for Xiaohongshu scraping. Learn responsible practices and ensure compliance with XHS Terms of Service. Download MediaCrawler for ethical data collection.

- Tags: best-practices
- Published: 2026-07-03

### [What Research and Tools Power the XHS Scraping Logic in MediaCrawler](/NanmiCoder/MediaCrawler/research-tools-used-build-xhs-scraping-logic)

Learn how MediaCrawler's XHS scraping logic uses xhshow, Playwright, and custom tools to reverse-engineer XiaoHongShu's private API for efficient data extraction.

- Tags: how-to-guide
- Published: 2026-07-03

### [How to Contribute New Platform Crawlers to MediaCrawler: A Step-by-Step Guide](/NanmiCoder/MediaCrawler/contribute-development-new-platform-crawlers)

Learn how to contribute new platform crawlers to the MediaCrawler project. This step-by-step guide covers implementing the AbstractCrawler interface and registering your crawler.

- Tags: how-to-guide
- Published: 2026-07-03

### [XHS Crawler Performance Characteristics: How MediaCrawler Handles Async Throughput and Resource Management](/NanmiCoder/MediaCrawler/performance-characteristics-xhs-crawler)

Discover XHS crawler performance: MediaCrawler achieves 10-15 items/sec with async I/O, ideal for steady, long-running crawls. Learn about its resource management & throughput.

- Tags: performance
- Published: 2026-07-03

### [How NanmiCoder/MediaCrawler Handles Proxy Rotation for XHS Scraping](/NanmiCoder/MediaCrawler/handle-proxy-rotation-scraping-xhs-nanmicoder-mediacrawler)

Learn how NanmiCoder/MediaCrawler handles XHS scraping with automatic proxy rotation. Discover its three-layer architecture: ProxyIpPool, ProxyRefreshMixin, and XiaoHongShuClient.

- Tags: how-to-guide
- Published: 2026-07-03

### [Xiaohongshu (XHS) Crawler Limitations and Known Issues in MediaCrawler](/NanmiCoder/MediaCrawler/limitations-known-issues-xiaohongshu-xhs-crawler)

Explore Xiaohongshu crawler limitations in MediaCrawler including authentication, pagination, and rate-limiting issues. Learn how these restrict large-scale XHS data scraping.

- Tags: known-issues
- Published: 2026-07-03

### [How MediaCrawler Manages Configuration for Different Platform Scrapers](/NanmiCoder/MediaCrawler/configuration-management-platform-scrapers-nanmicoder-mediacrawler)

Discover how MediaCrawler effectively manages configuration for diverse platform scrapers using a centralized config package and platform-specific modules. Optimize your scraping setups.

- Tags: configuration-management
- Published: 2026-07-03

### [Anti-Scraping Measures in the MediaCrawler XHS Module: A Technical Implementation Guide](/NanmiCoder/MediaCrawler/anti-scraping-measures-xhs-module-developers-aware)

Uncover MediaCrawler XHS module's anti-scraping measures: cryptographic signing, CAPTCHA detection, IP monitoring, retry logic, delays, proxy rotation, and data anonymization. Enhance your scraping resilience.

- Tags: how-to-guide
- Published: 2026-07-03

### [Logging Mechanisms for Debugging Platform-Specific Issues in MediaCrawler](/NanmiCoder/MediaCrawler/logging-mechanisms-debugging-platform-specific-issues)

Debug platform-specific issues in MediaCrawler with its centralized Python logging. Trace execution paths across browser automation, data storage, and async processing for efficient troubleshooting.

- Tags: how-to-guide
- Published: 2026-07-03

### [How NanmiCoder/MediaCrawler Ensures the Integrity of Scraped Data from XHS](/NanmiCoder/MediaCrawler/ensure-integrity-scraped-data-xhs-nanmicoder-mediacrawler)

Learn how NanmiCoder MediaCrawler guarantees XHS scraped data integrity with hashing, validated persistence, idempotent up-serts, and thread-safe writes. Prevent duplicates effortlessly.

- Tags: how-to-guide
- Published: 2026-07-03

### [Maintaining the XHS Crawler as the Platform Evolves: A Technical Guide to Long-Term Stability](/NanmiCoder/MediaCrawler/considerations-maintaining-xhs-crawler-platform-evolves)

Keep your XHS crawler stable and up-to-date. Learn to monitor API changes, manage tokens, adjust rate limits, and sync data models as Xiaohongshu evolves.

- Tags: how-to-guide
- Published: 2026-07-03

### [How to Add Support for a New Platform to NanmiCoder/MediaCrawler: A Step-by-Step Guide](/NanmiCoder/MediaCrawler/add-support-new-platform-nanmicoder-mediacrawler)

Easily add a new platform to MediaCrawler by implementing AbstractCrawler, creating client wrappers and store modules, and registering in CrawlerFactory. Follow our step-by-step guide.

- Tags: how-to-guide
- Published: 2026-07-03

### [How the XHS Crawler Handles Dynamic Content and JavaScript-Rendered Elements](/NanmiCoder/MediaCrawler/dynamic-content-javascript-rendered-elements-xhs-crawler)

Discover how the XHS crawler tackles dynamic content and JavaScript elements using Playwright headless browser and an HTML fallback parser for robust data extraction.

- Tags: how-to-guide
- Published: 2026-07-03

### [XHS Data Models in MediaCrawler: How Scraped Xiaohongshu Content Is Structured](/NanmiCoder/MediaCrawler/data-models-scraped-content-xhs-nanmicoder-mediacrawler)

Explore XHS data models in MediaCrawler. Learn how Pydantic models structure scraped Xiaohongshu content for efficient parsing and storage with the XhsNote ORM.

- Tags: architecture
- Published: 2026-07-03

### [Error Handling Strategies for Individual Platform Scrapers in MediaCrawler](/NanmiCoder/MediaCrawler/error-handling-strategies-platform-scrapers-nanmicoder-mediacrawler)

Discover MediaCrawler's robust error handling strategies for platform scrapers. Learn how custom exceptions, automatic retries, and graceful degradation ensure reliable data extraction.

- Tags: best-practices
- Published: 2026-07-03

### [How NanmiCoder/MediaCrawler Handles XiaoHongShu (XHS) Login Authentication](/NanmiCoder/MediaCrawler/handle-authentication-login-platforms-xhs-nanmicoder-mediacrawler)

Learn how MediaCrawler handles XiaoHongShu login authentication using QR code scanning, SMS verification, or cookie injection. Explore its secure signature generation.

- Tags: how-to-guide
- Published: 2026-07-03

### [Main Dependencies for the Xiaohongshu (XHS) Crawler Module in MediaCrawler](/NanmiCoder/MediaCrawler/dependencies-xiaohongshu-xhs-crawler-module)

Discover the essential dependencies for the Xiaohongshu crawler in MediaCrawler. Learn how httpx, playwright, tenacity, and xhshow power its functionality.

- Tags: deep-dive
- Published: 2026-07-03

### [How MediaCrawler Manages API Versions and Changes for Xiaohongshu (XHS)](/NanmiCoder/MediaCrawler/manage-api-versions-changes-platforms-xhs-nanmicoder-mediacrawler)

Learn how MediaCrawler handles XHS API versions and changes using package isolation, constants, and factory patterns. Keep your crawler logic stable and updated.

- Tags: best-practices
- Published: 2026-07-03

### [Common Interfaces and Base Classes for Platform Crawlers in MediaCrawler](/NanmiCoder/MediaCrawler/common-interfaces-base-classes-platform-crawlers-nanmicoder-mediacrawler)

Discover common interfaces and base classes for platform crawlers in MediaCrawler. Standardize your crawling, authentication, and API interactions with our comprehensive hierarchy.

- Tags: internals
- Published: 2026-07-03

### [Directory Structure for Platform-Specific Code in NanmiCoder/MediaCrawler](/NanmiCoder/MediaCrawler/directory-structure-platform-specific-code-nanmicoder-mediacrawler)

Explore the NanmiCoder/MediaCrawler directory structure for platform-specific code. Learn how to isolate each social media platform with consistent internal files and mirrored storage logic.

- Tags: architecture
- Published: 2026-07-03

### [Where to Find the Xiaohongshu (XHS) Scraping Implementation in MediaCrawler](/NanmiCoder/MediaCrawler/xiaohongshu-xhs-scraping-implementation-details-nanmicoder-mediacrawler)

Find the Xiaohongshu scraping implementation in NanmiCoder/MediaCrawler within the media_platform/xhs/ and store/xhs/ directories. Access entry points via store/xhs/__init__.py.

- Tags: how-to-guide
- Published: 2026-07-03

### [Understanding Platform-Specific Features in NanmiCoder/MediaCrawler: Key Files and Architecture](/NanmiCoder/MediaCrawler/key-files-understand-platform-specific-features-nanmicoder-mediacrawler)

Explore platform-specific features in NanmiCoder/MediaCrawler. Discover key files like corepy, clientpy, and loginpy within isolated packages to understand architecture and customization.

- Tags: architecture
- Published: 2026-07-03

### [How NanmiCoder/MediaCrawler Handles Xiaohongshu (XHS) Platform-Specific Implementation](/NanmiCoder/MediaCrawler/how-nanmicoder-mediacrawler-handle-platform-specific-implementations-xiaohongshu)

Discover how NanmiCoder/MediaCrawler manages Xiaohongshu XHS platform specifics. Explore its modular design featuring typed data models authentication, crypto signing, and three crawl modes.

- Tags: deep-dive
- Published: 2026-07-03

### [How Platform-Specific field.py Modules Define Data Schemas in MediaCrawler](/NanmiCoder/MediaCrawler/how-do-platform-specific-field-py-modules-define-data-schemas-in-media-crawler)

Discover how MediaCrawler uses platform-specific field.py modules to define precise data schemas for request parameters and API responses, ensuring efficient data handling.

- Tags: internals
- Published: 2026-07-02

### [How MediaCrawler Handles Async Concurrency for Parallel Crawling](/NanmiCoder/MediaCrawler/how-does-media-crawler-handle-async-concurrency-for-parallel-crawling)

Discover how MediaCrawler manages async concurrency for parallel crawling using asyncio Semaphore and gather. Optimize your crawls and avoid rate limiting.

- Tags: internals
- Published: 2026-07-02

### [ENABLE_CDP_MODE vs Playwright Headless Mode in MediaCrawler: Key Differences Explained](/NanmiCoder/MediaCrawler/what-is-the-difference-between-enable-cdp-mode-and-playwright-headless-mode-in-media-crawler)

Understand ENABLE_CDP_MODE vs Playwright headless mode in MediaCrawler. Learn how MediaCrawler connects to Chrome or uses Playwright's fresh instances for optimal control.

- Tags: deep-dive
- Published: 2026-07-02

### [How MediaCrawler Extracts Cookies from Browser Context Using Playwright](/NanmiCoder/MediaCrawler/how-does-media-crawler-extract-cookies-from-browser-context)

Learn how MediaCrawler extracts browser cookies using Playwright. Discover the efficient methods employed for obtaining and converting cookie data for your needs. Get the details now.

- Tags: how-to-guide
- Published: 2026-07-02

### [What Is the Role of tools/user_hash.py in MediaCrawler?](/NanmiCoder/MediaCrawler/what-is-the-role-of-tools-user-hash-py-in-media-crawler)

Discover how tools/user_hash.py in MediaCrawler anonymizes user IDs and masks nicknames to safeguard personal data while maintaining analytical insights.

- Tags: internals
- Published: 2026-07-02

### [How to Add a New Media Platform to MediaCrawler's Crawler Factory](/NanmiCoder/MediaCrawler/how-to-add-a-new-media-platform-to-media-crawlers-crawler-factory)

Easily add new media platforms to MediaCrawler. Learn how to create a new crawler class, register it, and integrate it seamlessly into the Crawler Factory with this guide.

- Tags: how-to-guide
- Published: 2026-07-02

### [How MediaCrawler Uses Redis Cache to Manage Session State and Deduplication](/NanmiCoder/MediaCrawler/how-does-the-redis-cache-manage-session-state-and-deduplication-in-media-crawler)

Discover how MediaCrawler leverages Redis cache for efficient session state management and request deduplication. Learn about Redis sets and TTL expiration for optimized performance.

- Tags: performance
- Published: 2026-07-02

### [How ExcelStoreBase Flushes Data to Excel Files in MediaCrawler](/NanmiCoder/MediaCrawler/how-does-excelstorebase-flush-data-to-excel-files-in-media-crawler)

Discover how MediaCrawler's ExcelStoreBase flushes data to Excel. Learn about its singleton pattern, automatic column adjustment, and efficient saving for clean, organized spreadsheets.

- Tags: internals
- Published: 2026-07-02

### [How MediaCrawler Handles Login Failures and Switches Login Methods Automatically](/NanmiCoder/MediaCrawler/how-does-media-crawler-handle-login-failures-and-switch-login-methods)

Learn how MediaCrawler handles login failures by automatically switching authentication methods from QR code to mobile to cookie for a stable session.

- Tags: how-to-guide
- Published: 2026-07-02

### [What Crawler Types Does MediaCrawler's CLI Support?](/NanmiCoder/MediaCrawler/what-are-the-different-crawler-types-in-media-crawlers-cli)

Discover the MediaCrawler CLI's supported crawler types. Learn how to perform keyword-based content discovery across social media platforms with the search crawler.

- Tags: how-to-guide
- Published: 2026-07-02

### [How MediaCrawler Handles Comment Pagination with Nested Comments in Tieba](/NanmiCoder/MediaCrawler/how-does-media-crawler-handle-comment-pagination-with-nested-comments)

Discover how MediaCrawler tackles comment pagination and nested comments in Tieba. Learn about its two-stage system using total replay page and sub comment count for efficient data retrieval.

- Tags: how-to-guide
- Published: 2026-07-02

### [What Is cdp_browser.py in MediaCrawler? CDP Browser Management Explained](/NanmiCoder/MediaCrawler/what-is-the-purpose-of-cdp-browser-py-in-media-crawler)

Discover the CDP Browser Manager in MediaCrawler's cdp_browser.py. Learn how it launches browsers, manages CDP connections, and handles Playwright instances efficiently.

- Tags: internals
- Published: 2026-07-02

### [How the FastAPI WebUI Integrates with the MediaCrawler Core Engine](/NanmiCoder/MediaCrawler/how-does-the-fastapi-webui-in-media-crawler-integrate-with-the-crawler)

Explore how the FastAPI WebUI integrates with the MediaCrawler core engine. This guide details its role in managing the crawler's lifecycle, logs, and status through REST and WebSockets.

- Tags: architecture
- Published: 2026-07-02

### [How MediaCrawler Implements Rate Limiting: Pagination, Sleep Intervals, and Jitter](/NanmiCoder/MediaCrawler/how-does-media-crawler-implement-rate-limiting)

Discover how MediaCrawler implements effective rate limiting using pagination, sleep intervals, and jitter. Optimize your scraping with these advanced techniques.

- Tags: deep-dive
- Published: 2026-07-02

### [How MediaCrawler Generates Wordclouds from Scraped Comments: A Technical Deep Dive](/NanmiCoder/MediaCrawler/how-does-media-crawler-generate-wordclouds-from-scraped-comments)

Discover how MediaCrawler crafts wordclouds from scraped comments. Learn about its async pipeline, token frequency calculation, and PNG generation with the Python wordcloud library.

- Tags: deep-dive
- Published: 2026-07-02

### [How MediaCrawler Handles Sliding Captcha Verification: OpenCV and Track Simulation](/NanmiCoder/MediaCrawler/how-does-media-crawler-handle-sliding-captcha-verification)

MediaCrawler defeats sliding captcha with OpenCV template matching and track simulation. Learn how it automates slider verification for seamless web scraping.

- Tags: how-to-guide
- Published: 2026-07-02

### [MediaCrawler Data Storage Backends: JSONL, SQLite, MySQL, and PostgreSQL Configuration Guide](/NanmiCoder/MediaCrawler/what-are-the-data-storage-backends-in-media-crawler-jsonl-sqlite-mysql-postgres)

Configure MediaCrawler's data storage backends JSONL, SQLite, MySQL, and PostgreSQL. Learn how to set up your preferred database for efficient data management.

- Tags: configuration-guide
- Published: 2026-07-02

### [How MediaCrawler's Proxy IP Pool Works with Multiple Providers](/NanmiCoder/MediaCrawler/how-does-media-crawlers-proxy-ip-pool-system-work-with-different-providers)

Discover how MediaCrawler's proxy IP pool seamlessly integrates with multiple providers, automatically managing and refreshing IPs for efficient web crawling.

- Tags: internals
- Published: 2026-07-02

### [How MediaCrawler Uses CDP Mode to Bypass Bot Detection](/NanmiCoder/MediaCrawler/how-does-media-crawler-use-cdp-mode-to-bypass-bot-detection)

Learn how MediaCrawler bypasses bot detection using CDP mode. This advanced technique leverages real browser instances, profiles, extensions, and cookies for stealthy automation.

- Tags: how-to-guide
- Published: 2026-07-02

### [MediaCrawler Command-Line Arguments: Complete CLI Reference](/NanmiCoder/MediaCrawler/what-are-the-command-line-arguments-for-running-mediacrawler)

Explore MediaCrawler command-line arguments for platform selection, login, crawling scope, storage, and proxy settings. Get the complete CLI reference for efficient media crawling.

- Tags: api-reference
- Published: 2026-07-01

### [How to Use MediaCrawler for Scraping Data from Bilibili: A Complete Technical Guide](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-for-scraping-data-from-bilibili)

Learn how to scrape Bilibili data with NanmiCoder MediaCrawler. Extract videos, comments, and profiles using this async, Playwright-based framework. Supports CSV, MongoDB, SQLite, and Excel.

- Tags: how-to-guide
- Published: 2026-07-01

### [How to Generate Word Clouds from Scraped Data Using MediaCrawler](/NanmiCoder/MediaCrawler/how-can-i-generate-word-clouds-from-scraped-data-using-mediacrawler)

Easily generate word clouds from scraped data with MediaCrawler. Follow simple steps to enable comment and word cloud generation for instant data visualization.

- Tags: how-to-guide
- Published: 2026-07-01

### [What is stealth.js and How MediaCrawler Uses It for Anti-Detection Scraping](/NanmiCoder/MediaCrawler/what-is-stealthjs-and-how-is-it-used-in-mediacrawler)

Discover stealth.js, a JavaScript tool that bypasses bot detection. Learn how MediaCrawler employs stealth.js for effective anti-detection web scraping of Chinese platforms.

- Tags: deep-dive
- Published: 2026-07-01

### [How to Use Cookie-Based Login in MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/how-to-use-cookie-based-login-in-mediacrawler)

Master cookie-based login in MediaCrawler with this guide. Learn to set LOGIN_TYPE and provide cookies via CLI or config for seamless authentication bypass. Get started today.

- Tags: how-to-guide
- Published: 2026-07-01

### [How to Customize MediaCrawler's Settings Using the Configuration System](/NanmiCoder/MediaCrawler/how-can-i-customize-mediacrawlers-settings-using-the-configuration-system)

Customize MediaCrawler settings easily. Learn to modify runtime configurations via file edits, command-line options, or programmatic changes using its Python module-level variables.

- Tags: how-to-guide
- Published: 2026-07-01

### [How to Use Playwright CDP Mode with MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/how-can-i-use-playwrights-cdp-mode-with-mediacrawler)

Integrate Playwright CDP mode with MediaCrawler for enhanced network control. This guide shows how to leverage CdpBrowser to send raw CDP commands while using Playwright APIs.

- Tags: how-to-guide
- Published: 2026-07-01

### [Which Platforms Does MediaCrawler Support for Web Scraping?](/NanmiCoder/MediaCrawler/which-platforms-does-mediacrawler-support-for-web-scraping)

Scrape data from 7 major Chinese platforms with MediaCrawler. Supports Zhihu, RED, Douyin, Bilibili, Kuaishou, Weibo, and Baidu Tieba easily and efficiently.

- Tags: getting-started
- Published: 2026-07-01

### [How to Install MediaCrawler and Its Dependencies: Complete Setup Guide](/NanmiCoder/MediaCrawler/how-to-install-mediacrawler-and-its-dependencies)

Easily install MediaCrawler and its dependencies. Follow our complete setup guide to clone the repo, sync Python dependencies, and install Node.js prerequisites for seamless operation.

- Tags: getting-started
- Published: 2026-07-01

### [How Redis Caching Improves MediaCrawler Performance: Architecture and Implementation](/NanmiCoder/MediaCrawler/how-does-redis-caching-improve-mediacrawler-performance)

Discover how Redis caching boosts MediaCrawler performance by reducing network requests and latency. Learn about its architecture and implementation with configurable TTLs for efficient data storage.

- Tags: performance
- Published: 2026-06-30

### [How to Integrate Custom Proxy Providers with MediaCrawler: Static, Kuaidaili, and Wandouhttp Configuration](/NanmiCoder/MediaCrawler/how-to-integrate-custom-proxy-providers-with-mediacrawler)

Easily integrate custom proxy providers like static, Kuaidaili, and Wandouhttp with MediaCrawler for efficient request routing. Configure your proxy settings in base_config.py.

- Tags: how-to-guide
- Published: 2026-06-30

### [MediaCrawler Login Types Explained: QR-Code vs Phone vs Cookie Authentication](/NanmiCoder/MediaCrawler/what-are-the-differences-between-qrcode-phone-and-cookie-login-types-in-mediacrawler)

Discover QR-code, phone, and cookie authentication in MediaCrawler. Learn how to choose the right login type for your needs and integrate seamlessly with NanmiCoder MediaCrawler.

- Tags: deep-dive
- Published: 2026-06-30

### [How to Configure an IP Proxy Pool with the KuaiDaili Provider in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-ip-proxy-pool-with-kuaidaili-provider)

Learn how to configure an IP proxy pool with the KuaiDaili provider in MediaCrawler. Easily set up credentials and environment variables for efficient proxy usage.

- Tags: how-to-guide
- Published: 2026-06-30

### [How MediaCrawler's CDP Mode Works for Anti-Detection: A Technical Deep Dive](/NanmiCoder/MediaCrawler/how-does-mediacrawler-s-cdp-mode-work-for-anti-detection)

Discover how MediaCrawler's CDP mode bypasses anti-detection. Connect to real Chrome or Edge browsers, inherit user fingerprints & sessions, and mask automation for effective scraping.

- Tags: deep-dive
- Published: 2026-06-30

### [How to Debug Common MediaCrawler Issues: Browser Crashes, Login Failures, and Data Gaps](/NanmiCoder/MediaCrawler/mediacrawler-debug-common-issues)

Troubleshoot MediaCrawler problems like browser crashes, login errors, and missing data. Learn to debug CDP connections, QR code selectors, and async store writes effectively.

- Tags: how-to-guide
- Published: 2026-06-29

### [When to Use login_by_cookies Instead of QR Code Login in MediaCrawler](/NanmiCoder/MediaCrawler/mediacrawler-cookie-login-vs-qr-code-login)

Discover when to use login_by_cookies in MediaCrawler for headless automation, CI/CD deployments, and reusing sessions. Automate logins efficiently.

- Tags: best-practices
- Published: 2026-06-29

### [Headless vs Non-Headless Mode Trade-offs in MediaCrawler: Platform-Specific Guide](/NanmiCoder/MediaCrawler/mediacrawler-headless-non-headless-mode-tradeoffs)

Explore headless vs non-headless mode trade-offs for MediaCrawler. Understand resource use, anti-bot risks, and interactive logins for platform-specific deployment and verification.

- Tags: deep-dive
- Published: 2026-06-29

### [How to Configure Rate Limiting and Sleep Intervals Using CRAWLER_MAX_SLEEP_SEC in MediaCrawler](/NanmiCoder/MediaCrawler/mediacrawler-rate-limiting-sleep-intervals-crawler-max-sleep-sec)

Learn to configure rate limiting and sleep intervals in MediaCrawler by adjusting CRAWLER_MAX_SLEEP_SEC. Control pause duration between HTTP requests for better crawling. Find out how in this guide.

- Tags: how-to-guide
- Published: 2026-06-29

### [How to Configure a Custom Browser Path for CDP Mode in MediaCrawler](/NanmiCoder/MediaCrawler/mediacrawler-cdp-mode-custom-browser-path)

Easily configure MediaCrawler's CDP Mode with a custom browser path. Learn how to set CUSTOM_BROWSER_PATH in config base_config.py for Chrome or Edge.

- Tags: how-to-guide
- Published: 2026-06-29

### [How MediaCrawler Implements Second-Level Comment Crawling with ENABLE_GET_SUB_COMMENTS](/NanmiCoder/MediaCrawler/mediacrawler-second-level-comment-crawling-enable-get-sub-comments)

Learn how MediaCrawler enables second-level comment crawling by setting ENABLE_GET_SUB_COMMENTS to True. This activates the get_comments_all_sub_comments method for nested replies on Zhihu, Douyin, and Bilibili.

- Tags: deep-dive
- Published: 2026-06-29

### [How to Manage Browser Contexts in MediaCrawler to Prevent Memory Leaks](/NanmiCoder/MediaCrawler/mediacrawler-browser-context-management-memory-leaks)

Learn how MediaCrawler prevents memory leaks by managing browser contexts with CDPBrowserManager. Ensure proper context closure to free memory and optimize performance.

- Tags: performance
- Published: 2026-06-29

### [How to Generate a Word Cloud from Comments in MediaCrawler with Custom and Stop Words](/NanmiCoder/MediaCrawler/mediacrawler-word-cloud-comments-custom-stop-words)

Learn how to generate a word cloud from MediaCrawler comments. Easily customize word clouds using custom and stop words with AsyncWordCloudGenerator.

- Tags: how-to-guide
- Published: 2026-06-29

### [How to Control Concurrency in MediaCrawler Using MAX_CONCURRENCY_NUM](/NanmiCoder/MediaCrawler/mediacrawler-concurrency-control-max-concurrency-num)

Optimize MediaCrawler throughput by controlling concurrency. Learn to set MAX_CONCURRENCY_NUM via config or CLI to manage parallel requests and avoid rate limits effectively.

- Tags: best-practices
- Published: 2026-06-29

### [How MediaCrawler Handles Slider Verification in Automated Crawling](/NanmiCoder/MediaCrawler/mediacrawler-handle-slider-verification)

Learn how MediaCrawler effectively handles slider verification in automated crawling. Discover its image processing and human-like mouse trajectory generation techniques.

- Tags: how-to-guide
- Published: 2026-06-29

### [When to Use Playwright versus CDP Mode in MediaCrawler](/NanmiCoder/MediaCrawler/mediacrawler-playwright-vs-cdp-mode)

Decide between Playwright and CDP Mode in MediaCrawler. Use Playwright for isolated, stateless crawls with proxies, or CDP Mode to reuse Chrome profiles and bypass bot detection.

- Tags: deep-dive
- Published: 2026-06-29

### [How to Use CDP Mode in MediaCrawler to Connect to an Existing Chrome Browser](/NanmiCoder/MediaCrawler/mediacrawler-cdp-mode-existing-chrome)

Learn how to use MediaCrawler's CDP Mode to connect to an existing Chrome browser. Preserve sessions and skip launch time with the Chrome DevTools Protocol.

- Tags: how-to-guide
- Published: 2026-06-29

### [Where Are MediaCrawler Logs Stored? Configuration and File Paths Explained](/NanmiCoder/MediaCrawler/where-are-logs-stored-for-mediacrawler)

Discover where MediaCrawler logs are stored. Learn how to configure file paths for log persistence and understand default stderr behavior for the NanmiCoder/MediaCrawler repository.

- Tags: how-to-guide
- Published: 2026-06-28

### [How MediaCrawler Handles Images, Videos, and Text: A Complete Technical Guide](/NanmiCoder/MediaCrawler/how-does-mediacrawler-handle-different-media-types)

Learn how MediaCrawler handles images, videos, and text with platform-specific dispatch methods and specialized storage helpers. A complete technical guide to the NanmiCoder/MediaCrawler repo.

- Tags: deep-dive
- Published: 2026-06-28

### [Limitations of MediaCrawler: What to Know Before Scraping Chinese Social Media](/NanmiCoder/MediaCrawler/what-are-the-limitations-of-mediacrawler)

Explore MediaCrawler limitations: legal restrictions, browser automation needs, QR code auth, and limited comment depth. Learn what to know before scraping Chinese social media.

- Tags: deep-dive
- Published: 2026-06-28

### [How to Parse the Data Extracted by MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/how-to-parse-data-extracted-by-mediacrawler)

Master parsing MediaCrawler data. Learn how this Python tool transforms raw social media content into clean, typed objects with its efficient three-layer pipeline.

- Tags: how-to-guide
- Published: 2026-06-28

### [Is MediaCrawler Actively Maintained? Current Development Status and Architecture Deep Dive](/NanmiCoder/MediaCrawler/is-mediacrawler-actively-maintained)

Discover if MediaCrawler is actively maintained. Explore recent development, architecture, and documentation updates for this actively maintained project. Learn more now.

- Tags: deep-dive
- Published: 2026-06-28

### [How to Set Up Proxy Servers for MediaCrawler: A Complete Configuration Guide](/NanmiCoder/MediaCrawler/how-to-set-up-proxy-servers-for-mediacrawler)

Learn how to set up proxy servers for MediaCrawler with this complete configuration guide. Easily integrate KuaiDaili, Wandou HTTP, or static proxies for enhanced performance.

- Tags: how-to-guide
- Published: 2026-06-28

### [What Kind of Data Can Be Extracted by MediaCrawler: Content, Comments, and Creator Profiles](/NanmiCoder/MediaCrawler/what-kind-of-data-can-be-extracted-by-mediacrawler)

Discover the data MediaCrawler extracts: content, comments, and creator profiles from Chinese social media. Normalize and store data in CSV, JSON, SQL, or NoSQL.

- Tags: getting-started
- Published: 2026-06-28

### [How to Filter Crawled Content by Date or Keywords in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-filter-crawled-content-by-date-keywords)

Learn how to filter crawled content by date or keywords in MediaCrawler. Easily modify configuration files to refine your search and get relevant results faster.

- Tags: how-to-guide
- Published: 2026-06-28

### [How MediaCrawler Bypasses IP Bans: Proxy Pool Architecture Explained](/NanmiCoder/MediaCrawler/can-mediacrawler-bypass-ip-bans)

MediaCrawler bypasses IP bans using its advanced proxy pool architecture. Learn how it automatically refreshes IPs for seamless scraping and avoid getting blocked.

- Tags: architecture
- Published: 2026-06-28

### [How to Use MediaCrawler with Docker: Complete Containerization Guide](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-with-docker)

Containerize MediaCrawler with Docker. This guide shows how to build a custom Dockerfile, install dependencies, and run the crawler for efficient media management.

- Tags: how-to-guide
- Published: 2026-06-28

### [Ethical Considerations When Using MediaCrawler: Compliance and Best Practices](/NanmiCoder/MediaCrawler/what-are-ethical-considerations-when-using-mediacrawler)

Learn ethical considerations for MediaCrawler. Ensure compliance with laws, respect Terms of Service, protect data, and operate responsibly for educational and research use.

- Tags: best-practices
- Published: 2026-06-28

### [How to Contribute to the MediaCrawler Open-Source Project](/NanmiCoder/MediaCrawler/how-to-contribute-to-mediacrawler-project)

Learn how to contribute to the MediaCrawler open-source project. Fork the repo, set up your environment, and follow the modular architecture to add new features or fix bugs.

- Tags: getting-started
- Published: 2026-06-28

### [Troubleshooting Common MediaCrawler Errors: A Complete Guide](/NanmiCoder/MediaCrawler/troubleshooting-common-mediacrawler-errors)

Troubleshoot MediaCrawler errors by fixing Playwright injection, validating login cookies with pong(), and ensuring proxy rotation via ProxyIpPool checks. Resolve runtime failures now.

- Tags: how-to-guide
- Published: 2026-06-28

### [How to Update MediaCrawler to the Latest Version: A Complete Upgrade Guide](/NanmiCoder/MediaCrawler/how-to-update-mediacrawler-to-latest-version)

Easily update MediaCrawler to the latest version. Learn to pull commits, sync dependencies, and reinstall browsers to keep your MediaCrawler current and functional.

- Tags: how-to-guide
- Published: 2026-06-28

### [What Are the Dependencies for MediaCrawler? Complete Requirements Guide](/NanmiCoder/MediaCrawler/what-are-the-dependencies-for-mediacrawler)

MediaCrawler needs Python 3.x and over 30 runtime packages like httpx playwright and sqlalchemy. Learn about all MediaCrawler dependencies to get started.

- Tags: how-to-guide
- Published: 2026-06-28

### [How to Schedule MediaCrawler Tasks](/NanmiCoder/MediaCrawler/how-to-schedule-mediacrawler-tasks)

Schedule MediaCrawler tasks effortlessly using its built-in asyncio scheduler. Run recurring background crawling jobs at fixed intervals without external dependencies.

- Tags: how-to-guide
- Published: 2026-06-28

### [How to Customize MediaCrawler's Behavior: A Complete Configuration Guide](/NanmiCoder/MediaCrawler/how-to-customize-crawler-behavior)

Customize MediaCrawler behavior easily. Learn how to configure this powerful crawling framework using config values, command-line arguments, or environment variables without code changes.

- Tags: how-to-guide
- Published: 2026-06-28

### [MediaCrawler Output Format: Configuration Guide for JSONL, CSV, Excel, and Databases](/NanmiCoder/MediaCrawler/what-is-the-output-format-of-mediacrawler)

Explore MediaCrawler output formats. Learn to configure JSONL, CSV, Excel, and databases like SQL & MongoDB for your data needs. Get your data your way.

- Tags: configuration-guide
- Published: 2026-06-28

### [How to Handle Rate Limits with MediaCrawler: Configuration and Best Practices](/NanmiCoder/MediaCrawler/how-to-handle-rate-limits-with-mediacrawler)

Learn to handle rate limits with MediaCrawler. Discover configurations for fixed spacing, automatic retries, and IP proxy rotation to optimize your web scraping.

- Tags: best-practices
- Published: 2026-06-28

### [Can MediaCrawler Crawl Private Social Media Accounts? A Technical Deep Dive](/NanmiCoder/MediaCrawler/can-mediacrawler-crawl-private-social-media-accounts)

Discover if MediaCrawler can access private social media accounts. Learn how authentication impacts access and platform privacy limitations. Get the technical details now.

- Tags: deep-dive
- Published: 2026-06-28

### [How to Configure MediaCrawler for a Specific Platform: Complete Setup Guide](/NanmiCoder/MediaCrawler/how-to-configure-mediacrawler-for-specific-platform)

Configure MediaCrawler for any platform by creating a custom Python config file. Follow this setup guide to launch the crawler with targeted settings for your needs.

- Tags: how-to-guide
- Published: 2026-06-28

### [What Platforms Does MediaCrawler Support? A Complete Guide to the 7 Chinese Social Media Platforms](/NanmiCoder/MediaCrawler/what-platforms-does-mediacrawler-support)

Explore the 7 Chinese social media platforms MediaCrawler supports including XiaoHongShu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. Get the complete list!

- Tags: getting-started
- Published: 2026-06-28

### [How to Install MediaCrawler: Complete Setup Guide for the NanmiCoder Repository](/NanmiCoder/MediaCrawler/how-to-install-mediacrawler)

Learn how to install MediaCrawler with our complete NanmiCoder repository setup guide. Follow these easy steps to get your media management up and running quickly.

- Tags: how-to-guide
- Published: 2026-06-28

### [MediaCrawler Headless Browser Options: Playwright and CDP Configuration Guide](/NanmiCoder/MediaCrawler/what-headless-browser-options-are-available-in-media-crawler)

Explore MediaCrawler's headless browser options: Playwright and CDP. Configure headless mode using config settings or CLI flags for efficient web scraping.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Configure Rate Limiting and Sleep Intervals in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-rate-limiting-and-sleep-intervals-in-media-crawler)

Learn how to configure rate limiting and sleep intervals in MediaCrawler. Control request throttling efficiently using CRAWLER_MAX_SLEEP_SEC for consistent intervals.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Configure Media Downloads for Videos in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-media-downloads-for-videos-in-media-crawler)

Learn how to configure media downloads for videos in MediaCrawler. Enable video fetching and set your download path for automatic storage.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Configure Media Downloads for Images in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-media-downloads-for-images-in-media-crawler)

Learn how to configure media downloads for images in MediaCrawler. Enable image grabbing by setting ENABLE_GET_MEIDAS to True and customize save paths in configbase_configpy.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Set Up and Use the MediaCrawler WebUI API Server](/NanmiCoder/MediaCrawler/how-to-set-up-and-use-the-media-crawler-webui-api-server)

Learn how to set up and use the MediaCrawler WebUI API server. Control crawler processes, access data, and stream logs with this FastAPI REST service. Start it with uvicorn.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Handle Verification Challenges in MediaCrawler: Slider Captchas and SMS Verification](/NanmiCoder/MediaCrawler/how-to-handle-verification-challenges-in-media-crawler)

Learn how to handle verification challenges in MediaCrawler. Automate slider captchas with image processing and SMS verification via FastAPI webhooks for seamless operation.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Handle Sliding CAPTCHA in MediaCrawler: A Complete Technical Guide](/NanmiCoder/MediaCrawler/how-to-handle-sliding-captcha-in-media-crawler)

Learn how to handle sliding CAPTCHA in MediaCrawler! This guide details image processing, OpenCV template matching, and realistic mouse track generation for automated solutions.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Configure Custom Stop Words for Word Clouds in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-custom-stop-words-for-word-clouds-in-media-crawler)

Learn how to configure custom stop words for word clouds in MediaCrawler. Easily set the STOP_WORDS_FILE constant to customize your word cloud generation. Improve your text analysis.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Generate Word Clouds from MediaCrawler Comment Data](/NanmiCoder/MediaCrawler/how-to-generate-word-clouds-from-media-crawler-comment-data)

Easily generate word clouds from MediaCrawler comment data. Enable comment gathering and word cloud generation in your config, then run the script to visualize comment frequency.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Crawl Second-Level (Nested) Comments with MediaCrawler](/NanmiCoder/MediaCrawler/how-to-crawl-second-level-nested-comments-with-media-crawler)

Learn how to crawl second-level comments with MediaCrawler. Extract parent comments and their IDs to fetch nested replies efficiently. Improve your data. Get started now.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Crawl First-Level Comments with MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/how-to-crawl-first-level-comments-with-media-crawler)

Learn how to crawl first-level comments using MediaCrawler. This guide explains how to use the TieBaExtractor method for efficient data extraction of user details, content, and metadata.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Configure Concurrent Crawling in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-concurrent-crawling-in-media-crawler)

Learn how to configure concurrent crawling in MediaCrawler. Adjust the MAX_CONCURRENCY_NUM setting to speed up your crawling process efficiently.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Set Up Static Proxy Providers in MediaCrawler: Complete Configuration Guide](/NanmiCoder/MediaCrawler/how-to-set-up-static-proxy-providers-in-media-crawler)

Learn how to set up static proxy providers in MediaCrawler with our complete guide. Configure IP proxy settings easily for enhanced crawling.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Set Up Proxy Rotation with KuaiDaili in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-set-up-proxy-rotation-with-kuaidaili-in-media-crawler)

Easily set up proxy rotation with KuaiDaili in MediaCrawler. Enable IP proxy, add credentials to .env, and let ProxyIpPool manage proxies automatically for seamless scraping.

- Tags: how-to-guide
- Published: 2026-06-27

### [MediaCrawler Crawler Modes: Search, Detail, and Creator Explained](/NanmiCoder/MediaCrawler/what-are-the-different-crawler-modes-in-media-crawler-search-detail-creator)

Explore MediaCrawler's three modes: Search for keywords, Detail for specific IDs, and Creator for full profiles. Efficiently gather media data with this powerful tool.

- Tags: deep-dive
- Published: 2026-06-27

### [How to Use Cookie-Based Login in MediaCrawler: Complete Guide](/NanmiCoder/MediaCrawler/how-to-use-cookie-based-login-in-media-crawler)

Learn how to use cookie-based login in MediaCrawler with our complete guide. Bypass verification using CLI flags or config settings for easier access.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Use Phone Number Login in MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/how-to-use-phone-number-login-in-media-crawler)

Learn how to use phone number login in MediaCrawler with our complete guide. Configure CLI and environment variables for seamless access. NanmiCoder/MediaCrawler

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Use QR Code Login in MediaCrawler: Complete Implementation Guide](/NanmiCoder/MediaCrawler/how-to-use-qr-code-login-in-media-crawler)

Learn how to use QR code login in MediaCrawler with our complete guide. Implement seamless authentication using the --lt qrcode CLI flag for efficient scraping. Get started today!

- Tags: how-to-guide
- Published: 2026-06-27

### [What Social Media Platforms Are Supported by MediaCrawler: The Complete 2024 Guide](/NanmiCoder/MediaCrawler/what-social-media-platforms-are-supported-by-media-crawler)

Discover the social media platforms MediaCrawler supports in 2024. Access XiaoHongShu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu with this comprehensive guide.

- Tags: getting-started
- Published: 2026-06-27

### [MediaCrawler Data Storage Options: CSV, JSON, MongoDB, Excel, and Relational Databases](/NanmiCoder/MediaCrawler/what-are-the-data-storage-options-available-in-media-crawler)

Explore MediaCrawler data storage options including CSV, JSON, MongoDB, Excel, and relational databases. Choose the best format for your project.

- Tags: architecture
- Published: 2026-06-27

### [How to Configure CDP Mode for Stealth Browsing in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-configure-cdp-mode-for-stealth-browsing-in-media-crawler)

Learn how to configure CDP mode for stealth browsing in MediaCrawler. This guide explains connecting to a remote debugging endpoint and injecting anti-fingerprinting scripts for enhanced privacy.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Select Platform-Specific Crawlers at Runtime Using Factory Pattern in MediaCrawler](/NanmiCoder/MediaCrawler/how-to-select-platform-specific-crawlers-at-runtime-using-factory-pattern-in-media-crawler)

Learn how to select platform-specific crawlers at runtime using the Factory pattern in MediaCrawler. This guide details dynamic crawler selection via CLI arguments.

- Tags: how-to-guide
- Published: 2026-06-27

### [How to Extend MediaCrawler's Platform Support: A Complete Developer Guide](/NanmiCoder/MediaCrawler/how-to-extend-mediacrawlers-platform-support)

Learn how to extend MediaCrawler's platform support by adding new platforms. This developer guide covers platform enumeration, configuration, store implementation, and factory registration. Boost your MediaCrawler integration t...

- Tags: how-to-guide
- Published: 2026-06-26

### [MediaCrawler Structure and Abstract Base Classes: A Deep Dive into the Modular Architecture](/NanmiCoder/MediaCrawler/mediacrawler-structure-and-abstract-base-classes)

Explore the MediaCrawler structure and abstract base classes. Understand its modular architecture, factory pattern, and how it enables rapid extension to new platforms.

- Tags: deep-dive
- Published: 2026-06-26

### [How to Run MediaCrawler API Server: Complete Setup and Configuration Guide](/NanmiCoder/MediaCrawler/how-to-run-mediacrawler-api-server)

Learn how to run the MediaCrawler API server with this complete setup and configuration guide. Install dependencies, configure .env, and launch the server easily.

- Tags: how-to-guide
- Published: 2026-06-26

### [MediaCrawler IP Proxy Pool Integration: Architecture and Implementation Guide](/NanmiCoder/MediaCrawler/mediacrawler-ip-proxy-pool-integration)

Learn how to integrate MediaCrawler's IP proxy pool. This guide details the architecture and implementation, covering automatic rotation, validation, and caching of HTTP proxies.

- Tags: architecture
- Published: 2026-06-26

### [How to Store Crawled Data in MongoDB with MediaCrawler: A Complete Guide](/NanmiCoder/MediaCrawler/how-to-store-crawled-data-in-mongodb-with-mediacrawler)

Learn to store crawled data in MongoDB using MediaCrawler. This guide details configuring your mongodb_config for seamless data persistence with Motor async connections and automatic upserts.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Store Crawled Data in MySQL with MediaCrawler: Complete Configuration Guide](/NanmiCoder/MediaCrawler/how-to-store-crawled-data-in-mysql-with-mediacrawler)

Learn how to store crawled data in MySQL with MediaCrawler. Configure the storage option to db for efficient data persistence using SQLAlchemy and asyncmy.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Store Crawled Data in SQLite with MediaCrawler](/NanmiCoder/MediaCrawler/how-to-store-crawled-data-in-sqlite-with-mediacrawler)

Learn to store crawled data in SQLite using MediaCrawler. This guide shows you how to leverage the built-in SQLite backend for efficient local data persistence without a database server.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Store Crawled Data in JSON with MediaCrawler](/NanmiCoder/MediaCrawler/how-to-store-crawled-data-in-json-with-mediacrawler)

Learn how to store crawled data in JSON using MediaCrawler. Simply configure SAVE_DATA_OPTION to json or jsonl and MediaCrawler handles the rest automatically.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Store Crawled Data in CSV with MediaCrawler: Complete Configuration Guide](/NanmiCoder/MediaCrawler/how-to-store-crawled-data-in-csv-with-mediacrawler)

Learn how to store crawled data in CSV with MediaCrawler. Configure your settings and export content comments and creator data effortlessly. Get started today.

- Tags: how-to-guide
- Published: 2026-06-26

### [MediaCrawler Data Storage Options Explained: CSV, JSON, SQLite, MongoDB, and Excel](/NanmiCoder/MediaCrawler/mediacrawler-data-storage-options-explained)

Explore MediaCrawler data storage options including CSV, JSON, SQLite, MongoDB, and Excel. Learn how to manage your crawled data effectively with this guide.

- Tags: deep-dive
- Published: 2026-06-26

### [How to Implement a Custom Crawler in MediaCrawler: A Complete Developer Guide](/NanmiCoder/MediaCrawler/how-to-implement-a-custom-crawler-in-mediacrawler)

Implement a custom crawler in MediaCrawler by subclassing AbstractCrawler. Follow this developer guide to easily add your own crawlers and extend MediaCrawler's capabilities.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Use MediaCrawler CDP Mode with an Existing Chrome Browser](/NanmiCoder/MediaCrawler/using-mediacrawler-cdp-mode-with-existing-chrome)

Connect MediaCrawler to your existing Chrome browser using CDP mode. Reuse sessions and cookies, avoid detection, and speed up your scraping tasks. Learn how now.

- Tags: how-to-guide
- Published: 2026-06-26

### [MediaCrawler Playwright Automation Setup: Complete CDP Browser Management Guide](/NanmiCoder/MediaCrawler/mediacrawler-playwright-automation-setup)

Master MediaCrawler Playwright automation setup with our complete CDP browser management guide. Automate authentication and data extraction on Zhihu and Xiaohongshu effortlessly.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Use MediaCrawler for Tieba Data: A Complete Guide to Baidu Tieba Scraping](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-for-tieba-data)

Learn how to use MediaCrawler for Tieba data. Scrape Baidu Tieba content with Playwright automation. Extract threads, search keywords, and get creator info easily.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Use MediaCrawler for Weibo Data: A Complete Guide to Scraping Weibo Posts](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-for-weibo-data)

Learn how to use MediaCrawler for Weibo data with this complete guide. Scrape Weibo posts efficiently using this powerful Python framework and its three-layer architecture.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Use MediaCrawler for Bilibili Data: Complete Guide to Video and Comment Scraping](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-for-bilibili-data)

Learn to scrape Bilibili videos and comments with MediaCrawler's CLI or Python API. This complete guide covers data extraction, storage, and headless browser automation for NanmiCoder users.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Use MediaCrawler for Kuaishou Data: Configuration, Modes, and Implementation](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-for-kuaishou-data)

Learn how to use MediaCrawler for Kuaishou data. Explore configuration, crawling modes, and async Playwright & GraphQL implementation for video metadata, comments, and creator profiles.

- Tags: how-to-guide
- Published: 2026-06-26

### [How to Use MediaCrawler for Douyin Data: Architecture and Implementation Guide](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-for-douyin-data)

Learn how to use MediaCrawler for Douyin data with this architecture and implementation guide. Our Python client handles anti-scraping, sessions, and proxies for seamless data extraction.

- Tags: getting-started
- Published: 2026-06-26

### [How to Use MediaCrawler for Xiaohongshu Data: Complete Setup Guide](/NanmiCoder/MediaCrawler/how-to-use-mediacrawler-for-xiaohongshu-data)

Learn how to use MediaCrawler for Xiaohongshu data. This complete setup guide covers QR-code authentication, URL parsing, and async data retrieval for notes and profiles.

- Tags: getting-started
- Published: 2026-06-26

