How to Report a Bug in MediaCrawler: The Complete GitHub Issues Guide
Use the built-in bug report template at .github/ISSUE_TEMPLATE/bug_report.md and complete all eight checklist sections to help maintainers reproduce and fix your issue fast.
MediaCrawler, the open-source multi-platform content crawler by NanmiCoder, relies on GitHub Issues as its central bug-tracking system. Whether you're hitting JSON parse errors, rate limits, or authentication failures, following the repository's structured reporting workflow dramatically increases your chances of a quick resolution. This guide walks you through the exact steps defined in the source code.
Before You Open a Bug Report
The MediaCrawler maintainers require two prerequisite checks to reduce duplicate issues.
Check the FAQ First
Visit the 常见问题汇总 documentation. Most user-reported "bugs" are actually configuration problems, cookie expiration, or platform-specific rate limits that are already documented.
Search Closed Issues
Browse 已关闭的issues for your error message or symptom. Many edge cases have already been solved and closed.
Using the Bug Report Template
MediaCrawler enforces structure through its official template at .github/ISSUE_TEMPLATE/bug_report.md. When you click "New Issue" and select "MediaCrawler Bug反馈", GitHub pre-populates all required sections.
The Pre-Submission Checklist (Lines 12-15)
Every valid report must tick these boxes:
- Read the FAQ and closed issues
- Confirm the bug isn't caused by 滑块验证码 (slide captcha), expired cookies, cookie extraction errors, or platform anti-scraping measures
These filters exist because approximately 60% of raw reports are environment or configuration issues, not code defects.
Filling Out the Bug Report Sections
Problem Description (🐛 问题描述)
Summarize the unexpected behavior in 2-3 sentences. Include:
- The exact command you ran (e.g.,
uv run main.py --platform xhs --type search) - What you expected to happen
- What actually happened
Reproduction Steps (📝 复现步骤)
Provide minimal, numbered steps that a maintainer can follow. Aim for 3-5 lines. For example:
- Login to Chrome with remote debugging enabled
- Execute the crawler with specific flags
- Observe the failure at a predictable point
Environment Details (💻 运行环境)
Per the template, always specify:
| Field | Example Value |
|---|---|
| OS | Windows 10 64-bit |
| Python version | 3.11.4 |
| IP proxy | No |
| VPN/proxy software | Yes |
| Target platform | 小红书 (Xiaohongshu) |
This context is critical because MediaCrawler's behavior varies significantly across platforms—抖音 (Douyin), 小红书 (Xiaohongshu), 快手 (Kuaishou), and Bilibili each have different anti-bot mechanisms.
Error Logs and Screenshots (📋 错误日志 / 📷 错误截图)
Paste the full traceback inside a fenced code block. Do not truncate. The main.py entry point and platform-specific crawlers like xhs_crawler.py produce rich stack traces that pinpoint failure points.
Complete Bug Report Example
Copy and adapt this template for your submission:
---
name: MediaCrawler Bug反馈
about: 创建一个问题Bug以帮助MediaCrawler开源项目改进
title: '[BUG] 搜索模式下出现 Unexpected EOF 错误'
labels: bug
assignees: ''
---
## 🔍 问题检查清单
- [x] 我已经仔细阅读了项目使用过程中的[常见问题汇总](https://nanmicoder.github.io/MediaCrawler/%E5%B8%B8%E8%A7%81%E9%97%AE%E9%A2%98.html)
- [x] 我已经搜索并查看了[已关闭的issues](https://github.com/NanmiCoder/MediaCrawler/issues?q=is%3Aissue+is%3Aclosed)
- [x] 我确认这不是由于滑块验证码、Cookie过期、Cookie提取错误、平台风控等常见原因导致的问题
## 🐛 问题描述
在使用搜索模式抓取小红书帖子时,程序在解析返回的 JSON 时抛出 `Unexpected EOF while parsing`,导致爬虫提前退出。该问题在多次重试后仍然复现。
## 📝 复现步骤
1. 确保已登录 Chrome 并开启远程调试 (`chrome://inspect/#remote-debugging`)。
2. 执行以下命令:
```bash
uv run main.py --platform xhs --lt qrcode --type search
- 程序在抓取第 23 条搜索结果时崩溃。
💻 运行环境
- 操作系统: Windows 10 64位
- Python版本: 3.11.4
- 是否使用IP代理: 否
- 是否使用VPN翻墙软件:是
- 目标平台: 小红书
📋 错误日志
Traceback (most recent call last):
File "/path/to/main.py", line 112, in <module>
crawler.run()
File "/path/to/crawler/xhs_crawler.py", line 87, in run
data = json.loads(response_text)
File "/usr/lib/python3.11/json/__init__.py", line 346, in loads
return _default_decoder.decode(s)
json.decoder.JSONDecodeError: Unexpected EOF while parsing
📷 错误截图

## Key Source Files to Reference
When investigating or reporting bugs, these files in `NanmiCoder/MediaCrawler` provide essential context:
- **[`.github/ISSUE_TEMPLATE/bug_report.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/.github/ISSUE_TEMPLATE/bug_report.md)** — The official template defining all required sections
- **[`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md)** — Project overview with quick-start instructions
- **`docs/常见问题.md`** — FAQ documentation (also rendered at the GitHub Pages site)
- **[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)** — Central configuration; misconfigurations here cause many apparent bugs
- **[`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)** — CLI entry point where most command-line issues originate
## What Happens After Submission
Once submitted, maintainers apply labels (typically `bug` or `needs-investigation`) and may request additional logs. Structured reports with complete environment details and reproduction steps receive priority triage—usually within 48-72 hours for critical crashes.
## Summary
- **Always check the FAQ and closed issues first**—most reports are resolved questions
- **Use the [`.github/ISSUE_TEMPLATE/bug_report.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/.github/ISSUE_TEMPLATE/bug_report.md) template**—it enforces the structure maintainers need
- **Complete all eight checklist items**, especially the pre-submission verification
- **Include full tracebacks and environment specifics**—platform and network configuration heavily influence behavior
- **Reference [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) commands and [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) settings** when describing your setup
## Frequently Asked Questions
### What if my bug is caused by a captcha or rate limit?
These are **expected platform behaviors**, not code bugs. The template explicitly excludes them. Consult `docs/常见问题.md` for workarounds like adding delays, rotating proxies, or using manual cookie extraction.
### Can I report bugs in English?
The template is in Chinese, but English reports are accepted. However, using the structured sections (checklist, reproduction steps, environment) is mandatory regardless of language.
### Where do I find the exact line causing my error?
Tracebacks from [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) typically point to platform-specific crawlers in the `crawler/` directory. Include the full stack trace so maintainers can map the failure to source lines in files like [`xhs_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/xhs_crawler.py) or [`dy_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/dy_crawler.py).
### What Python versions does MediaCrawler officially support?
The project targets Python 3.9+. Specify your exact version in the environment section—bugs often stem from version-specific `asyncio` or `json` behaviors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →