MediaCrawler Logging: How the Open-Source Crawler Handles Python Logging
MediaCrawler centralizes its logging setup in tools/utils.py, providing a named "MediaCrawler" logger with structured console output, configurable log levels, and built-in noise suppression for third-party libraries like httpx and jieba.
The NanmiCoder/MediaCrawler repository implements a lightweight, extensible MediaCrawler logging system using Python's standard library. This design prioritizes out-of-the-box functionality while allowing developers to extend handlers and formats without modifying core crawler logic.
Centralized Configuration in tools/utils.py
The entire logging system initializes through a single function: init_loging_config() in tools/utils.py. When the module is imported, this function executes automatically and returns a configured logger instance.
# tools/utils.py (lines 29-40)
def init_loging_config():
level = logging.INFO
logging.basicConfig(
level=level,
format="%(asctime)s %(name)s %(levelname)s (%(filename)s:%(lineno)d) - %(message)s",
datefmt='%Y-%m-%d %H:%M:%S'
)
_logger = logging.getLogger("MediaCrawler")
_logger.setLevel(level)
# Disable httpx INFO level logs
logging.getLogger("httpx").setLevel(logging.WARNING)
return _logger
logger = init_loging_config()
The default log format includes timestamps, logger name, severity level, source file, and line number—critical for debugging distributed crawl operations.
Key Logging Features
Named Logger for Selective Control
The code creates a dedicated "MediaCrawler" logger via logging.getLogger("MediaCrawler"). This naming convention enables:
- Runtime level changes without affecting other loggers
- Filter rules in external logging configurations
- Structured routing to different handlers or aggregation services
Noise Suppression for Third-Party Libraries
MediaCrawler deliberately quiets verbose dependencies. In tools/utils.py, httpx is forced to WARNING level. Additionally, tools/words.py silences jieba:
# tools/words.py
logging.getLogger('jieba').setLevel(logging.WARNING)
This keeps console output focused on actual crawler events rather than HTTP client internals or text-segmentation debug messages.
Zero External Dependencies
The implementation uses only Python's standard logging module. No loguru, structlog, or custom handlers are required. The default StreamHandler prints directly to console, making the system portable across environments.
Logging Usage Across the Codebase
All functional modules import the shared logger via from tools import utils and call standard methods: utils.logger.info(), utils.logger.warning(), utils.logger.error(), and utils.logger.debug().
Word-Cloud Generation
tools/words.py logs lock acquisition and cloud generation events (lines 60-73), tracking when the multiprocessing lock is held and when visualization completes.
CDP Browser Management
tools/cdp_browser.py contains extensive instrumentation for browser lifecycle events—connection attempts, page navigation, errors, and cleanup operations.
Data Store Operations
Database implementations log successful persistence. For example:
store/zhihu/_store_impl.py(lines 218-235): Logs MongoDB insertions for Zhihu contentstore/weibo/_store_impl.py(lines 268-285): Logs Weibo store updatesstore/tieba/_store_impl.py: Similar patterns for Tieba data
Practical Code Examples
Basic Module Integration
from tools import utils
def fetch_data(url: str):
utils.logger.info(f"Fetching data from {url}")
try:
# ... request logic ...
utils.logger.debug("Response headers: %s", response.headers)
except Exception as exc:
utils.logger.error(f"Failed to fetch {url}: {exc}")
raise
Runtime Debug Activation
from tools import utils
import logging
utils.logger.setLevel(logging.DEBUG)
utils.logger.debug("Debug verbosity enabled for troubleshooting")
Adding Persistent File Logs
import logging
from tools import utils
file_handler = logging.FileHandler("media_crawler.log", encoding="utf-8")
file_handler.setFormatter(logging.Formatter(
"%(asctime)s %(levelname)s %(message)s"
))
utils.logger.addHandler(file_handler)
utils.logger.info("File handler attached—logs now persist to media_crawler.log")
Core Logging Files Reference
| File | Purpose |
|---|---|
tools/utils.py |
Core logging configuration and shared utils.logger instance |
tools/words.py |
Demonstrates jieba suppression and lock-event logging |
tools/cdp_browser.py |
Browser lifecycle and error instrumentation |
store/zhihu/_store_impl.py |
Database operation logging for Zhihu |
store/weibo/_store_impl.py |
Database operation logging for Weibo |
main.py |
Application entry—logging active via utils import |
Summary
- MediaCrawler logging uses a single initialization function in
tools/utils.pyfor uniform configuration - A named
"MediaCrawler"logger enables selective filtering and extension - Third-party noise suppression keeps
httpxandjiebaatWARNINGlevel - Standard library only—no external dependencies required
- Console-first design with clear paths to add file handlers or structured formats
Frequently Asked Questions
How do I change the MediaCrawler log level to DEBUG?
Import utils from tools and call setLevel() on the logger instance: from tools import utils; utils.logger.setLevel(logging.DEBUG). This takes effect immediately without restarting the crawler.
Where does MediaCrawler save log files by default?
It doesn't. The default configuration uses only a console StreamHandler. To persist logs, attach a FileHandler to utils.logger as shown in the examples above.
Why am I seeing httpx logs despite setting INFO level?
The initialization code explicitly sets httpx to WARNING. If you're seeing INFO-level httpx output, verify you're importing tools.utils before any other modules that might trigger httpx imports.
Can I use MediaCrawler's logger with a third-party logging framework?
Yes. The utils.logger object is a standard logging.Logger instance compatible with loguru, structlog, or centralized aggregators. You can replace handlers or wrap the logger without modifying tools/utils.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →