MediaCrawler Logging: How the Open-Source Crawler Handles Python Logging

MediaCrawler centralizes its logging setup in tools/utils.py, providing a named "MediaCrawler" logger with structured console output, configurable log levels, and built-in noise suppression for third-party libraries like httpx and jieba.

The NanmiCoder/MediaCrawler repository implements a lightweight, extensible MediaCrawler logging system using Python's standard library. This design prioritizes out-of-the-box functionality while allowing developers to extend handlers and formats without modifying core crawler logic.

Centralized Configuration in tools/utils.py

The entire logging system initializes through a single function: init_loging_config() in tools/utils.py. When the module is imported, this function executes automatically and returns a configured logger instance.


# tools/utils.py (lines 29-40)

def init_loging_config():
    level = logging.INFO
    logging.basicConfig(
        level=level,
        format="%(asctime)s %(name)s %(levelname)s (%(filename)s:%(lineno)d) - %(message)s",
        datefmt='%Y-%m-%d %H:%M:%S'
    )
    _logger = logging.getLogger("MediaCrawler")
    _logger.setLevel(level)

    # Disable httpx INFO level logs

    logging.getLogger("httpx").setLevel(logging.WARNING)

    return _logger

logger = init_loging_config()

The default log format includes timestamps, logger name, severity level, source file, and line number—critical for debugging distributed crawl operations.

Key Logging Features

Named Logger for Selective Control

The code creates a dedicated "MediaCrawler" logger via logging.getLogger("MediaCrawler"). This naming convention enables:

  • Runtime level changes without affecting other loggers
  • Filter rules in external logging configurations
  • Structured routing to different handlers or aggregation services

Noise Suppression for Third-Party Libraries

MediaCrawler deliberately quiets verbose dependencies. In tools/utils.py, httpx is forced to WARNING level. Additionally, tools/words.py silences jieba:


# tools/words.py

logging.getLogger('jieba').setLevel(logging.WARNING)

This keeps console output focused on actual crawler events rather than HTTP client internals or text-segmentation debug messages.

Zero External Dependencies

The implementation uses only Python's standard logging module. No loguru, structlog, or custom handlers are required. The default StreamHandler prints directly to console, making the system portable across environments.

Logging Usage Across the Codebase

All functional modules import the shared logger via from tools import utils and call standard methods: utils.logger.info(), utils.logger.warning(), utils.logger.error(), and utils.logger.debug().

Word-Cloud Generation

tools/words.py logs lock acquisition and cloud generation events (lines 60-73), tracking when the multiprocessing lock is held and when visualization completes.

CDP Browser Management

tools/cdp_browser.py contains extensive instrumentation for browser lifecycle events—connection attempts, page navigation, errors, and cleanup operations.

Data Store Operations

Database implementations log successful persistence. For example:

Practical Code Examples

Basic Module Integration

from tools import utils

def fetch_data(url: str):
    utils.logger.info(f"Fetching data from {url}")
    try:
        # ... request logic ...

        utils.logger.debug("Response headers: %s", response.headers)
    except Exception as exc:
        utils.logger.error(f"Failed to fetch {url}: {exc}")
        raise

Runtime Debug Activation

from tools import utils
import logging

utils.logger.setLevel(logging.DEBUG)
utils.logger.debug("Debug verbosity enabled for troubleshooting")

Adding Persistent File Logs

import logging
from tools import utils

file_handler = logging.FileHandler("media_crawler.log", encoding="utf-8")
file_handler.setFormatter(logging.Formatter(
    "%(asctime)s %(levelname)s %(message)s"
))
utils.logger.addHandler(file_handler)

utils.logger.info("File handler attached—logs now persist to media_crawler.log")

Core Logging Files Reference

File Purpose
tools/utils.py Core logging configuration and shared utils.logger instance
tools/words.py Demonstrates jieba suppression and lock-event logging
tools/cdp_browser.py Browser lifecycle and error instrumentation
store/zhihu/_store_impl.py Database operation logging for Zhihu
store/weibo/_store_impl.py Database operation logging for Weibo
main.py Application entry—logging active via utils import

Summary

  • MediaCrawler logging uses a single initialization function in tools/utils.py for uniform configuration
  • A named "MediaCrawler" logger enables selective filtering and extension
  • Third-party noise suppression keeps httpx and jieba at WARNING level
  • Standard library only—no external dependencies required
  • Console-first design with clear paths to add file handlers or structured formats

Frequently Asked Questions

How do I change the MediaCrawler log level to DEBUG?

Import utils from tools and call setLevel() on the logger instance: from tools import utils; utils.logger.setLevel(logging.DEBUG). This takes effect immediately without restarting the crawler.

Where does MediaCrawler save log files by default?

It doesn't. The default configuration uses only a console StreamHandler. To persist logs, attach a FileHandler to utils.logger as shown in the examples above.

Why am I seeing httpx logs despite setting INFO level?

The initialization code explicitly sets httpx to WARNING. If you're seeing INFO-level httpx output, verify you're importing tools.utils before any other modules that might trigger httpx imports.

Can I use MediaCrawler's logger with a third-party logging framework?

Yes. The utils.logger object is a standard logging.Logger instance compatible with loguru, structlog, or centralized aggregators. You can replace handlers or wrap the logger without modifying tools/utils.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →