# MediaCrawler Logging: How the Open-Source Crawler Handles Python Logging

> Explore MediaCrawler's robust Python logging capabilities. Discover its centralized setup, structured output, configurable levels, and noise suppression for seamless debugging.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-12

---

**MediaCrawler centralizes its logging setup in [`tools/utils.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/utils.py), providing a named `"MediaCrawler"` logger with structured console output, configurable log levels, and built-in noise suppression for third-party libraries like `httpx` and `jieba`.**

The [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) repository implements a lightweight, extensible **MediaCrawler logging** system using Python's standard library. This design prioritizes out-of-the-box functionality while allowing developers to extend handlers and formats without modifying core crawler logic.

## Centralized Configuration in tools/utils.py

The entire logging system initializes through a single function: `init_loging_config()` in [`tools/utils.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/utils.py). When the module is imported, this function executes automatically and returns a configured logger instance.

```python

# tools/utils.py (lines 29-40)

def init_loging_config():
    level = logging.INFO
    logging.basicConfig(
        level=level,
        format="%(asctime)s %(name)s %(levelname)s (%(filename)s:%(lineno)d) - %(message)s",
        datefmt='%Y-%m-%d %H:%M:%S'
    )
    _logger = logging.getLogger("MediaCrawler")
    _logger.setLevel(level)

    # Disable httpx INFO level logs

    logging.getLogger("httpx").setLevel(logging.WARNING)

    return _logger

logger = init_loging_config()

```

The default log format includes **timestamps**, **logger name**, **severity level**, **source file**, and **line number**—critical for debugging distributed crawl operations.

## Key Logging Features

### Named Logger for Selective Control

The code creates a dedicated `"MediaCrawler"` logger via `logging.getLogger("MediaCrawler")`. This naming convention enables:

- Runtime level changes without affecting other loggers
- Filter rules in external logging configurations
- Structured routing to different handlers or aggregation services

### Noise Suppression for Third-Party Libraries

MediaCrawler deliberately quiets verbose dependencies. In [`tools/utils.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/utils.py), `httpx` is forced to `WARNING` level. Additionally, [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) silences `jieba`:

```python

# tools/words.py

logging.getLogger('jieba').setLevel(logging.WARNING)

```

This keeps console output focused on actual crawler events rather than HTTP client internals or text-segmentation debug messages.

### Zero External Dependencies

The implementation uses only Python's standard `logging` module. No `loguru`, `structlog`, or custom handlers are required. The default `StreamHandler` prints directly to console, making the system portable across environments.

## Logging Usage Across the Codebase

All functional modules import the shared logger via `from tools import utils` and call standard methods: `utils.logger.info()`, `utils.logger.warning()`, `utils.logger.error()`, and `utils.logger.debug()`.

### Word-Cloud Generation

[`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) logs lock acquisition and cloud generation events (lines 60-73), tracking when the multiprocessing lock is held and when visualization completes.

### CDP Browser Management

[`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) contains extensive instrumentation for browser lifecycle events—connection attempts, page navigation, errors, and cleanup operations.

### Data Store Operations

Database implementations log successful persistence. For example:

- [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py) (lines 218-235): Logs MongoDB insertions for Zhihu content
- [`store/weibo/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/weibo/_store_impl.py) (lines 268-285): Logs Weibo store updates
- [`store/tieba/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/tieba/_store_impl.py): Similar patterns for Tieba data

## Practical Code Examples

### Basic Module Integration

```python
from tools import utils

def fetch_data(url: str):
    utils.logger.info(f"Fetching data from {url}")
    try:
        # ... request logic ...

        utils.logger.debug("Response headers: %s", response.headers)
    except Exception as exc:
        utils.logger.error(f"Failed to fetch {url}: {exc}")
        raise

```

### Runtime Debug Activation

```python
from tools import utils
import logging

utils.logger.setLevel(logging.DEBUG)
utils.logger.debug("Debug verbosity enabled for troubleshooting")

```

### Adding Persistent File Logs

```python
import logging
from tools import utils

file_handler = logging.FileHandler("media_crawler.log", encoding="utf-8")
file_handler.setFormatter(logging.Formatter(
    "%(asctime)s %(levelname)s %(message)s"
))
utils.logger.addHandler(file_handler)

utils.logger.info("File handler attached—logs now persist to media_crawler.log")

```

## Core Logging Files Reference

| File | Purpose |
|------|---------|
| [`tools/utils.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/utils.py) | Core logging configuration and shared `utils.logger` instance |
| [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) | Demonstrates `jieba` suppression and lock-event logging |
| [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) | Browser lifecycle and error instrumentation |
| [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py) | Database operation logging for Zhihu |
| [`store/weibo/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/weibo/_store_impl.py) | Database operation logging for Weibo |
| [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) | Application entry—logging active via `utils` import |

## Summary

- **MediaCrawler logging** uses a single initialization function in [`tools/utils.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/utils.py) for uniform configuration
- A **named `"MediaCrawler"` logger** enables selective filtering and extension
- **Third-party noise suppression** keeps `httpx` and `jieba` at `WARNING` level
- **Standard library only**—no external dependencies required
- **Console-first design** with clear paths to add file handlers or structured formats

## Frequently Asked Questions

### How do I change the MediaCrawler log level to DEBUG?

Import `utils` from `tools` and call `setLevel()` on the logger instance: `from tools import utils; utils.logger.setLevel(logging.DEBUG)`. This takes effect immediately without restarting the crawler.

### Where does MediaCrawler save log files by default?

It doesn't. The default configuration uses only a console `StreamHandler`. To persist logs, attach a `FileHandler` to `utils.logger` as shown in the examples above.

### Why am I seeing httpx logs despite setting INFO level?

The initialization code explicitly sets `httpx` to `WARNING`. If you're seeing INFO-level httpx output, verify you're importing `tools.utils` before any other modules that might trigger httpx imports.

### Can I use MediaCrawler's logger with a third-party logging framework?

Yes. The `utils.logger` object is a standard `logging.Logger` instance compatible with `loguru`, `structlog`, or centralized aggregators. You can replace handlers or wrap the logger without modifying [`tools/utils.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/utils.py).