Where Are MediaCrawler Logs Stored? Configuration and File Paths Explained
MediaCrawler writes all log output to standard error (stderr) by default, meaning logs appear in the console and are not persisted to files unless you manually configure a file handler.
MediaCrawler is an open-source social media crawling framework that relies on Python’s standard logging module for diagnostic output. According to the source code in the NanmiCoder/MediaCrawler repository, the application does not create log files automatically—all diagnostic messages stream directly to the terminal. Understanding this logging behavior is critical for monitoring long-running crawl jobs or troubleshooting failed extractions.
Default Log Destination in MediaCrawler
Console-Only Output via basicConfig
The logging system is initialized in tools/utils.py by the function init_loging_config(). This function calls logging.basicConfig() at lines 29-35 without specifying a filename parameter, which forces Python to use the default stream handler pointing to stderr (standard error).
# tools/utils.py (lines 29-35)
def init_loging_config():
logging.basicConfig(
level=level,
format="%(asctime)s %(name)s %(levelname)s %(message)s",
datefmt="%Y-%m-%d %H:%M:%S"
)
Because no file handler is attached, the MediaCrawler logger only emits messages to the console. If you run the crawler in a terminal, you see output immediately; if you run it in the background or via Docker, logs disappear unless you redirect stderr.
The MediaCrawler Logger Instance
Immediately after configuration, the function creates and returns a logger named MediaCrawler at lines 36-38:
# tools/utils.py (lines 36-38)
logger = logging.getLogger("MediaCrawler")
return logger
This instance is imported throughout the codebase (e.g., by crawler implementations in the media_platform directory) via from tools.utils import logger. No other modules in the repository add file handlers, so all log output remains strictly console-based unless you modify the configuration.
How to Configure File-Based Logging
If you need persistent logs for audit trails or debugging, you must explicitly add a FileHandler to the existing logger or modify the initialization code.
Adding a FileHandler to the Existing Logger
The safest approach is to append a file handler to the global logger after importing it, preserving the existing console output while also writing to disk:
import logging
from tools.utils import logger
# Create a file handler
file_handler = logging.FileHandler("media_crawler.log", encoding="utf-8")
formatter = logging.Formatter("%(asctime)s %(name)s %(levelname)s %(message)s")
file_handler.setFormatter(formatter)
# Attach to the MediaCrawler logger
logger.addHandler(file_handler)
# Now this writes to both console and file
logger.info("Crawler session started")
Modifying the Initial Configuration
Alternatively, you can modify tools/utils.py to write logs to a file by default by adding the filename argument to basicConfig():
# Modified tools/utils.py
logging.basicConfig(
level=level,
format="%(asctime)s %(name)s %(levelname)s %(message)s",
datefmt="%Y-%m-%d %H:%M:%S",
filename="media_crawler.log", # NEW: persists logs to file
encoding="utf-8"
)
When filename is present, Python automatically redirects all output to the specified file instead of stderr.
Key Source Files for Log Configuration
tools/utils.py– Contains theinit_loging_config()function (lines 29-38) that establishes theMediaCrawlerlogger with console-only output.- Crawler modules (e.g.,
media_platform/xiaohongshu.py,media_platform/douyin.py) – Import the shared logger fromtools.utilsand emit diagnostic messages throughout the crawl lifecycle.
Summary
- MediaCrawler logs stream to stderr (console) by default, not to files.
- The logger is configured in
tools/utils.pyviainit_loging_config()usinglogging.basicConfig()without afilenameparameter. - To persist logs, manually add a
FileHandlerto theMediaCrawlerlogger or modify thebasicConfig()call intools/utils.pyto include afilenameargument. - No automatic log rotation or archival is implemented in the current codebase.
Frequently Asked Questions
Does MediaCrawler create log files automatically?
No. MediaCrawler does not write logs to disk by default. The init_loging_config() function in tools/utils.py omits the filename parameter in its basicConfig() call, forcing all output to the console (stderr).
How can I redirect MediaCrawler logs to a file?
You can redirect logs by either appending a logging.FileHandler to the existing MediaCrawler logger at runtime or by modifying tools/utils.py to pass filename="your_log.log" to basicConfig(). The first approach preserves console output while adding file persistence.
What log level does MediaCrawler use by default?
The default log level is determined by the level variable passed to basicConfig() in tools/utils.py. Typically, this is set to logging.INFO, but you can adjust it by changing the level parameter in the configuration function or by calling logger.setLevel() after initialization.
Where is the logger initialized in the MediaCrawler codebase?
The logger is initialized in tools/utils.py within the init_loging_config() function (lines 29-38). This function creates a logger named MediaCrawler using logging.getLogger("MediaCrawler") and returns it for use across all crawler modules.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →