How to Generate Word Clouds from Scraped Data Using MediaCrawler

Enable ENABLE_GET_COMMENTS and ENABLE_GET_WORDCLOUD in config/base_config.py, then run the crawler to automatically produce frequency JSON files and PNG visualizations in data/words/.

MediaCrawler is an open-source asynchronous scraping framework that extracts content from various social media platforms. When you need to visualize textual patterns from scraped comments, the built-in word cloud generator processes raw data into ready-to-use frequency maps and image files without requiring external tools.

Prerequisites and Configuration

Before generating word clouds, you must ensure data persistence and comment extraction are properly configured. In config/base_config.py, set SAVE_DATA_OPTION = "jsonl" (or "json") so the pipeline can read the scraped content. You must also set ENABLE_GET_COMMENTS = True to fetch the textual data that feeds the visualization engine.

The word cloud feature itself is controlled by the ENABLE_GET_WORDCLOUD boolean flag. When this is set to True, the crawler automatically invokes the AsyncWordCloudGenerator class after each run.

How the Word Cloud Pipeline Works

The implementation in tools/words.py follows a three-stage asynchronous pipeline that transforms raw comments into visual outputs.

Data Collection and Text Aggregation

The AsyncWordCloudGenerator class first aggregates all comment content from the scraped dataset. It extracts the content field from each item in the data array and concatenates them into a single string for processing.

Tokenization and Frequency Analysis

The generator uses jieba to segment Chinese text into individual tokens. It filters these tokens against a stop-word list (default: ./docs/hit_stopwords.txt) and builds a frequency Counter from the remaining words. The class also respects the CUSTOM_WORDS dictionary defined in your configuration, adding domain-specific terms to jieba's dictionary during initialization so multi-word phrases are treated as single tokens.

Image Rendering and Output

The top 20 words by frequency are passed to the wordcloud library's WordCloud.generate_from_frequencies() method. Matplotlib renders the visualization using the font specified in FONT_PATH (default: ./docs/STZHONGS.TTF) and saves a PNG file alongside the raw frequency JSON. An asyncio lock (plot_lock) prevents resource conflicts when running multiple concurrent crawls.

Output files are written to data/words/ with the following naming convention:

  • <crawl-id>_word_freq.json – Raw word frequency data
  • <crawl-id>_word_cloud.png – Rendered visualization

Key Configuration Options

Tune these settings in config/base_config.py to control the output:

Setting Description
ENABLE_GET_COMMENTS Must be True to scrape comment text
ENABLE_GET_WORDCLOUD Must be True to trigger generation
SAVE_DATA_OPTION Must be "jsonl" or "json" for the reader to parse files
CUSTOM_WORDS Dictionary mapping phrases to tags (e.g., {"机器学习": "tech"})
STOP_WORDS_FILE Path to line-separated stop-words file
FONT_PATH Path to TrueType font for Chinese character rendering

Code Examples

Running a Full Crawl with Word Cloud Output

This script executes the complete pipeline from configuration to visualization:

import asyncio
import config
from main import run

async def main():
    # Configure options before importing other modules

    config.ENABLE_GET_COMMENTS = True
    config.ENABLE_GET_WORDCLOUD = True
    config.SAVE_DATA_OPTION = "jsonl"
    
    # Run the crawler (platform settings defined in base_config)

    await run()
    
    print("Word cloud files saved to data/words/")

if __name__ == "__main__":
    asyncio.run(main())

The run() function in main/main.py orchestrates the entire process, automatically invoking AsyncWordCloudGenerator when the flags are enabled.

Generating Word Clouds from Existing Data

To re-process existing JSONL files without re-crawling:

import asyncio
import json
import aiofiles
from tools.words import AsyncWordCloudGenerator

async def gen_from_file(jsonl_path: str, prefix: str):
    # Load previously saved comments

    async with aiofiles.open(jsonl_path, 'r', encoding='utf-8') as f:
        lines = await f.readlines()
        data = [json.loads(line) for line in lines]
    
    generator = AsyncWordCloudGenerator()
    await generator.generate_word_frequency_and_cloud(data, prefix)

if __name__ == "__main__":
    asyncio.run(gen_from_file(
        "data/comments/example.jsonl",
        "data/words/example"
    ))

Customizing Stop Words and Tokenization

Add domain-specific vocabulary to improve tokenization accuracy:


# config/base_config.py

CUSTOM_WORDS = {
    "人工智能": "tech",
    "2025": "year",
    "深度学习": "tech"
}

Update your stop-words list in docs/hit_stopwords.txt (one word per line):


的
了
和

The generator calls jieba.add_word() during initialization, ensuring these custom terms are recognized in subsequent runs.

Summary

  • Enable three flags in config/base_config.py: ENABLE_GET_COMMENTS, ENABLE_GET_WORDCLOUD, and set SAVE_DATA_OPTION to "jsonl"
  • Process flow: AsyncWordCloudGenerator in tools/words.py aggregates text, tokenizes with jieba, filters stop-words, and renders via matplotlib
  • Output location: data/words/ directory contains both JSON frequency data and PNG images
  • Customization: Use CUSTOM_WORDS for phrase tokenization and STOP_WORDS_FILE to filter noise
  • Concurrency safe: The plot_lock prevents matplotlib conflicts during async operations

Frequently Asked Questions

Where are the word cloud files saved?

MediaCrawler writes output to data/words/<crawl-id>_word_freq.json and data/words/<crawl-id>_word_cloud.png. The prefix is generated from your crawl parameters and timestamp.

Can I generate word clouds from existing JSONL files without re-crawling?

Yes. Import AsyncWordCloudGenerator from tools/words.py and call generate_word_frequency_and_cloud() with your loaded data array and desired file prefix. This bypasses the scraping stage entirely.

How do I add custom words or phrases to the tokenization?

Define a CUSTOM_WORDS dictionary in config/base_config.py where keys are phrases and values are tags. The generator passes these to jieba.add_word() during initialization, ensuring multi-character terms are treated as single tokens.

Why is my word cloud showing incorrect characters or squares?

This indicates a font issue. Verify that FONT_PATH in config/base_config.py points to a valid Chinese TrueType font (default is ./docs/STZHONGS.TTF). The system must have read access to this file, and the font must support the character set in your scraped data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →