How to Generate a Word Cloud from Comments in MediaCrawler with Custom and Stop Words

You can generate a word cloud from comments in MediaCrawler by enabling ENABLE_GET_WORDCLOUD in config/base_config.py, then using the AsyncWordCloudGenerator class from tools/words.py to process comment data with custom jieba dictionary entries and filtered stop words.

MediaCrawler is a powerful open-source framework for scraping content from Chinese social media platforms like Xiaohongshu and Weibo. When analyzing scraped comments, visualizing term frequency through word clouds helps identify trending topics and user sentiment patterns. This guide explains how to generate a word cloud from comments in MediaCrawler with custom and stop words using the built-in asynchronous generator and configuration system.

Understanding the Word Cloud Architecture

The word cloud functionality in MediaCrawler is built around a specialized asynchronous class that handles text processing, tokenization, and image generation.

The AsyncWordCloudGenerator Class

The core implementation resides in tools/words.py, which defines the AsyncWordCloudGenerator class. This class orchestrates the entire pipeline: loading stop words, registering custom dictionary entries with jieba, computing term frequencies, and rendering the final visualization. It uses plot_lock to ensure thread-safe rendering when multiple crawler instances run concurrently.

Configuration Settings

All customizable parameters are centralized in config/base_config.py. The system relies on four key settings:

  • ENABLE_GET_WORDCLOUD: Boolean toggle to activate the feature
  • CUSTOM_WORDS: Dictionary mapping phrases to categories (e.g., {'机器学习': '技术'})
  • STOP_WORDS_FILE: Path to a plain-text file containing one stop word per line
  • FONT_PATH: Path to a Chinese-compatible TTF font file (default: ./docs/STZHONGS.TTF)

Step-by-Step Implementation Process

The generator follows a specific pipeline to transform raw comment text into a visual word cloud.

Initialization and Setup

When you instantiate AsyncWordCloudGenerator(), the __init__ method performs two critical operations:

  1. Loads the stop word list from config.STOP_WORDS_FILE into a Python set for O(1) lookup performance
  2. Iterates through config.CUSTOM_WORDS and registers each phrase using jieba.add_word(phrase), ensuring multi-character terms are treated as single tokens during segmentation

Processing Raw Comment Data

The primary entry point is generate_word_frequency_and_cloud(data, save_words_prefix). This method expects data to be a list of dictionaries, where each dictionary represents a comment and must contain a content key:

comments = [
    {"content": "零几是一个年份词,值得关注", "author": "user1"},
    {"content": "高频词在技术文档里经常出现", "like_count": 5}
]

The method concatenates all content values into a single text string for batch processing.

Tokenization and Filtering

The generator processes text using the following sequence:

  1. Tokenization: Calls jieba.lcut(all_text) to segment Chinese text into individual tokens
  2. Filtering: Removes any token that exists in the stop word set or contains only whitespace
  3. Frequency Analysis: Uses collections.Counter to calculate occurrence statistics for each remaining token

Generating Output Files

The method produces two artifacts:

  • JSON frequency file: Named <prefix>_word_freq.json, containing the complete word frequency mapping written asynchronously using aiofiles
  • PNG word cloud: Named <prefix>_word_cloud.png, generated using the wordcloud library with the top 20 most frequent terms

The image generation acquires plot_lock before rendering to prevent matplotlib threading conflicts, then releases it immediately after saving the file.

Configuring Custom Words and Stop Words

Fine-tuning the word cloud requires modifying the configuration files to control exactly which terms appear in your visualization.

Adding Custom Words

Edit config/base_config.py to define phrases that jieba should preserve as single tokens:

CUSTOM_WORDS = {
    "机器学习": "技术领域",
    "零几": "年份",
    "深度学习": "技术领域",
    "高频词": "专业术语"
}

Each key is automatically registered via jieba.add_word() during generator initialization, preventing the tokenizer from splitting these phrases into separate characters.

Managing Stop Words

Create or edit the file specified by STOP_WORDS_FILE (commonly ./docs/hit_stopwords.txt). Populate it with one word per line to exclude from analysis:


的
了
在
是
我
你

These words are filtered out immediately after tokenization, ensuring common particles and pronouns do not dominate your visualization.

Font Configuration for Chinese Characters

Ensure the file specified by FONT_PATH exists and supports Chinese glyphs. The default STZHONGS.TTF provides proper rendering for Simplified Chinese characters. Without a valid Chinese font, the word cloud will display squares or gibberish instead of text.

Practical Code Examples

Standalone Usage

You can use the generator outside the main crawler pipeline for testing or batch processing existing comment data:

import asyncio
from tools.words import AsyncWordCloudGenerator
import config

# Sample comment data structure

comments = [
    {"content": "零几是一个年份词,值得关注"},
    {"content": "高频词在技术文档里经常出现"},
    {"content": "普通词也会被统计"},
]

# Enable feature and configure

config.ENABLE_GET_WORDCLOUD = True
config.CUSTOM_WORDS = {"零几": "年份", "高频词": "术语"}

async def main():
    gen = AsyncWordCloudGenerator()
    await gen.generate_word_frequency_and_cloud(
        data=comments, 
        save_words_prefix="demo"
    )
    print("Files generated: demo_word_freq.json and demo_word_cloud.png")

asyncio.run(main())

Integration Within the Crawler Pipeline

Inside platform-specific crawler implementations (typically after collecting comments in extractor.py or client.py):

from tools.words import AsyncWordCloudGenerator
import config
import asyncio

async def process_comments(comments, video_id):
    if config.ENABLE_GET_WORDCLOUD:
        gen = AsyncWordCloudGenerator()
        await gen.generate_word_frequency_and_cloud(
            data=comments,
            save_words_prefix=f"output/{video_id}"
        )

# Schedule during crawl completion

asyncio.create_task(process_comments(comments, video_id))

This ensures word clouds are generated asynchronously without blocking the main crawling logic.

Summary

  • Enable the feature by setting ENABLE_GET_WORDCLOUD = True in config/base_config.py
  • Customize vocabulary by adding entries to CUSTOM_WORDS to preserve multi-character phrases as single tokens
  • Filter noise by populating the stop words file referenced by STOP_WORDS_FILE with terms to exclude
  • Use the generator by importing AsyncWordCloudGenerator from tools/words.py and calling generate_word_frequency_and_cloud() with your comment data
  • Ensure Chinese support by verifying the font file at FONT_PATH exists and contains Chinese glyphs
  • Thread safety is handled automatically via plot_lock when rendering images in concurrent crawling scenarios

Frequently Asked Questions

How do I prevent specific phrases from being split into separate words?

Add the phrases to the CUSTOM_WORDS dictionary in config/base_config.py. Each key is automatically registered with jieba.add_word() during generator initialization, forcing the tokenizer to treat the entire phrase as a single token. For example, adding "机器学习": "tech" ensures "机器学习" is counted as one term rather than separate characters.

Where should I place my custom stop words file?

Specify the absolute or relative path in config/base_config.py using the STOP_WORDS_FILE variable. The default configuration typically points to ./docs/hit_stopwords.txt. The file should contain plain text with one stop word per line. The generator loads these into a set during initialization to efficiently filter them during tokenization.

Why does my word cloud display squares or incorrect characters instead of Chinese text?

This indicates the font file specified by FONT_PATH in config/base_config.py is missing or lacks Chinese glyph support. Ensure the TTF file exists at the specified path and contains Chinese characters. The repository defaults to ./docs/STZHONGS.TTF, which you must provide or replace with another Chinese-compatible font like SimHei or Source Han Sans.

Can I generate word clouds from existing comment data without running the full crawler?

Yes. Import AsyncWordCloudGenerator from tools/words.py directly into your own script. Prepare your data as a list of dictionaries with a content key, ensure ENABLE_GET_WORDCLOUD is set to True in your config, and call generate_word_frequency_and_cloud() asynchronously. This allows offline processing of previously scraped comment JSON files or data from other sources.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →