How to Generate a Word Cloud from Comments in MediaCrawler with Custom and Stop Words
You can generate a word cloud from comments in MediaCrawler by enabling ENABLE_GET_WORDCLOUD in config/base_config.py, then using the AsyncWordCloudGenerator class from tools/words.py to process comment data with custom jieba dictionary entries and filtered stop words.
MediaCrawler is a powerful open-source framework for scraping content from Chinese social media platforms like Xiaohongshu and Weibo. When analyzing scraped comments, visualizing term frequency through word clouds helps identify trending topics and user sentiment patterns. This guide explains how to generate a word cloud from comments in MediaCrawler with custom and stop words using the built-in asynchronous generator and configuration system.
Understanding the Word Cloud Architecture
The word cloud functionality in MediaCrawler is built around a specialized asynchronous class that handles text processing, tokenization, and image generation.
The AsyncWordCloudGenerator Class
The core implementation resides in tools/words.py, which defines the AsyncWordCloudGenerator class. This class orchestrates the entire pipeline: loading stop words, registering custom dictionary entries with jieba, computing term frequencies, and rendering the final visualization. It uses plot_lock to ensure thread-safe rendering when multiple crawler instances run concurrently.
Configuration Settings
All customizable parameters are centralized in config/base_config.py. The system relies on four key settings:
ENABLE_GET_WORDCLOUD: Boolean toggle to activate the featureCUSTOM_WORDS: Dictionary mapping phrases to categories (e.g.,{'机器学习': '技术'})STOP_WORDS_FILE: Path to a plain-text file containing one stop word per lineFONT_PATH: Path to a Chinese-compatible TTF font file (default:./docs/STZHONGS.TTF)
Step-by-Step Implementation Process
The generator follows a specific pipeline to transform raw comment text into a visual word cloud.
Initialization and Setup
When you instantiate AsyncWordCloudGenerator(), the __init__ method performs two critical operations:
- Loads the stop word list from
config.STOP_WORDS_FILEinto a Pythonsetfor O(1) lookup performance - Iterates through
config.CUSTOM_WORDSand registers each phrase usingjieba.add_word(phrase), ensuring multi-character terms are treated as single tokens during segmentation
Processing Raw Comment Data
The primary entry point is generate_word_frequency_and_cloud(data, save_words_prefix). This method expects data to be a list of dictionaries, where each dictionary represents a comment and must contain a content key:
comments = [
{"content": "零几是一个年份词,值得关注", "author": "user1"},
{"content": "高频词在技术文档里经常出现", "like_count": 5}
]
The method concatenates all content values into a single text string for batch processing.
Tokenization and Filtering
The generator processes text using the following sequence:
- Tokenization: Calls
jieba.lcut(all_text)to segment Chinese text into individual tokens - Filtering: Removes any token that exists in the stop word set or contains only whitespace
- Frequency Analysis: Uses
collections.Counterto calculate occurrence statistics for each remaining token
Generating Output Files
The method produces two artifacts:
- JSON frequency file: Named
<prefix>_word_freq.json, containing the complete word frequency mapping written asynchronously usingaiofiles - PNG word cloud: Named
<prefix>_word_cloud.png, generated using thewordcloudlibrary with the top 20 most frequent terms
The image generation acquires plot_lock before rendering to prevent matplotlib threading conflicts, then releases it immediately after saving the file.
Configuring Custom Words and Stop Words
Fine-tuning the word cloud requires modifying the configuration files to control exactly which terms appear in your visualization.
Adding Custom Words
Edit config/base_config.py to define phrases that jieba should preserve as single tokens:
CUSTOM_WORDS = {
"机器学习": "技术领域",
"零几": "年份",
"深度学习": "技术领域",
"高频词": "专业术语"
}
Each key is automatically registered via jieba.add_word() during generator initialization, preventing the tokenizer from splitting these phrases into separate characters.
Managing Stop Words
Create or edit the file specified by STOP_WORDS_FILE (commonly ./docs/hit_stopwords.txt). Populate it with one word per line to exclude from analysis:
的
了
在
是
我
你
These words are filtered out immediately after tokenization, ensuring common particles and pronouns do not dominate your visualization.
Font Configuration for Chinese Characters
Ensure the file specified by FONT_PATH exists and supports Chinese glyphs. The default STZHONGS.TTF provides proper rendering for Simplified Chinese characters. Without a valid Chinese font, the word cloud will display squares or gibberish instead of text.
Practical Code Examples
Standalone Usage
You can use the generator outside the main crawler pipeline for testing or batch processing existing comment data:
import asyncio
from tools.words import AsyncWordCloudGenerator
import config
# Sample comment data structure
comments = [
{"content": "零几是一个年份词,值得关注"},
{"content": "高频词在技术文档里经常出现"},
{"content": "普通词也会被统计"},
]
# Enable feature and configure
config.ENABLE_GET_WORDCLOUD = True
config.CUSTOM_WORDS = {"零几": "年份", "高频词": "术语"}
async def main():
gen = AsyncWordCloudGenerator()
await gen.generate_word_frequency_and_cloud(
data=comments,
save_words_prefix="demo"
)
print("Files generated: demo_word_freq.json and demo_word_cloud.png")
asyncio.run(main())
Integration Within the Crawler Pipeline
Inside platform-specific crawler implementations (typically after collecting comments in extractor.py or client.py):
from tools.words import AsyncWordCloudGenerator
import config
import asyncio
async def process_comments(comments, video_id):
if config.ENABLE_GET_WORDCLOUD:
gen = AsyncWordCloudGenerator()
await gen.generate_word_frequency_and_cloud(
data=comments,
save_words_prefix=f"output/{video_id}"
)
# Schedule during crawl completion
asyncio.create_task(process_comments(comments, video_id))
This ensures word clouds are generated asynchronously without blocking the main crawling logic.
Summary
- Enable the feature by setting
ENABLE_GET_WORDCLOUD = Trueinconfig/base_config.py - Customize vocabulary by adding entries to
CUSTOM_WORDSto preserve multi-character phrases as single tokens - Filter noise by populating the stop words file referenced by
STOP_WORDS_FILEwith terms to exclude - Use the generator by importing
AsyncWordCloudGeneratorfromtools/words.pyand callinggenerate_word_frequency_and_cloud()with your comment data - Ensure Chinese support by verifying the font file at
FONT_PATHexists and contains Chinese glyphs - Thread safety is handled automatically via
plot_lockwhen rendering images in concurrent crawling scenarios
Frequently Asked Questions
How do I prevent specific phrases from being split into separate words?
Add the phrases to the CUSTOM_WORDS dictionary in config/base_config.py. Each key is automatically registered with jieba.add_word() during generator initialization, forcing the tokenizer to treat the entire phrase as a single token. For example, adding "机器学习": "tech" ensures "机器学习" is counted as one term rather than separate characters.
Where should I place my custom stop words file?
Specify the absolute or relative path in config/base_config.py using the STOP_WORDS_FILE variable. The default configuration typically points to ./docs/hit_stopwords.txt. The file should contain plain text with one stop word per line. The generator loads these into a set during initialization to efficiently filter them during tokenization.
Why does my word cloud display squares or incorrect characters instead of Chinese text?
This indicates the font file specified by FONT_PATH in config/base_config.py is missing or lacks Chinese glyph support. Ensure the TTF file exists at the specified path and contains Chinese characters. The repository defaults to ./docs/STZHONGS.TTF, which you must provide or replace with another Chinese-compatible font like SimHei or Source Han Sans.
Can I generate word clouds from existing comment data without running the full crawler?
Yes. Import AsyncWordCloudGenerator from tools/words.py directly into your own script. Prepare your data as a list of dictionaries with a content key, ensure ENABLE_GET_WORDCLOUD is set to True in your config, and call generate_word_frequency_and_cloud() asynchronously. This allows offline processing of previously scraped comment JSON files or data from other sources.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →