Adding Custom Words and Chinese Fonts for MediaCrawler Word Clouds: A Complete Guide
To add custom words and Chinese fonts in MediaCrawler, define CUSTOM_WORDS and FONT_PATH in config/base_config.py, then ensure ENABLE_GET_WORDCLOUD is set to True before running the crawler.
MediaCrawler generates word-cloud visualizations from crawled comment data, but accurate Chinese text segmentation and proper character rendering require specific configuration. By customizing the vocabulary mapping and font path, you can ensure domain-specific phrases remain intact and Chinese characters display correctly in the final PNG output.
How Word Cloud Generation Works in MediaCrawler
The word-cloud pipeline is orchestrated by AsyncWordCloudGenerator in tools/words.py. When comment collection completes, the generator processes text through the following stages:
- Custom word injection – The generator reads
CUSTOM_WORDSfromconfig/base_config.pyand registers each phrase with jieba viajieba.add_word(word), forcing the segmenter to treat multi-character terms as single tokens. - Text segmentation – Comments are split using jieba, filtered against the stop-word list (
STOP_WORDS_FILE), and counted using aCounterobject. - Frequency serialization – Word frequencies are saved as JSON to
data/<platform>/words/<filename>_word_freq.json. - Image rendering – A
WordCloudinstance is created withfont_path=config.FONT_PATH, generating a PNG visualization saved as<filename>_word_cloud.png.
# tools/words.py – core implementation
self.custom_words = config.CUSTOM_WORDS
for word, group in self.custom_words.items():
jieba.add_word(word) # Preserve phrases as single tokens
wordcloud = WordCloud(
font_path=config.FONT_PATH, # Critical for Chinese rendering
width=800,
height=400,
background_color='white',
max_words=200,
stopwords=self.stop_words,
colormap='viridis'
).generate_from_frequencies(top_20_word_freq)
Configuring Custom Words and Chinese Fonts
Setting Up CUSTOM_WORDS in base_config.py
Domain-specific terminology often gets split incorrectly by default segmentation. Use CUSTOM_WORDS to map phrases to grouping labels, which simultaneously adds them to jieba's dictionary.
Edit config/base_config.py:
# config/base_config.py
CUSTOM_WORDS = {
"零几": "年份", # Treats "零几" as one token under "年份" group
"高频词": "专业术语",
"机器学习": "技术", # Your custom domain phrase here
"深度学习": "技术",
}
Each key becomes an indivisible token during segmentation, improving accuracy for technical jargon, brand names, or slang.
Configuring FONT_PATH for Chinese Character Rendering
The WordCloud library requires a TrueType font supporting Chinese Unicode glyphs. MediaCrawler provides a default font at docs/STZHONGS.TTF, but you can substitute any compatible .ttf file.
Update config/base_config.py:
# config/base_config.py
FONT_PATH = "./docs/STZHONGS.TTF" # Absolute or project-relative path
Place your chosen font file in the docs/ directory or reference an absolute system path. Without this configuration, Chinese characters render as rectangular placeholders (tofu).
Running the Crawler with Word Cloud Enabled
Word-cloud generation is disabled by default. Set ENABLE_GET_WORDCLOUD to True to trigger the pipeline after comment aggregation:
# Option 1: Environment variable
export ENABLE_GET_WORDCLOUD=True
python -m main
# Option 2: Direct config edit
# In config/base_config.py: ENABLE_GET_WORDCLOUD = True
After execution completes, check the output directory:
data/<platform>/words/
├── <crawler_type>_comments_2024-08-15_word_freq.json
└── <crawler_type>_comments_2024-08-15_word_cloud.png
Programmatic Usage Without Full Crawling
You can invoke the word-cloud generator independently for existing comment data. This script demonstrates the API using custom configuration values:
# example_generate_wordcloud.py
import asyncio
from tools.words import AsyncWordCloudGenerator
from config import CUSTOM_WORDS, FONT_PATH, STOP_WORDS_FILE
# Mock comment data – normally extracted by the crawler
comments = [
{"content": "机器学习 是 当下 最 热门 的 技术"},
{"content": "我 在 学习 机器学习 和 深度学习"},
{"content": "零几 年 的 发展 很 快"},
]
async def main():
generator = AsyncWordCloudGenerator()
prefix = "example/2024-08-15_demo"
await generator.generate_word_frequency_and_cloud(comments, prefix)
print(f"Files written: {prefix}_word_freq.json and {prefix}_word_cloud.png")
if __name__ == "__main__":
asyncio.run(main())
This produces the frequency JSON and PNG image while respecting your CUSTOM_WORDS mappings and FONT_PATH settings.
Summary
- Custom vocabulary is defined in
config/base_config.pyviaCUSTOM_WORDS, which injects phrases into jieba usingjieba.add_word()intools/words.py. - Chinese font rendering requires setting
FONT_PATHto a Unicode-compatible.ttffile (default:docs/STZHONGS.TTF). - Stop-word filtering uses the file specified by
STOP_WORDS_FILE(default:docs/hit_stopwords.txt). - Enable generation by setting
ENABLE_GET_WORDCLOUD=Truebefore runningpython -m main. - Output files include a JSON frequency table and a PNG word cloud saved to
data/<platform>/words/.
Frequently Asked Questions
Why do Chinese characters appear as boxes in my word cloud?
This occurs when the WordCloud instance cannot locate a valid Chinese TrueType font. Ensure FONT_PATH in config/base_config.py points to a valid .ttf file containing Chinese glyphs, such as the provided docs/STZHONGS.TTF or system fonts like /System/Library/Fonts/PingFang.ttc on macOS.
How do I prevent specific phrases from being split into separate words?
Add the phrase to the CUSTOM_WORDS dictionary in config/base_config.py. During initialization, AsyncWordCloudGenerator calls jieba.add_word(word) for each entry, ensuring the segmenter treats the entire phrase as a single token regardless of default dictionary rules.
Can I use the word-cloud generator without running the full MediaCrawler?
Yes. Import AsyncWordCloudGenerator from tools/words.py and call generate_word_frequency_and_cloud(comments, prefix) with a list of comment dictionaries. The method handles segmentation, frequency calculation, and image rendering using the same configuration values defined in base_config.py.
Where are the generated word-cloud files saved?
MediaCrawler writes output to data/<platform>/words/, creating two files: a *_word_freq.json containing the raw frequency data and a *_word_cloud.png containing the visualization. The exact filename incorporates the crawler type and current date.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →