Adding Custom Words and Chinese Fonts for MediaCrawler Word Clouds: A Complete Guide

To add custom words and Chinese fonts in MediaCrawler, define CUSTOM_WORDS and FONT_PATH in config/base_config.py, then ensure ENABLE_GET_WORDCLOUD is set to True before running the crawler.

MediaCrawler generates word-cloud visualizations from crawled comment data, but accurate Chinese text segmentation and proper character rendering require specific configuration. By customizing the vocabulary mapping and font path, you can ensure domain-specific phrases remain intact and Chinese characters display correctly in the final PNG output.

How Word Cloud Generation Works in MediaCrawler

The word-cloud pipeline is orchestrated by AsyncWordCloudGenerator in tools/words.py. When comment collection completes, the generator processes text through the following stages:

  1. Custom word injection – The generator reads CUSTOM_WORDS from config/base_config.py and registers each phrase with jieba via jieba.add_word(word), forcing the segmenter to treat multi-character terms as single tokens.
  2. Text segmentation – Comments are split using jieba, filtered against the stop-word list (STOP_WORDS_FILE), and counted using a Counter object.
  3. Frequency serialization – Word frequencies are saved as JSON to data/<platform>/words/<filename>_word_freq.json.
  4. Image rendering – A WordCloud instance is created with font_path=config.FONT_PATH, generating a PNG visualization saved as <filename>_word_cloud.png.

# tools/words.py – core implementation

self.custom_words = config.CUSTOM_WORDS
for word, group in self.custom_words.items():
    jieba.add_word(word)  # Preserve phrases as single tokens

wordcloud = WordCloud(
    font_path=config.FONT_PATH,  # Critical for Chinese rendering

    width=800,
    height=400,
    background_color='white',
    max_words=200,
    stopwords=self.stop_words,
    colormap='viridis'
).generate_from_frequencies(top_20_word_freq)

Configuring Custom Words and Chinese Fonts

Setting Up CUSTOM_WORDS in base_config.py

Domain-specific terminology often gets split incorrectly by default segmentation. Use CUSTOM_WORDS to map phrases to grouping labels, which simultaneously adds them to jieba's dictionary.

Edit config/base_config.py:


# config/base_config.py

CUSTOM_WORDS = {
    "零几": "年份",           # Treats "零几" as one token under "年份" group

    "高频词": "专业术语",
    "机器学习": "技术",        # Your custom domain phrase here

    "深度学习": "技术",
}

Each key becomes an indivisible token during segmentation, improving accuracy for technical jargon, brand names, or slang.

Configuring FONT_PATH for Chinese Character Rendering

The WordCloud library requires a TrueType font supporting Chinese Unicode glyphs. MediaCrawler provides a default font at docs/STZHONGS.TTF, but you can substitute any compatible .ttf file.

Update config/base_config.py:


# config/base_config.py

FONT_PATH = "./docs/STZHONGS.TTF"  # Absolute or project-relative path

Place your chosen font file in the docs/ directory or reference an absolute system path. Without this configuration, Chinese characters render as rectangular placeholders (tofu).

Running the Crawler with Word Cloud Enabled

Word-cloud generation is disabled by default. Set ENABLE_GET_WORDCLOUD to True to trigger the pipeline after comment aggregation:


# Option 1: Environment variable

export ENABLE_GET_WORDCLOUD=True
python -m main

# Option 2: Direct config edit

# In config/base_config.py: ENABLE_GET_WORDCLOUD = True

After execution completes, check the output directory:


data/<platform>/words/
├── <crawler_type>_comments_2024-08-15_word_freq.json
└── <crawler_type>_comments_2024-08-15_word_cloud.png

Programmatic Usage Without Full Crawling

You can invoke the word-cloud generator independently for existing comment data. This script demonstrates the API using custom configuration values:


# example_generate_wordcloud.py

import asyncio
from tools.words import AsyncWordCloudGenerator
from config import CUSTOM_WORDS, FONT_PATH, STOP_WORDS_FILE

# Mock comment data – normally extracted by the crawler

comments = [
    {"content": "机器学习 是 当下 最 热门 的 技术"},
    {"content": "我 在 学习 机器学习 和 深度学习"},
    {"content": "零几 年 的 发展 很 快"},
]

async def main():
    generator = AsyncWordCloudGenerator()
    prefix = "example/2024-08-15_demo"
    await generator.generate_word_frequency_and_cloud(comments, prefix)
    print(f"Files written: {prefix}_word_freq.json and {prefix}_word_cloud.png")

if __name__ == "__main__":
    asyncio.run(main())

This produces the frequency JSON and PNG image while respecting your CUSTOM_WORDS mappings and FONT_PATH settings.

Summary

  • Custom vocabulary is defined in config/base_config.py via CUSTOM_WORDS, which injects phrases into jieba using jieba.add_word() in tools/words.py.
  • Chinese font rendering requires setting FONT_PATH to a Unicode-compatible .ttf file (default: docs/STZHONGS.TTF).
  • Stop-word filtering uses the file specified by STOP_WORDS_FILE (default: docs/hit_stopwords.txt).
  • Enable generation by setting ENABLE_GET_WORDCLOUD=True before running python -m main.
  • Output files include a JSON frequency table and a PNG word cloud saved to data/<platform>/words/.

Frequently Asked Questions

Why do Chinese characters appear as boxes in my word cloud?

This occurs when the WordCloud instance cannot locate a valid Chinese TrueType font. Ensure FONT_PATH in config/base_config.py points to a valid .ttf file containing Chinese glyphs, such as the provided docs/STZHONGS.TTF or system fonts like /System/Library/Fonts/PingFang.ttc on macOS.

How do I prevent specific phrases from being split into separate words?

Add the phrase to the CUSTOM_WORDS dictionary in config/base_config.py. During initialization, AsyncWordCloudGenerator calls jieba.add_word(word) for each entry, ensuring the segmenter treats the entire phrase as a single token regardless of default dictionary rules.

Can I use the word-cloud generator without running the full MediaCrawler?

Yes. Import AsyncWordCloudGenerator from tools/words.py and call generate_word_frequency_and_cloud(comments, prefix) with a list of comment dictionaries. The method handles segmentation, frequency calculation, and image rendering using the same configuration values defined in base_config.py.

Where are the generated word-cloud files saved?

MediaCrawler writes output to data/<platform>/words/, creating two files: a *_word_freq.json containing the raw frequency data and a *_word_cloud.png containing the visualization. The exact filename incorporates the crawler type and current date.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →