# Adding Custom Words and Chinese Fonts for MediaCrawler Word Clouds: A Complete Guide

> Easily add custom words and Chinese fonts to MediaCrawler word clouds. Configure CUSTOM_WORDS and FONT_PATH in base_config.py, enable word clouds, and run the crawler for stunning visualizations.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**To add custom words and Chinese fonts in MediaCrawler, define `CUSTOM_WORDS` and `FONT_PATH` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), then ensure `ENABLE_GET_WORDCLOUD` is set to `True` before running the crawler.**

MediaCrawler generates word-cloud visualizations from crawled comment data, but accurate Chinese text segmentation and proper character rendering require specific configuration. By customizing the vocabulary mapping and font path, you can ensure domain-specific phrases remain intact and Chinese characters display correctly in the final PNG output.

## How Word Cloud Generation Works in MediaCrawler

The word-cloud pipeline is orchestrated by **`AsyncWordCloudGenerator`** in [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py). When comment collection completes, the generator processes text through the following stages:

1. **Custom word injection** – The generator reads `CUSTOM_WORDS` from [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and registers each phrase with **jieba** via `jieba.add_word(word)`, forcing the segmenter to treat multi-character terms as single tokens.
2. **Text segmentation** – Comments are split using jieba, filtered against the stop-word list (`STOP_WORDS_FILE`), and counted using a `Counter` object.
3. **Frequency serialization** – Word frequencies are saved as JSON to `data/<platform>/words/<filename>_word_freq.json`.
4. **Image rendering** – A `WordCloud` instance is created with `font_path=config.FONT_PATH`, generating a PNG visualization saved as `<filename>_word_cloud.png`.

```python

# tools/words.py – core implementation

self.custom_words = config.CUSTOM_WORDS
for word, group in self.custom_words.items():
    jieba.add_word(word)  # Preserve phrases as single tokens

wordcloud = WordCloud(
    font_path=config.FONT_PATH,  # Critical for Chinese rendering

    width=800,
    height=400,
    background_color='white',
    max_words=200,
    stopwords=self.stop_words,
    colormap='viridis'
).generate_from_frequencies(top_20_word_freq)

```

## Configuring Custom Words and Chinese Fonts

### Setting Up CUSTOM_WORDS in base_config.py

Domain-specific terminology often gets split incorrectly by default segmentation. Use **`CUSTOM_WORDS`** to map phrases to grouping labels, which simultaneously adds them to jieba's dictionary.

Edit [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py):

```python

# config/base_config.py

CUSTOM_WORDS = {
    "零几": "年份",           # Treats "零几" as one token under "年份" group

    "高频词": "专业术语",
    "机器学习": "技术",        # Your custom domain phrase here

    "深度学习": "技术",
}

```

Each key becomes an indivisible token during segmentation, improving accuracy for technical jargon, brand names, or slang.

### Configuring FONT_PATH for Chinese Character Rendering

The **WordCloud** library requires a TrueType font supporting Chinese Unicode glyphs. MediaCrawler provides a default font at `docs/STZHONGS.TTF`, but you can substitute any compatible `.ttf` file.

Update [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py):

```python

# config/base_config.py

FONT_PATH = "./docs/STZHONGS.TTF"  # Absolute or project-relative path

```

Place your chosen font file in the `docs/` directory or reference an absolute system path. Without this configuration, Chinese characters render as rectangular placeholders (tofu).

## Running the Crawler with Word Cloud Enabled

Word-cloud generation is disabled by default. Set **`ENABLE_GET_WORDCLOUD`** to `True` to trigger the pipeline after comment aggregation:

```bash

# Option 1: Environment variable

export ENABLE_GET_WORDCLOUD=True
python -m main

# Option 2: Direct config edit

# In config/base_config.py: ENABLE_GET_WORDCLOUD = True

```

After execution completes, check the output directory:

```

data/<platform>/words/
├── <crawler_type>_comments_2024-08-15_word_freq.json
└── <crawler_type>_comments_2024-08-15_word_cloud.png

```

## Programmatic Usage Without Full Crawling

You can invoke the word-cloud generator independently for existing comment data. This script demonstrates the API using custom configuration values:

```python

# example_generate_wordcloud.py

import asyncio
from tools.words import AsyncWordCloudGenerator
from config import CUSTOM_WORDS, FONT_PATH, STOP_WORDS_FILE

# Mock comment data – normally extracted by the crawler

comments = [
    {"content": "机器学习 是 当下 最 热门 的 技术"},
    {"content": "我 在 学习 机器学习 和 深度学习"},
    {"content": "零几 年 的 发展 很 快"},
]

async def main():
    generator = AsyncWordCloudGenerator()
    prefix = "example/2024-08-15_demo"
    await generator.generate_word_frequency_and_cloud(comments, prefix)
    print(f"Files written: {prefix}_word_freq.json and {prefix}_word_cloud.png")

if __name__ == "__main__":
    asyncio.run(main())

```

This produces the frequency JSON and PNG image while respecting your `CUSTOM_WORDS` mappings and `FONT_PATH` settings.

## Summary

- **Custom vocabulary** is defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) via `CUSTOM_WORDS`, which injects phrases into jieba using `jieba.add_word()` in [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py).
- **Chinese font rendering** requires setting `FONT_PATH` to a Unicode-compatible `.ttf` file (default: `docs/STZHONGS.TTF`).
- **Stop-word filtering** uses the file specified by `STOP_WORDS_FILE` (default: [`docs/hit_stopwords.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/docs/hit_stopwords.txt)).
- **Enable generation** by setting `ENABLE_GET_WORDCLOUD=True` before running `python -m main`.
- Output files include a JSON frequency table and a PNG word cloud saved to `data/<platform>/words/`.

## Frequently Asked Questions

### Why do Chinese characters appear as boxes in my word cloud?

This occurs when the **WordCloud** instance cannot locate a valid Chinese TrueType font. Ensure `FONT_PATH` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) points to a valid `.ttf` file containing Chinese glyphs, such as the provided `docs/STZHONGS.TTF` or system fonts like `/System/Library/Fonts/PingFang.ttc` on macOS.

### How do I prevent specific phrases from being split into separate words?

Add the phrase to the **`CUSTOM_WORDS`** dictionary in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). During initialization, `AsyncWordCloudGenerator` calls `jieba.add_word(word)` for each entry, ensuring the segmenter treats the entire phrase as a single token regardless of default dictionary rules.

### Can I use the word-cloud generator without running the full MediaCrawler?

Yes. Import **`AsyncWordCloudGenerator`** from [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) and call `generate_word_frequency_and_cloud(comments, prefix)` with a list of comment dictionaries. The method handles segmentation, frequency calculation, and image rendering using the same configuration values defined in [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py).

### Where are the generated word-cloud files saved?

MediaCrawler writes output to `data/<platform>/words/`, creating two files: a `*_word_freq.json` containing the raw frequency data and a `*_word_cloud.png` containing the visualization. The exact filename incorporates the crawler type and current date.