# How to Generate a Word Cloud from Comments in MediaCrawler with Custom and Stop Words

> Learn how to generate a word cloud from MediaCrawler comments. Easily customize word clouds using custom and stop words with AsyncWordCloudGenerator.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-06-29

---

**You can generate a word cloud from comments in MediaCrawler by enabling `ENABLE_GET_WORDCLOUD` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), then using the `AsyncWordCloudGenerator` class from [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) to process comment data with custom jieba dictionary entries and filtered stop words.**

MediaCrawler is a powerful open-source framework for scraping content from Chinese social media platforms like Xiaohongshu and Weibo. When analyzing scraped comments, visualizing term frequency through word clouds helps identify trending topics and user sentiment patterns. This guide explains how to generate a word cloud from comments in MediaCrawler with custom and stop words using the built-in asynchronous generator and configuration system.

## Understanding the Word Cloud Architecture

The word cloud functionality in MediaCrawler is built around a specialized asynchronous class that handles text processing, tokenization, and image generation.

### The AsyncWordCloudGenerator Class

The core implementation resides in **[`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py)**, which defines the `AsyncWordCloudGenerator` class. This class orchestrates the entire pipeline: loading stop words, registering custom dictionary entries with jieba, computing term frequencies, and rendering the final visualization. It uses `plot_lock` to ensure thread-safe rendering when multiple crawler instances run concurrently.

### Configuration Settings

All customizable parameters are centralized in **[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)**. The system relies on four key settings:

- **`ENABLE_GET_WORDCLOUD`**: Boolean toggle to activate the feature
- **`CUSTOM_WORDS`**: Dictionary mapping phrases to categories (e.g., `{'机器学习': '技术'}`)
- **`STOP_WORDS_FILE`**: Path to a plain-text file containing one stop word per line
- **`FONT_PATH`**: Path to a Chinese-compatible TTF font file (default: `./docs/STZHONGS.TTF`)

## Step-by-Step Implementation Process

The generator follows a specific pipeline to transform raw comment text into a visual word cloud.

### Initialization and Setup

When you instantiate `AsyncWordCloudGenerator()`, the `__init__` method performs two critical operations:

1. Loads the stop word list from `config.STOP_WORDS_FILE` into a Python `set` for O(1) lookup performance
2. Iterates through `config.CUSTOM_WORDS` and registers each phrase using `jieba.add_word(phrase)`, ensuring multi-character terms are treated as single tokens during segmentation

### Processing Raw Comment Data

The primary entry point is `generate_word_frequency_and_cloud(data, save_words_prefix)`. This method expects `data` to be a list of dictionaries, where each dictionary represents a comment and must contain a `content` key:

```python
comments = [
    {"content": "零几是一个年份词，值得关注", "author": "user1"},
    {"content": "高频词在技术文档里经常出现", "like_count": 5}
]

```

The method concatenates all `content` values into a single text string for batch processing.

### Tokenization and Filtering

The generator processes text using the following sequence:

1. **Tokenization**: Calls `jieba.lcut(all_text)` to segment Chinese text into individual tokens
2. **Filtering**: Removes any token that exists in the stop word set or contains only whitespace
3. **Frequency Analysis**: Uses `collections.Counter` to calculate occurrence statistics for each remaining token

### Generating Output Files

The method produces two artifacts:

- **JSON frequency file**: Named `<prefix>_word_freq.json`, containing the complete word frequency mapping written asynchronously using `aiofiles`
- **PNG word cloud**: Named `<prefix>_word_cloud.png`, generated using the `wordcloud` library with the top 20 most frequent terms

The image generation acquires `plot_lock` before rendering to prevent matplotlib threading conflicts, then releases it immediately after saving the file.

## Configuring Custom Words and Stop Words

Fine-tuning the word cloud requires modifying the configuration files to control exactly which terms appear in your visualization.

### Adding Custom Words

Edit **[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)** to define phrases that jieba should preserve as single tokens:

```python
CUSTOM_WORDS = {
    "机器学习": "技术领域",
    "零几": "年份",
    "深度学习": "技术领域",
    "高频词": "专业术语"
}

```

Each key is automatically registered via `jieba.add_word()` during generator initialization, preventing the tokenizer from splitting these phrases into separate characters.

### Managing Stop Words

Create or edit the file specified by `STOP_WORDS_FILE` (commonly [`./docs/hit_stopwords.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/./docs/hit_stopwords.txt)). Populate it with one word per line to exclude from analysis:

```

的
了
在
是
我
你

```

These words are filtered out immediately after tokenization, ensuring common particles and pronouns do not dominate your visualization.

### Font Configuration for Chinese Characters

Ensure the file specified by `FONT_PATH` exists and supports Chinese glyphs. The default `STZHONGS.TTF` provides proper rendering for Simplified Chinese characters. Without a valid Chinese font, the word cloud will display squares or gibberish instead of text.

## Practical Code Examples

### Standalone Usage

You can use the generator outside the main crawler pipeline for testing or batch processing existing comment data:

```python
import asyncio
from tools.words import AsyncWordCloudGenerator
import config

# Sample comment data structure

comments = [
    {"content": "零几是一个年份词，值得关注"},
    {"content": "高频词在技术文档里经常出现"},
    {"content": "普通词也会被统计"},
]

# Enable feature and configure

config.ENABLE_GET_WORDCLOUD = True
config.CUSTOM_WORDS = {"零几": "年份", "高频词": "术语"}

async def main():
    gen = AsyncWordCloudGenerator()
    await gen.generate_word_frequency_and_cloud(
        data=comments, 
        save_words_prefix="demo"
    )
    print("Files generated: demo_word_freq.json and demo_word_cloud.png")

asyncio.run(main())

```

### Integration Within the Crawler Pipeline

Inside platform-specific crawler implementations (typically after collecting comments in [`extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/extractor.py) or [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py)):

```python
from tools.words import AsyncWordCloudGenerator
import config
import asyncio

async def process_comments(comments, video_id):
    if config.ENABLE_GET_WORDCLOUD:
        gen = AsyncWordCloudGenerator()
        await gen.generate_word_frequency_and_cloud(
            data=comments,
            save_words_prefix=f"output/{video_id}"
        )

# Schedule during crawl completion

asyncio.create_task(process_comments(comments, video_id))

```

This ensures word clouds are generated asynchronously without blocking the main crawling logic.

## Summary

- **Enable the feature** by setting `ENABLE_GET_WORDCLOUD = True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)
- **Customize vocabulary** by adding entries to `CUSTOM_WORDS` to preserve multi-character phrases as single tokens
- **Filter noise** by populating the stop words file referenced by `STOP_WORDS_FILE` with terms to exclude
- **Use the generator** by importing `AsyncWordCloudGenerator` from [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) and calling `generate_word_frequency_and_cloud()` with your comment data
- **Ensure Chinese support** by verifying the font file at `FONT_PATH` exists and contains Chinese glyphs
- **Thread safety** is handled automatically via `plot_lock` when rendering images in concurrent crawling scenarios

## Frequently Asked Questions

### How do I prevent specific phrases from being split into separate words?

Add the phrases to the `CUSTOM_WORDS` dictionary in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). Each key is automatically registered with `jieba.add_word()` during generator initialization, forcing the tokenizer to treat the entire phrase as a single token. For example, adding `"机器学习": "tech"` ensures "机器学习" is counted as one term rather than separate characters.

### Where should I place my custom stop words file?

Specify the absolute or relative path in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) using the `STOP_WORDS_FILE` variable. The default configuration typically points to [`./docs/hit_stopwords.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/./docs/hit_stopwords.txt). The file should contain plain text with one stop word per line. The generator loads these into a set during initialization to efficiently filter them during tokenization.

### Why does my word cloud display squares or incorrect characters instead of Chinese text?

This indicates the font file specified by `FONT_PATH` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) is missing or lacks Chinese glyph support. Ensure the TTF file exists at the specified path and contains Chinese characters. The repository defaults to `./docs/STZHONGS.TTF`, which you must provide or replace with another Chinese-compatible font like SimHei or Source Han Sans.

### Can I generate word clouds from existing comment data without running the full crawler?

Yes. Import `AsyncWordCloudGenerator` from [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) directly into your own script. Prepare your data as a list of dictionaries with a `content` key, ensure `ENABLE_GET_WORDCLOUD` is set to `True` in your config, and call `generate_word_frequency_and_cloud()` asynchronously. This allows offline processing of previously scraped comment JSON files or data from other sources.