# How to Generate Word Clouds from Scraped Data Using MediaCrawler

> Easily generate word clouds from scraped data with MediaCrawler. Follow simple steps to enable comment and word cloud generation for instant data visualization.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-01

---

**Enable `ENABLE_GET_COMMENTS` and `ENABLE_GET_WORDCLOUD` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), then run the crawler to automatically produce frequency JSON files and PNG visualizations in `data/words/`.**

MediaCrawler is an open-source asynchronous scraping framework that extracts content from various social media platforms. When you need to visualize textual patterns from scraped comments, the built-in word cloud generator processes raw data into ready-to-use frequency maps and image files without requiring external tools.

## Prerequisites and Configuration

Before generating word clouds, you must ensure data persistence and comment extraction are properly configured. In [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), set `SAVE_DATA_OPTION = "jsonl"` (or `"json"`) so the pipeline can read the scraped content. You must also set `ENABLE_GET_COMMENTS = True` to fetch the textual data that feeds the visualization engine.

The word cloud feature itself is controlled by the `ENABLE_GET_WORDCLOUD` boolean flag. When this is set to `True`, the crawler automatically invokes the `AsyncWordCloudGenerator` class after each run.

## How the Word Cloud Pipeline Works

The implementation in [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) follows a three-stage asynchronous pipeline that transforms raw comments into visual outputs.

### Data Collection and Text Aggregation

The `AsyncWordCloudGenerator` class first aggregates all comment content from the scraped dataset. It extracts the `content` field from each item in the data array and concatenates them into a single string for processing.

### Tokenization and Frequency Analysis

The generator uses **jieba** to segment Chinese text into individual tokens. It filters these tokens against a stop-word list (default: [`./docs/hit_stopwords.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/./docs/hit_stopwords.txt)) and builds a frequency `Counter` from the remaining words. The class also respects the `CUSTOM_WORDS` dictionary defined in your configuration, adding domain-specific terms to jieba's dictionary during initialization so multi-word phrases are treated as single tokens.

### Image Rendering and Output

The top 20 words by frequency are passed to the **wordcloud** library's `WordCloud.generate_from_frequencies()` method. Matplotlib renders the visualization using the font specified in `FONT_PATH` (default: `./docs/STZHONGS.TTF`) and saves a PNG file alongside the raw frequency JSON. An `asyncio` lock (`plot_lock`) prevents resource conflicts when running multiple concurrent crawls.

Output files are written to `data/words/` with the following naming convention:

- `<crawl-id>_word_freq.json` – Raw word frequency data
- `<crawl-id>_word_cloud.png` – Rendered visualization

## Key Configuration Options

Tune these settings in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to control the output:

| Setting | Description |
|---------|-------------|
| `ENABLE_GET_COMMENTS` | Must be `True` to scrape comment text |
| `ENABLE_GET_WORDCLOUD` | Must be `True` to trigger generation |
| `SAVE_DATA_OPTION` | Must be `"jsonl"` or `"json"` for the reader to parse files |
| `CUSTOM_WORDS` | Dictionary mapping phrases to tags (e.g., `{"机器学习": "tech"}`) |
| `STOP_WORDS_FILE` | Path to line-separated stop-words file |
| `FONT_PATH` | Path to TrueType font for Chinese character rendering |

## Code Examples

### Running a Full Crawl with Word Cloud Output

This script executes the complete pipeline from configuration to visualization:

```python
import asyncio
import config
from main import run

async def main():
    # Configure options before importing other modules

    config.ENABLE_GET_COMMENTS = True
    config.ENABLE_GET_WORDCLOUD = True
    config.SAVE_DATA_OPTION = "jsonl"
    
    # Run the crawler (platform settings defined in base_config)

    await run()
    
    print("Word cloud files saved to data/words/")

if __name__ == "__main__":
    asyncio.run(main())

```

The `run()` function in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py) orchestrates the entire process, automatically invoking `AsyncWordCloudGenerator` when the flags are enabled.

### Generating Word Clouds from Existing Data

To re-process existing JSONL files without re-crawling:

```python
import asyncio
import json
import aiofiles
from tools.words import AsyncWordCloudGenerator

async def gen_from_file(jsonl_path: str, prefix: str):
    # Load previously saved comments

    async with aiofiles.open(jsonl_path, 'r', encoding='utf-8') as f:
        lines = await f.readlines()
        data = [json.loads(line) for line in lines]
    
    generator = AsyncWordCloudGenerator()
    await generator.generate_word_frequency_and_cloud(data, prefix)

if __name__ == "__main__":
    asyncio.run(gen_from_file(
        "data/comments/example.jsonl",
        "data/words/example"
    ))

```

### Customizing Stop Words and Tokenization

Add domain-specific vocabulary to improve tokenization accuracy:

```python

# config/base_config.py

CUSTOM_WORDS = {
    "人工智能": "tech",
    "2025": "year",
    "深度学习": "tech"
}

```

Update your stop-words list in [`docs/hit_stopwords.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/docs/hit_stopwords.txt) (one word per line):

```

的
了
和

```

The generator calls `jieba.add_word()` during initialization, ensuring these custom terms are recognized in subsequent runs.

## Summary

- **Enable three flags** in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py): `ENABLE_GET_COMMENTS`, `ENABLE_GET_WORDCLOUD`, and set `SAVE_DATA_OPTION` to `"jsonl"`
- **Process flow**: `AsyncWordCloudGenerator` in [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) aggregates text, tokenizes with jieba, filters stop-words, and renders via matplotlib
- **Output location**: `data/words/` directory contains both JSON frequency data and PNG images
- **Customization**: Use `CUSTOM_WORDS` for phrase tokenization and `STOP_WORDS_FILE` to filter noise
- **Concurrency safe**: The `plot_lock` prevents matplotlib conflicts during async operations

## Frequently Asked Questions

### Where are the word cloud files saved?

MediaCrawler writes output to `data/words/<crawl-id>_word_freq.json` and `data/words/<crawl-id>_word_cloud.png`. The prefix is generated from your crawl parameters and timestamp.

### Can I generate word clouds from existing JSONL files without re-crawling?

Yes. Import `AsyncWordCloudGenerator` from [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) and call `generate_word_frequency_and_cloud()` with your loaded data array and desired file prefix. This bypasses the scraping stage entirely.

### How do I add custom words or phrases to the tokenization?

Define a `CUSTOM_WORDS` dictionary in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) where keys are phrases and values are tags. The generator passes these to `jieba.add_word()` during initialization, ensuring multi-character terms are treated as single tokens.

### Why is my word cloud showing incorrect characters or squares?

This indicates a font issue. Verify that `FONT_PATH` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) points to a valid Chinese TrueType font (default is `./docs/STZHONGS.TTF`). The system must have read access to this file, and the font must support the character set in your scraped data.