# How to Generate a Word Cloud from Crawled Comments Using MediaCrawler

> Learn to generate a word cloud from crawled comments using MediaCrawler. This guide covers tokenization, stopword filtering, and visualization with Python's wordcloud library.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: tutorial
- Published: 2026-08-14

---

**To generate a word cloud from crawled comments, run `uv run tools/words.py` after collecting comments with MediaCrawler's platform-specific crawler—the script tokenizes Chinese text with jieba, filters stopwords, and renders the visualization using the wordcloud library.**

MediaCrawler is an open-source scraping framework that collects content from platforms like XiaoHongShu (XHS), Douyin, Kuaishou, Bilibili, Weibo, Baidu Tieba, and Zhihu. One of its built-in utilities transforms raw comment data into visual word clouds, making it easy to identify trending topics and sentiment patterns. This guide walks through the complete workflow from data collection to image generation.

## Prerequisites and Architecture Overview

The word cloud generation pipeline consists of three decoupled layers. Understanding this architecture helps you customize each stage independently.

| Layer | Component | Responsibility |
|-------|-----------|---------------|
| Crawler | Platform-specific store implementations ([`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py), [`store/douyin/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/_store_impl.py), etc.) | Fetch posts and comment threads from target platforms |
| Storage | CSV/JSON/SQLite/MySQL/MongoDB adapters | Persist comment text in your preferred format |
| Visualization | [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) | Tokenize, filter, and render the word cloud image |

The storage format is **agnostic** to the visualization layer. The word cloud tool reads only the comment text field, regardless of whether you exported to CSV or inserted into a database.

## Step 1: Crawl Comments from Your Target Platform

Before generating a word cloud, you need collected comment data. MediaCrawler supports multiple login types and search strategies.

```bash

# Example: Crawl XiaoHongShu comments using QR code authentication

uv run main.py --platform xhs --lt qrcode --type search

# Example: Crawl Douyin comment threads

uv run main.py --platform dy --lt cookie --type detail

# Example: Crawl Zhihu answers and their comments

uv run main.py --platform zhihu --lt qrcode --type search

```

Comments are automatically saved based on your `SAVE_DATA_OPTION` environment variable. The default location for CSV output is `data/` relative to your project root.

## Step 2: Configure Word Cloud Parameters

MediaCrawler's word cloud configuration is documented in `docs/词云图使用配置.md`. This file controls:

- **Stopwords**: Chinese particles and common words to exclude (的, 了, 在, 是, 我, etc.)
- **Font path**: System font for Chinese character rendering (default: SimHei.ttf)
- **Mask image**: Optional shape template (PNG with white background)
- **Visual styling**: Background color, max words, width, height, color scheme

Create or edit your configuration file to match your analysis needs:

```markdown

# docs/词云图使用配囨.md (example structure)

## 停用词配置

stopwords:
  - 的
  - 了
  - 在
  - 是
  - 我
  - 你
  - 有
  - 和
  - 就
  - 不

## 视觉配置

font_path: fonts/SimHei.ttf
background_color: white
max_words: 200
width: 800
height: 600
colormap: viridis

```

## Step 3: Generate the Word Cloud Image

The [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) script orchestrates the entire visualization pipeline: loading comments, segmenting text with **jieba**, filtering stopwords, computing frequencies, and rendering with **wordcloud**.

```bash

# Basic usage with default configuration

uv run tools/words.py \
    --input data/comments_xhs.csv \
    --output output/wordcloud.png

# With custom configuration file

uv run tools/words.py \
    --input data/comments_xhs.csv \
    --output output/custom_wordcloud.png \
    --config docs/词云图使用配置.md

```

### Understanding the Core Implementation

The [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) script follows this processing flow as implemented in the [MediaCrawler repository](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py):

```python
import argparse
import pandas as pd
import jieba
from wordcloud import WordCloud


def load_comments(path: str) -> list[str]:
    """Load comment text from CSV storage."""
    df = pd.read_csv(path)
    return df["comment"].dropna().tolist()


def build_frequency_dict(
    comments: list[str], 
    stopwords: set[str]
) -> dict[str, int]:
    """
    Tokenize Chinese comments with jieba and compute word frequencies.
    Filters out stopwords and whitespace-only tokens.
    """
    freq: dict[str, int] = {}
    
    for text in comments:
        for token in jieba.cut(text):
            word = token.strip()
            if word and word not in stopwords:
                freq[word] = freq.get(word, 0) + 1
    
    return freq


def create_wordcloud(
    frequencies: dict[str, int],
    config: dict
) -> WordCloud:
    """Initialize WordCloud with configuration parameters."""
    return WordCloud(
        font_path=config.get("font_path", "fonts/SimHei.ttf"),
        background_color=config.get("background_color", "white"),
        max_words=config.get("max_words", 200),
        width=config.get("width", 800),
        height=config.get("height", 600),
        colormap=config.get("colormap", "viridis"),
        mask=config.get("mask_image")  # Optional numpy array

    )


def main():
    parser = argparse.ArgumentParser(
        description="Generate word cloud from crawled comments"
    )
    parser.add_argument(
        "--input", 
        required=True, 
        help="Path to CSV file containing comments"
    )
    parser.add_argument(
        "--output", 
        required=True, 
        help="Output path for word cloud image"
    )
    parser.add_argument(
        "--config",
        default="docs/词云图使用配置.md",
        help="Path to configuration markdown file"
    )
    args = parser.parse_args()
    
    # Parse configuration (simplified—actual implementation parses YAML frontmatter)

    config = parse_config(args.config)
    stopwords = load_stopwords(config)
    
    # Process comments

    comments = load_comments(args.input)
    frequencies = build_frequency_dict(comments, stopwords)
    
    # Generate and save visualization

    wc = create_wordcloud(frequencies, config)
    wc.generate_from_frequencies(frequencies)
    wc.to_file(args.output)
    
    print(f"Word cloud saved to {args.output}")


if __name__ == "__main__":
    main()

```

## Advanced Customization Options

### Using a Custom Mask Shape

For branded or thematic visualizations, specify a mask image in your configuration:

```python
import numpy as np
from PIL import Image

# Inside create_wordcloud()

mask_image_path = config.get("mask_path")
if mask_image_path:
    mask = np.array(Image.open(mask_image_path))
    wc_config["mask"] = mask

```

The mask image must be a PNG with white background—white areas will be excluded from the word placement.

### Processing Database-Stored Comments

If you configured `SAVE_DATA_OPTION` to use SQLite or MySQL, modify the loader function:

```python
import sqlite3

def load_comments_from_db(db_path: str, table: str = "comments") -> list[str]:
    conn = sqlite3.connect(db_path)
    cursor = conn.execute(f"SELECT content FROM {table}")
    comments = [row[0] for row in cursor.fetchall() if row[0]]
    conn.close()
    return comments

```

Then invoke with database parameters:

```bash
uv run tools/words.py \
    --input data/crawler.db \
    --output output/wordcloud.png \
    --source-type sqlite

```

## Troubleshooting Common Issues

- **Missing Chinese characters**: Ensure `font_path` points to a valid Chinese TrueType font (SimHei, Microsoft YaHei, or WenQuanYi Micro Hei)
- **Empty word cloud**: Verify your CSV has a `comment` column; check that stopwords haven't excluded all meaningful tokens
- **jieba tokenization errors**: Update jieba with `uv pip install -U jieba` and consider adding a custom dictionary for domain-specific terms
- **Memory errors with large datasets**: Pre-filter comments by date or sample randomly before processing

## Summary

- MediaCrawler's **three-layer architecture** separates crawling, storage, and visualization concerns
- Comments are collected via platform-specific crawlers in `store/` and saved to your chosen format
- [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) performs **Chinese word segmentation with jieba** and renders via the **wordcloud** library
- Configuration in `docs/词云图使用配置.md` controls stopwords, fonts, colors, and mask shapes
- The pipeline accepts CSV, JSON, or database sources without code changes to the visualization layer

## Frequently Asked Questions

### What file formats does the word cloud generator support for input?

The default implementation reads **CSV files** with a `comment` column. You can extend [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) to support JSON, Excel, or database sources by modifying the `load_comments()` function—the visualization logic remains unchanged regardless of input format.

### Can I use the word cloud tool with comments I crawled previously?

Yes. As long as your historical data contains comment text in a parseable structure, point `--input` to that file. The tool has no dependency on when the data was collected, only on the schema of the comment field.

### How do I add domain-specific stopwords for my industry?

Edit `docs/词云图使用配置.md` and append terms to the stopwords list. For programmatic control, you can also pass a custom stopwords file via `--stopwords-path` if you modify [`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py) to accept this parameter.

### Why are some high-frequency words not appearing in my word cloud?

Check three things: (1) the word may be in your stopwords list, (2) jieba may be segmenting it differently than expected—test with `jieba.cut()` directly, or (3) `max_words` may be truncating the result—increase this value in your configuration.