How to Generate a Word Cloud from Crawled Comments Using MediaCrawler

To generate a word cloud from crawled comments, run uv run tools/words.py after collecting comments with MediaCrawler's platform-specific crawler—the script tokenizes Chinese text with jieba, filters stopwords, and renders the visualization using the wordcloud library.

MediaCrawler is an open-source scraping framework that collects content from platforms like XiaoHongShu (XHS), Douyin, Kuaishou, Bilibili, Weibo, Baidu Tieba, and Zhihu. One of its built-in utilities transforms raw comment data into visual word clouds, making it easy to identify trending topics and sentiment patterns. This guide walks through the complete workflow from data collection to image generation.

Prerequisites and Architecture Overview

The word cloud generation pipeline consists of three decoupled layers. Understanding this architecture helps you customize each stage independently.

Layer Component Responsibility
Crawler Platform-specific store implementations (store/xhs/_store_impl.py, store/douyin/_store_impl.py, etc.) Fetch posts and comment threads from target platforms
Storage CSV/JSON/SQLite/MySQL/MongoDB adapters Persist comment text in your preferred format
Visualization tools/words.py Tokenize, filter, and render the word cloud image

The storage format is agnostic to the visualization layer. The word cloud tool reads only the comment text field, regardless of whether you exported to CSV or inserted into a database.

Step 1: Crawl Comments from Your Target Platform

Before generating a word cloud, you need collected comment data. MediaCrawler supports multiple login types and search strategies.


# Example: Crawl XiaoHongShu comments using QR code authentication

uv run main.py --platform xhs --lt qrcode --type search

# Example: Crawl Douyin comment threads

uv run main.py --platform dy --lt cookie --type detail

# Example: Crawl Zhihu answers and their comments

uv run main.py --platform zhihu --lt qrcode --type search

Comments are automatically saved based on your SAVE_DATA_OPTION environment variable. The default location for CSV output is data/ relative to your project root.

Step 2: Configure Word Cloud Parameters

MediaCrawler's word cloud configuration is documented in docs/词云图使用配置.md. This file controls:

  • Stopwords: Chinese particles and common words to exclude (的, 了, 在, 是, 我, etc.)
  • Font path: System font for Chinese character rendering (default: SimHei.ttf)
  • Mask image: Optional shape template (PNG with white background)
  • Visual styling: Background color, max words, width, height, color scheme

Create or edit your configuration file to match your analysis needs:


# docs/词云图使用配囨.md (example structure)

## 停用词配置

stopwords:
  - 的
  - 了
  - 在
  - 是
  - 我
  - 你
  - 有
  - 和
  - 就
  - 不

## 视觉配置

font_path: fonts/SimHei.ttf
background_color: white
max_words: 200
width: 800
height: 600
colormap: viridis

Step 3: Generate the Word Cloud Image

The tools/words.py script orchestrates the entire visualization pipeline: loading comments, segmenting text with jieba, filtering stopwords, computing frequencies, and rendering with wordcloud.


# Basic usage with default configuration

uv run tools/words.py \
    --input data/comments_xhs.csv \
    --output output/wordcloud.png

# With custom configuration file

uv run tools/words.py \
    --input data/comments_xhs.csv \
    --output output/custom_wordcloud.png \
    --config docs/词云图使用配置.md

Understanding the Core Implementation

The tools/words.py script follows this processing flow as implemented in the MediaCrawler repository:

import argparse
import pandas as pd
import jieba
from wordcloud import WordCloud


def load_comments(path: str) -> list[str]:
    """Load comment text from CSV storage."""
    df = pd.read_csv(path)
    return df["comment"].dropna().tolist()


def build_frequency_dict(
    comments: list[str], 
    stopwords: set[str]
) -> dict[str, int]:
    """
    Tokenize Chinese comments with jieba and compute word frequencies.
    Filters out stopwords and whitespace-only tokens.
    """
    freq: dict[str, int] = {}
    
    for text in comments:
        for token in jieba.cut(text):
            word = token.strip()
            if word and word not in stopwords:
                freq[word] = freq.get(word, 0) + 1
    
    return freq


def create_wordcloud(
    frequencies: dict[str, int],
    config: dict
) -> WordCloud:
    """Initialize WordCloud with configuration parameters."""
    return WordCloud(
        font_path=config.get("font_path", "fonts/SimHei.ttf"),
        background_color=config.get("background_color", "white"),
        max_words=config.get("max_words", 200),
        width=config.get("width", 800),
        height=config.get("height", 600),
        colormap=config.get("colormap", "viridis"),
        mask=config.get("mask_image")  # Optional numpy array

    )


def main():
    parser = argparse.ArgumentParser(
        description="Generate word cloud from crawled comments"
    )
    parser.add_argument(
        "--input", 
        required=True, 
        help="Path to CSV file containing comments"
    )
    parser.add_argument(
        "--output", 
        required=True, 
        help="Output path for word cloud image"
    )
    parser.add_argument(
        "--config",
        default="docs/词云图使用配置.md",
        help="Path to configuration markdown file"
    )
    args = parser.parse_args()
    
    # Parse configuration (simplified—actual implementation parses YAML frontmatter)

    config = parse_config(args.config)
    stopwords = load_stopwords(config)
    
    # Process comments

    comments = load_comments(args.input)
    frequencies = build_frequency_dict(comments, stopwords)
    
    # Generate and save visualization

    wc = create_wordcloud(frequencies, config)
    wc.generate_from_frequencies(frequencies)
    wc.to_file(args.output)
    
    print(f"Word cloud saved to {args.output}")


if __name__ == "__main__":
    main()

Advanced Customization Options

Using a Custom Mask Shape

For branded or thematic visualizations, specify a mask image in your configuration:

import numpy as np
from PIL import Image

# Inside create_wordcloud()

mask_image_path = config.get("mask_path")
if mask_image_path:
    mask = np.array(Image.open(mask_image_path))
    wc_config["mask"] = mask

The mask image must be a PNG with white background—white areas will be excluded from the word placement.

Processing Database-Stored Comments

If you configured SAVE_DATA_OPTION to use SQLite or MySQL, modify the loader function:

import sqlite3

def load_comments_from_db(db_path: str, table: str = "comments") -> list[str]:
    conn = sqlite3.connect(db_path)
    cursor = conn.execute(f"SELECT content FROM {table}")
    comments = [row[0] for row in cursor.fetchall() if row[0]]
    conn.close()
    return comments

Then invoke with database parameters:

uv run tools/words.py \
    --input data/crawler.db \
    --output output/wordcloud.png \
    --source-type sqlite

Troubleshooting Common Issues

  • Missing Chinese characters: Ensure font_path points to a valid Chinese TrueType font (SimHei, Microsoft YaHei, or WenQuanYi Micro Hei)
  • Empty word cloud: Verify your CSV has a comment column; check that stopwords haven't excluded all meaningful tokens
  • jieba tokenization errors: Update jieba with uv pip install -U jieba and consider adding a custom dictionary for domain-specific terms
  • Memory errors with large datasets: Pre-filter comments by date or sample randomly before processing

Summary

  • MediaCrawler's three-layer architecture separates crawling, storage, and visualization concerns
  • Comments are collected via platform-specific crawlers in store/ and saved to your chosen format
  • tools/words.py performs Chinese word segmentation with jieba and renders via the wordcloud library
  • Configuration in docs/词云图使用配置.md controls stopwords, fonts, colors, and mask shapes
  • The pipeline accepts CSV, JSON, or database sources without code changes to the visualization layer

Frequently Asked Questions

What file formats does the word cloud generator support for input?

The default implementation reads CSV files with a comment column. You can extend tools/words.py to support JSON, Excel, or database sources by modifying the load_comments() function—the visualization logic remains unchanged regardless of input format.

Can I use the word cloud tool with comments I crawled previously?

Yes. As long as your historical data contains comment text in a parseable structure, point --input to that file. The tool has no dependency on when the data was collected, only on the schema of the comment field.

How do I add domain-specific stopwords for my industry?

Edit docs/词云图使用配置.md and append terms to the stopwords list. For programmatic control, you can also pass a custom stopwords file via --stopwords-path if you modify tools/words.py to accept this parameter.

Why are some high-frequency words not appearing in my word cloud?

Check three things: (1) the word may be in your stopwords list, (2) jieba may be segmenting it differently than expected—test with jieba.cut() directly, or (3) max_words may be truncating the result—increase this value in your configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →