How to Generate a Word Cloud from Crawled Comments Using MediaCrawler
To generate a word cloud from crawled comments, run uv run tools/words.py after collecting comments with MediaCrawler's platform-specific crawler—the script tokenizes Chinese text with jieba, filters stopwords, and renders the visualization using the wordcloud library.
MediaCrawler is an open-source scraping framework that collects content from platforms like XiaoHongShu (XHS), Douyin, Kuaishou, Bilibili, Weibo, Baidu Tieba, and Zhihu. One of its built-in utilities transforms raw comment data into visual word clouds, making it easy to identify trending topics and sentiment patterns. This guide walks through the complete workflow from data collection to image generation.
Prerequisites and Architecture Overview
The word cloud generation pipeline consists of three decoupled layers. Understanding this architecture helps you customize each stage independently.
| Layer | Component | Responsibility |
|---|---|---|
| Crawler | Platform-specific store implementations (store/xhs/_store_impl.py, store/douyin/_store_impl.py, etc.) |
Fetch posts and comment threads from target platforms |
| Storage | CSV/JSON/SQLite/MySQL/MongoDB adapters | Persist comment text in your preferred format |
| Visualization | tools/words.py |
Tokenize, filter, and render the word cloud image |
The storage format is agnostic to the visualization layer. The word cloud tool reads only the comment text field, regardless of whether you exported to CSV or inserted into a database.
Step 1: Crawl Comments from Your Target Platform
Before generating a word cloud, you need collected comment data. MediaCrawler supports multiple login types and search strategies.
# Example: Crawl XiaoHongShu comments using QR code authentication
uv run main.py --platform xhs --lt qrcode --type search
# Example: Crawl Douyin comment threads
uv run main.py --platform dy --lt cookie --type detail
# Example: Crawl Zhihu answers and their comments
uv run main.py --platform zhihu --lt qrcode --type search
Comments are automatically saved based on your SAVE_DATA_OPTION environment variable. The default location for CSV output is data/ relative to your project root.
Step 2: Configure Word Cloud Parameters
MediaCrawler's word cloud configuration is documented in docs/词云图使用配置.md. This file controls:
- Stopwords: Chinese particles and common words to exclude (的, 了, 在, 是, 我, etc.)
- Font path: System font for Chinese character rendering (default: SimHei.ttf)
- Mask image: Optional shape template (PNG with white background)
- Visual styling: Background color, max words, width, height, color scheme
Create or edit your configuration file to match your analysis needs:
# docs/词云图使用配囨.md (example structure)
## 停用词配置
stopwords:
- 的
- 了
- 在
- 是
- 我
- 你
- 有
- 和
- 就
- 不
## 视觉配置
font_path: fonts/SimHei.ttf
background_color: white
max_words: 200
width: 800
height: 600
colormap: viridis
Step 3: Generate the Word Cloud Image
The tools/words.py script orchestrates the entire visualization pipeline: loading comments, segmenting text with jieba, filtering stopwords, computing frequencies, and rendering with wordcloud.
# Basic usage with default configuration
uv run tools/words.py \
--input data/comments_xhs.csv \
--output output/wordcloud.png
# With custom configuration file
uv run tools/words.py \
--input data/comments_xhs.csv \
--output output/custom_wordcloud.png \
--config docs/词云图使用配置.md
Understanding the Core Implementation
The tools/words.py script follows this processing flow as implemented in the MediaCrawler repository:
import argparse
import pandas as pd
import jieba
from wordcloud import WordCloud
def load_comments(path: str) -> list[str]:
"""Load comment text from CSV storage."""
df = pd.read_csv(path)
return df["comment"].dropna().tolist()
def build_frequency_dict(
comments: list[str],
stopwords: set[str]
) -> dict[str, int]:
"""
Tokenize Chinese comments with jieba and compute word frequencies.
Filters out stopwords and whitespace-only tokens.
"""
freq: dict[str, int] = {}
for text in comments:
for token in jieba.cut(text):
word = token.strip()
if word and word not in stopwords:
freq[word] = freq.get(word, 0) + 1
return freq
def create_wordcloud(
frequencies: dict[str, int],
config: dict
) -> WordCloud:
"""Initialize WordCloud with configuration parameters."""
return WordCloud(
font_path=config.get("font_path", "fonts/SimHei.ttf"),
background_color=config.get("background_color", "white"),
max_words=config.get("max_words", 200),
width=config.get("width", 800),
height=config.get("height", 600),
colormap=config.get("colormap", "viridis"),
mask=config.get("mask_image") # Optional numpy array
)
def main():
parser = argparse.ArgumentParser(
description="Generate word cloud from crawled comments"
)
parser.add_argument(
"--input",
required=True,
help="Path to CSV file containing comments"
)
parser.add_argument(
"--output",
required=True,
help="Output path for word cloud image"
)
parser.add_argument(
"--config",
default="docs/词云图使用配置.md",
help="Path to configuration markdown file"
)
args = parser.parse_args()
# Parse configuration (simplified—actual implementation parses YAML frontmatter)
config = parse_config(args.config)
stopwords = load_stopwords(config)
# Process comments
comments = load_comments(args.input)
frequencies = build_frequency_dict(comments, stopwords)
# Generate and save visualization
wc = create_wordcloud(frequencies, config)
wc.generate_from_frequencies(frequencies)
wc.to_file(args.output)
print(f"Word cloud saved to {args.output}")
if __name__ == "__main__":
main()
Advanced Customization Options
Using a Custom Mask Shape
For branded or thematic visualizations, specify a mask image in your configuration:
import numpy as np
from PIL import Image
# Inside create_wordcloud()
mask_image_path = config.get("mask_path")
if mask_image_path:
mask = np.array(Image.open(mask_image_path))
wc_config["mask"] = mask
The mask image must be a PNG with white background—white areas will be excluded from the word placement.
Processing Database-Stored Comments
If you configured SAVE_DATA_OPTION to use SQLite or MySQL, modify the loader function:
import sqlite3
def load_comments_from_db(db_path: str, table: str = "comments") -> list[str]:
conn = sqlite3.connect(db_path)
cursor = conn.execute(f"SELECT content FROM {table}")
comments = [row[0] for row in cursor.fetchall() if row[0]]
conn.close()
return comments
Then invoke with database parameters:
uv run tools/words.py \
--input data/crawler.db \
--output output/wordcloud.png \
--source-type sqlite
Troubleshooting Common Issues
- Missing Chinese characters: Ensure
font_pathpoints to a valid Chinese TrueType font (SimHei, Microsoft YaHei, or WenQuanYi Micro Hei) - Empty word cloud: Verify your CSV has a
commentcolumn; check that stopwords haven't excluded all meaningful tokens - jieba tokenization errors: Update jieba with
uv pip install -U jiebaand consider adding a custom dictionary for domain-specific terms - Memory errors with large datasets: Pre-filter comments by date or sample randomly before processing
Summary
- MediaCrawler's three-layer architecture separates crawling, storage, and visualization concerns
- Comments are collected via platform-specific crawlers in
store/and saved to your chosen format tools/words.pyperforms Chinese word segmentation with jieba and renders via the wordcloud library- Configuration in
docs/词云图使用配置.mdcontrols stopwords, fonts, colors, and mask shapes - The pipeline accepts CSV, JSON, or database sources without code changes to the visualization layer
Frequently Asked Questions
What file formats does the word cloud generator support for input?
The default implementation reads CSV files with a comment column. You can extend tools/words.py to support JSON, Excel, or database sources by modifying the load_comments() function—the visualization logic remains unchanged regardless of input format.
Can I use the word cloud tool with comments I crawled previously?
Yes. As long as your historical data contains comment text in a parseable structure, point --input to that file. The tool has no dependency on when the data was collected, only on the schema of the comment field.
How do I add domain-specific stopwords for my industry?
Edit docs/词云图使用配置.md and append terms to the stopwords list. For programmatic control, you can also pass a custom stopwords file via --stopwords-path if you modify tools/words.py to accept this parameter.
Why are some high-frequency words not appearing in my word cloud?
Check three things: (1) the word may be in your stopwords list, (2) jieba may be segmenting it differently than expected—test with jieba.cut() directly, or (3) max_words may be truncating the result—increase this value in your configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →