What Kind of Data Does TrendRadar Process? A Complete Guide to Its Data Pipeline

TrendRadar processes two primary data types: hot-list news from the NewsNow API and RSS feed items, normalizing both into internal models before storage, analysis, and report generation.

TrendRadar is an open-source news-trend aggregation and analysis service. Understanding what kind of data TrendRadar processes is essential for customizing its behavior or extending its capabilities. This article breaks down the complete data pipeline from ingestion to final output.

Two Primary Data Sources

TrendRadar ingests data from two distinct external sources, each with its own raw format and parsing logic.

Hot-List News from NewsNow API

The NewsNow API provides ranked, trending news stories from various platforms. TrendRadar fetches these through the DataFetcher class in trendradar/crawler/fetcher.py (lines 20-46).

Raw JSON objects contain fields like title, url, rank, and platform metadata. These are normalized into NewsItem objects with the following structure:

  • title: Article headline
  • source_id: Platform identifier (e.g., "toutiao", "baidu")
  • rank_list: Historical ranking data
  • first_seen / last_seen: Appearance timestamps
  • occurrence_count: How many times the item appeared

RSS Feeds from Configurable Sources

TrendRadar also ingests standard RSS/Atom feeds through trendradar/crawler/rss/fetcher.py (class RSSFetcher, lines 15-66). This supports any feed URL you configure.

Raw feed items (XML or JSON) with title, link, published, and optional summary/author fields become RSSItem objects:

  • title: Article headline
  • feed_id: Configured feed identifier
  • url: Direct link to article
  • publish_time: Original publication timestamp
  • summary: Optional excerpt
  • author: Optional byline

Internal Data Models and Storage

Both NewsItem and RSSItem inherit from abstract contracts defined in trendradar/storage/base.py. This ensures consistent handling despite different source formats.

Collection Containers

Before storage, individual items are wrapped into collection objects:

Container Contents Metadata
NewsData source_id → List[NewsItem] crawl date, crawl time, source name mapping
RSSData feed_id → List[RSSItem] crawl date, crawl time, similar meta-fields

These structures appear in trendradar/storage/base.py at lines 22-38 (RSSData) and lines 77-92 (NewsData).

Pluggable Storage Backend

TrendRadar persists data through the StorageBackend abstract interface (trendradar/storage/base.py, lines 98-136). Concrete implementations include:

This modularity lets you swap storage technologies without touching ingestion or analysis code.

Processing Pipeline: From Raw Data to Insights

Once stored, TrendRadar applies a multi-stage analysis pipeline to transform raw items into actionable trend reports.

Stage 1: Loading Today's Data

The read_all_today_titles function in trendradar/core/data.py (lines 82-102) extracts the current day's hot-list titles from storage. This establishes the baseline for comparison.

Stage 2: Detecting New Titles

detect_latest_new_titles (same file, lines 123-158) compares the latest crawl batch against historical records. It identifies truly new items that haven't appeared before—critical for spotting emerging trends.

Stage 3: Statistical and AI-Enhanced Analysis

The count_word_frequency function in trendradar/core/analyzer.py (lines 93-140) performs the core analysis:

  • Groups titles by keyword groups (configured in config.yaml)
  • Applies ranking, frequency, and "hotness" weightings
  • Marks newly-detected titles for highlighted reporting

Optional AI filtering via trendradar/ai/filter.py and trendradar/ai/client.py can further refine results.

Stage 4: Report Generation and Notification

Final outputs are produced by:

Complete Data Flow Example

Here's how the full pipeline works in practice:


# Fetch and store hot-list data

from trendradar.crawler.fetcher import DataFetcher
from trendradar.storage.manager import StorageManager
from trendradar.storage.base import convert_crawl_results_to_news_data
from datetime import datetime

fetcher = DataFetcher(proxy_url=None)
ids = ["toutiao", ("baidu", "Baidu 热榜")]
results, id_to_name, failed = fetcher.crawl_websites(ids)

now = datetime.now()
news_data = convert_crawl_results_to_news_data(
    results, id_to_name, failed,
    crawl_time=now.strftime("%H:%M"),
    crawl_date=now.strftime("%Y-%m-%d"),
)

storage = StorageManager()
storage.save_news_data(news_data)

# Analyze and generate reports

from trendradar.core.data import read_all_today_titles, detect_latest_new_titles
from trendradar.core.analyzer import count_word_frequency
from trendradar.core.loader import load_config

all_titles, id_to_name, title_info = read_all_today_titles(storage)
new_titles = detect_latest_new_titles(storage, quiet=True)

cfg = load_config()
stats, total = count_word_frequency(
    results=all_titles,
    word_groups=cfg["word_groups"],
    filter_words=cfg["filter_words"],
    id_to_name=id_to_name,
    title_info=title_info,
    new_titles=new_titles,
    mode="daily",
    weight_config=cfg["weight_config"],
)

Summary

TrendRadar processes two distinct data types that power its trend detection pipeline:

  • Hot-list news from the NewsNow API, normalized into NewsItem objects with ranking history and appearance tracking
  • RSS feed items from configurable sources, normalized into RSSItem objects with publication metadata

Both stream through a modular pipeline: ingestion via DataFetcher and RSSFetcher, storage through the StorageBackend interface, analysis via count_word_frequency and detect_latest_new_titles, and final output through report generators and notification dispatchers.

Frequently Asked Questions

What data formats does TrendRadar accept as input?

TrendRadar accepts JSON from the NewsNow API and standard RSS/Atom feeds (XML or JSON). The DataFetcher class handles hot-list JSON in trendradar/crawler/fetcher.py, while RSSFetcher in trendradar/crawler/rss/fetcher.py parses RSS feeds using configurable feed definitions.

TrendRadar identifies new topics through the detect_latest_new_titles function in trendradar/core/data.py (lines 123-158). This compares the current crawl batch against historical records stored in the backend, returning only titles that have never appeared before. These newly detected items receive special highlighting in reports.

Can TrendRadar process custom RSS feeds beyond the defaults?

Yes. The RSSFetcher class reads feed configurations from a dictionary or YAML file, allowing arbitrary RSS/Atom URLs. Each feed requires an id, name, and url. The fetcher applies freshness filtering and deduplication automatically. Configure feeds in your config.yaml and pass them to RSSFetcher.from_config().

What storage backends does TrendRadar support?

TrendRadar uses a pluggable storage architecture via the StorageBackend abstract class in trendradar/storage/base.py. The reference implementation uses SQLite via trendradar/storage/sqlite_mixin.py, orchestrated by trendradar/storage/manager.py. You can implement custom backends (remote APIs, cloud databases) by subclassing StorageBackend and implementing save_news_data, load_news_data, and related methods.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →