# What Kind of Data Does TrendRadar Process? A Complete Guide to Its Data Pipeline

> Discover what data TrendRadar processes including NewsNow API hot-list news and RSS feeds. Learn about its comprehensive data pipeline for analysis and reporting.

- Repository: [sansan/TrendRadar](https://github.com/sansan0/TrendRadar)
- Tags: deep-dive
- Published: 2026-04-22

---

**TrendRadar processes two primary data types: hot-list news from the NewsNow API and RSS feed items, normalizing both into internal models before storage, analysis, and report generation.**

TrendRadar is an open-source news-trend aggregation and analysis service. Understanding what kind of data TrendRadar processes is essential for customizing its behavior or extending its capabilities. This article breaks down the complete data pipeline from ingestion to final output.

## Two Primary Data Sources

TrendRadar ingests data from two distinct external sources, each with its own raw format and parsing logic.

### Hot-List News from NewsNow API

The **NewsNow API** provides ranked, trending news stories from various platforms. TrendRadar fetches these through the `DataFetcher` class in [`trendradar/crawler/fetcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/crawler/fetcher.py) (lines 20-46).

Raw JSON objects contain fields like `title`, `url`, `rank`, and platform metadata. These are normalized into **`NewsItem`** objects with the following structure:

- `title`: Article headline
- `source_id`: Platform identifier (e.g., "toutiao", "baidu")
- `rank_list`: Historical ranking data
- `first_seen` / `last_seen`: Appearance timestamps
- `occurrence_count`: How many times the item appeared

### RSS Feeds from Configurable Sources

TrendRadar also ingests **standard RSS/Atom feeds** through [`trendradar/crawler/rss/fetcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/crawler/rss/fetcher.py) (class `RSSFetcher`, lines 15-66). This supports any feed URL you configure.

Raw feed items (XML or JSON) with `title`, `link`, `published`, and optional `summary`/`author` fields become **`RSSItem`** objects:

- `title`: Article headline
- `feed_id`: Configured feed identifier
- `url`: Direct link to article
- `publish_time`: Original publication timestamp
- `summary`: Optional excerpt
- `author`: Optional byline

## Internal Data Models and Storage

Both `NewsItem` and `RSSItem` inherit from abstract contracts defined in [`trendradar/storage/base.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/base.py). This ensures consistent handling despite different source formats.

### Collection Containers

Before storage, individual items are wrapped into collection objects:

| Container | Contents | Metadata |
|-----------|----------|----------|
| **`NewsData`** | `source_id → List[NewsItem]` | crawl date, crawl time, source name mapping |
| **`RSSData`** | `feed_id → List[RSSItem]` | crawl date, crawl time, similar meta-fields |

These structures appear in [`trendradar/storage/base.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/base.py) at lines 22-38 (RSSData) and lines 77-92 (NewsData).

### Pluggable Storage Backend

TrendRadar persists data through the **`StorageBackend`** abstract interface ([`trendradar/storage/base.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/base.py), lines 98-136). Concrete implementations include:

- **SQLite backend**: [`trendradar/storage/sqlite_mixin.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/sqlite_mixin.py)
- **Storage manager**: [`trendradar/storage/manager.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/manager.py) (orchestrates operations)

This modularity lets you swap storage technologies without touching ingestion or analysis code.

## Processing Pipeline: From Raw Data to Insights

Once stored, TrendRadar applies a multi-stage analysis pipeline to transform raw items into actionable trend reports.

### Stage 1: Loading Today's Data

The `read_all_today_titles` function in [`trendradar/core/data.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/data.py) (lines 82-102) extracts the current day's hot-list titles from storage. This establishes the baseline for comparison.

### Stage 2: Detecting New Titles

`detect_latest_new_titles` (same file, lines 123-158) compares the latest crawl batch against historical records. It identifies truly new items that haven't appeared before—critical for spotting emerging trends.

### Stage 3: Statistical and AI-Enhanced Analysis

The `count_word_frequency` function in [`trendradar/core/analyzer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/analyzer.py) (lines 93-140) performs the core analysis:

- Groups titles by **keyword groups** (configured in [`config.yaml`](https://github.com/sansan0/TrendRadar/blob/main/config.yaml))
- Applies **ranking**, **frequency**, and **"hotness"** weightings
- Marks newly-detected titles for highlighted reporting

Optional AI filtering via [`trendradar/ai/filter.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/filter.py) and [`trendradar/ai/client.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/client.py) can further refine results.

### Stage 4: Report Generation and Notification

Final outputs are produced by:

- **Report generator**: [`trendradar/report/generator.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/report/generator.py) (HTML/TXT snapshots)
- **Formatter**: [`trendradar/report/formatter.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/report/formatter.py) (styling and layout)
- **Notification dispatcher**: [`trendradar/notification/dispatcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/dispatcher.py) (push to DingTalk, WeChat, etc.)

## Complete Data Flow Example

Here's how the full pipeline works in practice:

```python

# Fetch and store hot-list data

from trendradar.crawler.fetcher import DataFetcher
from trendradar.storage.manager import StorageManager
from trendradar.storage.base import convert_crawl_results_to_news_data
from datetime import datetime

fetcher = DataFetcher(proxy_url=None)
ids = ["toutiao", ("baidu", "Baidu 热榜")]
results, id_to_name, failed = fetcher.crawl_websites(ids)

now = datetime.now()
news_data = convert_crawl_results_to_news_data(
    results, id_to_name, failed,
    crawl_time=now.strftime("%H:%M"),
    crawl_date=now.strftime("%Y-%m-%d"),
)

storage = StorageManager()
storage.save_news_data(news_data)

```

```python

# Analyze and generate reports

from trendradar.core.data import read_all_today_titles, detect_latest_new_titles
from trendradar.core.analyzer import count_word_frequency
from trendradar.core.loader import load_config

all_titles, id_to_name, title_info = read_all_today_titles(storage)
new_titles = detect_latest_new_titles(storage, quiet=True)

cfg = load_config()
stats, total = count_word_frequency(
    results=all_titles,
    word_groups=cfg["word_groups"],
    filter_words=cfg["filter_words"],
    id_to_name=id_to_name,
    title_info=title_info,
    new_titles=new_titles,
    mode="daily",
    weight_config=cfg["weight_config"],
)

```

## Summary

TrendRadar processes two distinct data types that power its trend detection pipeline:

- **Hot-list news** from the NewsNow API, normalized into `NewsItem` objects with ranking history and appearance tracking
- **RSS feed items** from configurable sources, normalized into `RSSItem` objects with publication metadata

Both stream through a modular pipeline: ingestion via `DataFetcher` and `RSSFetcher`, storage through the `StorageBackend` interface, analysis via `count_word_frequency` and `detect_latest_new_titles`, and final output through report generators and notification dispatchers.

## Frequently Asked Questions

### What data formats does TrendRadar accept as input?

TrendRadar accepts **JSON from the NewsNow API** and **standard RSS/Atom feeds** (XML or JSON). The `DataFetcher` class handles hot-list JSON in [`trendradar/crawler/fetcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/crawler/fetcher.py), while `RSSFetcher` in [`trendradar/crawler/rss/fetcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/crawler/rss/fetcher.py) parses RSS feeds using configurable feed definitions.

### How does TrendRadar identify new trending topics?

TrendRadar identifies new topics through the `detect_latest_new_titles` function in [`trendradar/core/data.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/data.py) (lines 123-158). This compares the current crawl batch against historical records stored in the backend, returning only titles that have never appeared before. These newly detected items receive special highlighting in reports.

### Can TrendRadar process custom RSS feeds beyond the defaults?

Yes. The `RSSFetcher` class reads feed configurations from a dictionary or YAML file, allowing arbitrary RSS/Atom URLs. Each feed requires an `id`, `name`, and `url`. The fetcher applies freshness filtering and deduplication automatically. Configure feeds in your [`config.yaml`](https://github.com/sansan0/TrendRadar/blob/main/config.yaml) and pass them to `RSSFetcher.from_config()`.

### What storage backends does TrendRadar support?

TrendRadar uses a pluggable storage architecture via the `StorageBackend` abstract class in [`trendradar/storage/base.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/base.py). The reference implementation uses **SQLite** via [`trendradar/storage/sqlite_mixin.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/sqlite_mixin.py), orchestrated by [`trendradar/storage/manager.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/manager.py). You can implement custom backends (remote APIs, cloud databases) by subclassing `StorageBackend` and implementing `save_news_data`, `load_news_data`, and related methods.