# TrendRadar Core Functionalities: A Complete Guide to the Modular Python News Aggregator

> Discover TrendRadar's core functionalities: automate news collection, filtering, ranking, AI enrichment, and notifications with this modular Python news aggregator.

- Repository: [sansan/TrendRadar](https://github.com/sansan0/TrendRadar)
- Tags: deep-dive
- Published: 2026-04-22

---

**TrendRadar is a modular Python tool that automates the collection, filtering, ranking, AI enrichment, and notification of trending news from multiple online platforms.**

TrendRadar orchestrates an end-to-end pipeline for discovering and distributing trending content. Built with a component-based architecture, it separates concerns across configuration management, scheduling, data crawling, storage, frequency-based analysis, optional AI processing, and multi-channel notifications. This article breaks down each core functionality with direct references to the source code in the [sansan0/TrendRadar](https://github.com/sansan0/TrendRadar) repository.

---

## Configuration Loading and Environment Management

Every TrendRadar execution begins with unified configuration assembly. The **configuration loader** in [`trendradar/core/loader.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/loader.py) reads YAML files from the `config/` directory, parses the [`frequency_words.txt`](https://github.com/sansan0/TrendRadar/blob/main/frequency_words.txt) rule file, and merges environment variable overrides into a single dictionary.

Key capabilities include:

- **Environment variable overrides** for `TIMEZONE`, `DEBUG`, `MAX_NEWS_PER_KEYWORD`, and other runtime settings
- **YAML merging** that combines [`config.yaml`](https://github.com/sansan0/TrendRadar/blob/main/config.yaml) with platform-specific and timeline configurations
- **Frequency word preprocessing** that converts text rules into structured matching directives

```python
from trendradar.core.loader import load_config

cfg = load_config()  # reads config/config.yaml by default

print("Enabled platforms:", cfg["PLATFORMS"])
print("Timezone:", cfg["TIMEZONE"])

```

---

## Timeline-Based Scheduling System

The **scheduler** in [`trendradar/core/scheduler.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/scheduler.py) interprets [`timeline.yaml`](https://github.com/sansan0/TrendRadar/blob/main/timeline.yaml) to determine which pipeline actions should execute at any given moment. It supports presets, custom periods, daily/weekly plans, and deduplication logic to prevent redundant runs.

The `Scheduler` class resolves three boolean flags:

- **`collect`**: Whether to crawl fresh data from platforms
- **`analyze`**: Whether to process and rank the collected data
- **`push`**: Whether to distribute notifications to configured channels

```python
from trendradar.core.scheduler import Scheduler
from trendradar.utils.time import get_configured_time

schedule = Scheduler(
    schedule_config=cfg["SCHEDULE"],
    timeline_data=cfg["_TIMELINE_DATA"],
    storage_backend=None,
    get_time_func=lambda: get_configured_time(cfg["TIMEZONE"])
)

resolved = schedule.resolve()
print(f"Collect? {resolved.collect}, Analyze? {resolved.analyze}, Push? {resolved.push}")

```

---

## Multi-Platform Data Crawling

The **crawler** in [`trendradar/crawler/__init__.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/crawler/__init__.py) (organized as the `DataFetcher` class) performs HTTP requests to configured platforms, respecting per-platform rate limits and optional proxy settings. It returns raw title-rank mappings that feed into downstream analysis.

Crawler capabilities include:

- **Request interval enforcement** to respect platform rate limits
- **Proxy configuration** for geographic or network flexibility
- **Platform-specific parsers** that normalize diverse response formats into consistent structures

---

## Persistent Storage and Incremental Detection

The **storage manager** in [`trendradar/storage/__init__.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/__init__.py) persists raw crawl results to SQLite and optional TXT/HTML snapshots. It provides read-only access to current and historical data, enabling incremental mode that detects newly added titles between runs.

Storage functions include:

- **SQLite persistence** for structured querying and deduplication
- **Snapshot generation** in TXT and HTML formats for manual inspection
- **Incremental detection** comparing current batches against stored history

---

## Frequency-Word Engine for Filtering and Ranking

The **frequency-word engine** in [`trendradar/core/frequency.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/frequency.py) parses [`frequency_words.txt`](https://github.com/sansan0/TrendRadar/blob/main/frequency_words.txt) with a flexible syntax supporting required terms (`+`), filters (`!`), maximum counts (`@`), regex patterns (`/…/`), and display name mapping (`=>`).

The engine exposes three key functions:

- **`load_frequency_words()`**: Parses the rule file into structured groups
- **`matches_word_groups()`**: Tests titles against all configured groups
- **`_word_matches()`**: Low-level matcher handling individual rule types

---

## Statistical Analysis and Weighted Ranking

The **analyzer** in [`trendradar/core/analyzer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/analyzer.py) aggregates title matches per keyword, applies multi-factor weighting, and produces final ranked news lists. It also handles RSS feeds through `count_rss_frequency` and related functions.

Weighting factors include:

- **Ranking weight**: Original platform position (lower rank = higher weight)
- **Frequency weight**: How often the keyword appears across platforms
- **Hotness weight**: Composite momentum score combining recency and velocity

```python
from trendradar.core.analyzer import count_word_frequency
from trendradar.core.frequency import load_frequency_words

word_groups, filter_words, global_filters = load_frequency_words()

stats, total = count_word_frequency(
    results,           # dict from crawler

    word_groups,
    filter_words,
    id_to_name,        # platform_id → readable name

    rank_threshold=3,
    max_news_per_keyword=5,
    sort_by_position_first=True,
    quiet=False
)

for group in stats:
    print(f'Keyword: {group["word"]} – matched {group["count"]} items')

```

---

## RSS Feed Integration

TrendRadar treats RSS sources as first-class citizens alongside hot-list platforms. The analyzer's RSS functions apply identical frequency-word filtering and weighting, converting feed items into the same data structure as crawled hot-lists.

This enables unified reporting that blends trending topics from social platforms with curated RSS content, all processed through the same ranking pipeline.

---

## AI-Powered Enrichment and Translation

The **AI processor** in [`trendradar/ai/processor.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/processor.py) (and referenced via `NewsAnalyzer._run_ai_analysis`) provides optional LiteLLM-compatible integration for intelligent content enhancement.

AI capabilities include:

- **Content analysis**: Summarization, sentiment scoring, and topic extraction
- **Translation**: Automated conversion to configured target languages
- **Intelligent filtering**: AI-assisted relevance scoring beyond rule-based matching

---

## Multi-Channel Notification Dispatch

The **notification subsystem** spans three files in `trendradar/notification/`:

- **[`renderer.py`](https://github.com/sansan0/TrendRadar/blob/main/renderer.py)**: Formats final reports in Markdown and HTML
- **[`splitter.py`](https://github.com/sansan0/TrendRadar/blob/main/splitter.py)**: Divides unified reports into channel-specific payloads
- **[`senders.py`](https://github.com/sansan0/TrendRadar/blob/main/senders.py)**: Implements all platform integrations

Supported channels include **Feishu**, **DingTalk**, **Telegram**, **Email**, and extensible webhook formats.

```python
from trendradar.notification.senders import FeishuSender

payload = {
    "title": "今日热点新闻",
    "content": "• 关键字 A: 5 条\n• 关键字 B: 3 条"
}

sender = FeishuSender(
    webhook_url="https://open.feishu.cn/open-apis/bot/v2/hook/your-token"
)
sender.send(payload)  # HTTP POST to webhook

```

---

## CLI Entry Point and Workflow Orchestration

The **main entry point** in [`trendradar/__main__.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/__main__.py) wires all components into a cohesive pipeline. It creates `AppContext`, executes the scheduler, conditionally runs collection/analysis/notification phases, and generates HTML reports.

Execution modes include:

- **Default/daily mode**: Full pipeline execution
- **`--mode=current`**: Only current ranking slice

```bash

# Full pipeline (default)

python -m trendradar

# Current ranking only

python -m trendradar --mode=current

```

---

## Summary

TrendRadar delivers comprehensive trending news automation through these core capabilities:

- **Modular configuration** with YAML and environment variable overrides
- **Timeline scheduling** supporting presets, custom periods, and deduplication
- **Multi-platform crawling** with rate limiting and proxy support
- **Persistent storage** with SQLite and incremental detection
- **Frequency-word engine** with flexible rule syntax for filtering and ranking
- **Statistical analysis** with multi-factor weighting (rank, frequency, hotness)
- **RSS integration** treated as first-class data sources
- **AI enrichment** via LiteLLM for summarization, translation, and intelligent filtering
- **Multi-channel notifications** supporting Feishu, DingTalk, Telegram, Email, and webhooks
- **CLI orchestration** with configurable execution modes

---

## Frequently Asked Questions

### How does TrendRadar prevent duplicate notifications?

The **scheduler** in [`trendradar/core/scheduler.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/scheduler.py) maintains execution records and implements "once-per-day" deduplication. When `push` is resolved as `True`, the scheduler checks whether a notification was already sent for the current period before allowing dispatch.

### What platforms can TrendRadar crawl for trending topics?

TrendRadar's **crawler** supports configurable HTTP-based platforms defined in YAML configuration. The extensible `DataFetcher` class in [`trendradar/crawler/__init__.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/crawler/__init__.py) handles per-platform request intervals, proxy settings, and response parsing. Specific platform implementations are loaded based on the `PLATFORMS` configuration key.

### How do frequency-word rules work in TrendRadar?

The **frequency-word engine** in [`trendradar/core/frequency.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/frequency.py) parses [`frequency_words.txt`](https://github.com/sansan0/TrendRadar/blob/main/frequency_words.txt) with syntax supporting: `+` (required terms), `!` (exclusion filters), `@` (maximum match counts), `/.../` (regex patterns), and `=>` (display name mapping). The `matches_word_groups` function tests titles against these rules to categorize and rank content.

### Can TrendRadar integrate with custom AI models?

Yes. The **AI processor** in [`trendradar/ai/processor.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/processor.py) uses **LiteLLM** for model compatibility, supporting any provider or self-hosted endpoint that LiteLLM recognizes. Configuration is loaded from YAML with environment-variable overrides for API keys and model selection, enabling integration with OpenAI, Anthropic, local Ollama instances, or custom endpoints.