# What Is the Architecture of TrendRadar? A Deep Dive into the Modular Python News Aggregator

> Explore the modular layered architecture of TrendRadar, a Python news aggregator. Discover how its pluggable components handle crawling, analysis, AI filtering, and notifications.

- Repository: [sansan/TrendRadar](https://github.com/sansan0/TrendRadar)
- Tags: architecture
- Published: 2026-04-22

---

**TrendRadar is a modular Python application with a layered architecture that separates crawling, storage, analysis, AI filtering, and notification into pluggable components coordinated through a central AppContext object.**

This article explains the complete architecture of [TrendRadar](https://github.com/sansan0/TrendRadar), an open-source news aggregation and reporting system. You'll learn how each layer functions, how data flows through the pipeline, and how the design enables extensibility without core code changes.

---

## Architectural Overview

TrendRadar follows a **layered, context-driven architecture** with clear separation of concerns. The system coordinates 11 distinct layers through a shared `AppContext` object that eliminates global state and enables testability.

| Layer | Responsibility | Main File Paths |
|-------|---------------|---------------|
| **Entry Point** | CLI/daemon startup, argument parsing | [`trendradar/__main__.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/__main__.py) |
| **Configuration & Context** | YAML/ENV loading, unified context object | [`trendradar/core/loader.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/loader.py), [`trendradar/context.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/context.py) |
| **Crawler** | Platform web scraping and RSS fetching | `trendradar/crawler/DataFetcher`, [`trendradar/crawler/rss.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/crawler/rss.py) |
| **Storage** | SQLite persistence, retention, remote sync | [`trendradar/storage/manager.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/manager.py), [`trendradar/storage/sqlite_mixin.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/sqlite_mixin.py) |
| **Analysis** | Keyword matching, weight calculation, aggregation | [`trendradar/core/analyzer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/analyzer.py), [`trendradar/core/frequency.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/frequency.py) |
| **AI Filter** | LLM-based news classification and tagging | [`trendradar/ai/filter.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/filter.py), [`trendradar/ai/client.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/client.py) |
| **Reporting** | HTML generation, data structure preparation | [`trendradar/report/__init__.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/report/__init__.py), [`trendradar/report/generator.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/report/generator.py) |
| **Notification** | Multi-channel dispatch with batching and translation | [`trendradar/notification/dispatcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/dispatcher.py), [`trendradar/notification/renderer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/renderer.py) |
| **Scheduler** | Timeline-based execution with "once" semantics | [`trendradar/core/scheduler.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/scheduler.py) |

---

## Core Design Patterns

### Context-Driven Dependency Injection

Every subsystem receives the same `AppContext` instance, making the codebase **testable and free of global state**. The `AppContext` class in [`trendradar/context.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/context.py) exposes:

- Configuration access
- Time utilities (`load_frequency_words`, `count_frequency`)
- Storage interface
- High-level workflow methods (`prepare_report`, `generate_html`, `run_ai_filter`)

### Pluggable Filtering Architecture

The `FILTER.METHOD` configuration toggles between **keyword matching** and **AI-driven tagging** without changing the surrounding pipeline. This is implemented through:

- [`trendradar/core/analyzer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/analyzer.py) – keyword-based word frequency and weight calculation
- [`trendradar/ai/filter.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/filter.py) – LLM classification with tag extraction and versioning

Both return compatible data structures that feed into the same reporting and notification layers.

---

## Data Flow Through the Pipeline

Understanding the complete **TrendRadar architecture** requires tracing how data moves from source to notification:

### 1. Startup and Configuration

```python
from trendradar.core.loader import load_config
from trendradar.context import AppContext

# Load configuration (uses CONFIG_PATH env var or defaults)

cfg = load_config()

# Create shared context

ctx = AppContext(cfg)

```

The `load_config()` function in [`trendradar/core/loader.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/loader.py) loads and validates the entire configuration tree from YAML files and environment variables.

### 2. Scheduling Resolution

```python

# Create scheduler and check if we should run

scheduler = ctx.create_scheduler()
should_run = scheduler.should_crawl()  # or should_analyze(), should_push()

```

The [`trendradar/core/scheduler.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/scheduler.py) module implements **timeline-based scheduling** with cron-like periods and "once" flags to prevent duplicate executions.

### 3. Data Crawling

```python

# Crawl platforms and RSS feeds

results, id_to_name, failed = ctx.create_scheduler().crawl_all()

```

The crawler layer in `trendradar/crawler/DataFetcher` and [`trendradar/crawler/rss.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/crawler/rss.py) fetches raw news from web platforms and RSS feeds, normalizing everything into a common data structure.

### 4. Storage and Deduplication

Crawled items persist through [`trendradar/storage/manager.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/manager.py), which provides:

- **SQLite** storage via [`sqlite_mixin.py`](https://github.com/sansan0/TrendRadar/blob/main/sqlite_mixin.py)
- **Retention policies** for automatic cleanup
- **Remote sync** capabilities via [`trendradar/storage/remote.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/remote.py)

New-title detection happens through `ctx.detect_new_titles()` and `ctx.detect_latest_new_titles()`, which identify titles appearing for the first time today.

### 5. Analysis and Filtering

**Keyword mode:**

```python
stats, total = ctx.count_frequency(
    results,
    *ctx.load_frequency_words(),
    id_to_name,
    title_info,
    new_titles=new_titles,
    mode="daily"
)

```

The [`trendradar/core/analyzer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/analyzer.py) module handles keyword-group matching, weight calculation, and platform-wise aggregation. Weights combine rank, frequency, and hotness scores.

**AI mode:**

```python

# Enable AI filtering

cfg["FILTER"]["METHOD"] = "ai"
ctx = AppContext(cfg)

# Run AI filter

ai_result = ctx.run_ai_filter(interests_file="my_interests.txt")

# Convert to report format

hotlist_stats, rss_stats = ctx.convert_ai_filter_to_report_data(ai_result, mode="daily")

```

The [`trendradar/ai/filter.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/filter.py) pipeline classifies items using large language models, handles tag extraction and versioning, and returns a unified result set compatible with the reporting layer.

### 6. Report Generation

```python

# Generate HTML report

html_path = ctx.generate_html(
    stats,
    total,
    new_titles=new_titles,
    id_to_name=id_to_name,
    mode="daily"
)

# Prepare language-agnostic report data

report_data = ctx.prepare_report(hotlist_stats, id_to_name=id_to_name, mode="daily")

```

The reporting layer in [`trendradar/report/__init__.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/report/__init__.py) and [`trendradar/report/generator.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/report/generator.py) builds data structures for HTML and text rendering. Templates reside in `docs/`.

### 7. Notification Dispatch

```python
dispatcher = ctx.create_notification_dispatcher()
dispatcher.dispatch_all(
    report_data=report_data,
    report_type="热点分析报告",
    update_info=None,
    proxy_url=None,
    mode="daily",
    html_file_path=None,
    rss_items=rss_stats,
    rss_new_items=None,
    ai_analysis=None,
    standalone_data=None,
    skip_translation=False,
)

```

The [`trendradar/notification/dispatcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/dispatcher.py) sends rendered reports to **10+ channels**: Feishu, DingTalk, WeWork, Telegram, Email, Ntfy, Bark, Slack, and generic webhooks. Key supporting modules:

- [`trendradar/notification/renderer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/renderer.py) – channel-specific payload formatting
- [`trendradar/notification/splitter.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/splitter.py) – batch splitting for large reports
- `AITranslator` – optional translation to reduce API costs

---

## Key Design Decisions

### Flexible Scheduling with Timeline Semantics

The [`timeline.yaml`](https://github.com/sansan0/TrendRadar/blob/main/timeline.yaml) configuration defines daily/weekly "periods" (e.g., "morning", "evening") with "once" flags to prevent duplicate executions. This is implemented in [`trendradar/core/scheduler.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/scheduler.py) through the `ResolvedSchedule` class.

### Storage Abstraction

`StorageManager` in [`trendradar/storage/manager.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/manager.py) hides implementation details behind a uniform API:

| Backend | Implementation |
|---------|---------------|
| Local SQLite | [`sqlite_mixin.py`](https://github.com/sansan0/TrendRadar/blob/main/sqlite_mixin.py) |
| Plain text | Direct file I/O |
| Remote S3 | [`trendradar/storage/remote.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/remote.py) |
| Pull sync | Sync protocol in [`remote.py`](https://github.com/sansan0/TrendRadar/blob/main/remote.py) |

This allows swapping backends without touching business logic.

### Modular Notification System

Each notification channel has its own **renderer** and **batch size configuration**. The dispatcher can skip translation for already-translated data, reducing API costs when multiple channels share the same content.

---

## Summary

- **TrendRadar architecture** is a **layered, modular Python application** with 11 distinct functional layers coordinated through `AppContext`
- **Context-driven dependency injection** eliminates global state and enables comprehensive testing
- **Pluggable filtering** supports both keyword-based and AI-driven analysis without pipeline changes
- **Timeline-based scheduling** with "once" semantics prevents duplicate executions
- **Storage abstraction** unifies SQLite, file, and remote S3 backends behind a common API
- **Multi-channel notifications** support 10+ platforms with batching, translation, and channel-specific rendering

---

## Frequently Asked Questions

### What programming pattern does TrendRadar use for dependency management?

TrendRadar uses **context-driven dependency injection**. Every subsystem receives the same `AppContext` instance passed through constructors or method parameters. This pattern eliminates global state, makes components individually testable with mocked contexts, and centralizes access to configuration, storage, and utilities in [`trendradar/context.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/context.py).

### How does TrendRadar switch between keyword and AI filtering?

TrendRadar toggles filtering methods through the `FILTER.METHOD` configuration key. Setting `"keyword"` activates the [`trendradar/core/analyzer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/core/analyzer.py) pipeline for word frequency and weight calculation. Setting `"ai"` activates [`trendradar/ai/filter.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/ai/filter.py) for LLM-based classification. Both pipelines return compatible data structures, so [`trendradar/report/__init__.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/report/__init__.py) and [`trendradar/notification/dispatcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/dispatcher.py) work unchanged regardless of filtering method.

### Can TrendRadar store data on remote servers instead of locally?

Yes. The `StorageManager` class in [`trendradar/storage/manager.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/manager.py) abstracts storage backends behind a uniform API. While local SQLite is the default via [`trendradar/storage/sqlite_mixin.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/sqlite_mixin.py), the system also supports remote S3 storage and pull-sync protocols through [`trendradar/storage/remote.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/remote.py). You can swap backends by changing configuration without modifying business logic in crawlers, analyzers, or notifiers.

### What notification channels does TrendRadar support?

TrendRadar supports **10 notification channels**: Feishu, DingTalk, WeWork, Telegram, Email, Ntfy, Bark, Slack, and generic webhooks. The [`trendradar/notification/dispatcher.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/dispatcher.py) coordinates delivery, while [`trendradar/notification/renderer.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/renderer.py) formats channel-specific payloads. The [`trendradar/notification/splitter.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/notification/splitter.py) handles batching for large reports, and optional `AITranslator` integration reduces API costs when multiple channels share content.