What Is the Architecture of TrendRadar? A Deep Dive into the Modular Python News Aggregator

TrendRadar is a modular Python application with a layered architecture that separates crawling, storage, analysis, AI filtering, and notification into pluggable components coordinated through a central AppContext object.

This article explains the complete architecture of TrendRadar, an open-source news aggregation and reporting system. You'll learn how each layer functions, how data flows through the pipeline, and how the design enables extensibility without core code changes.


Architectural Overview

TrendRadar follows a layered, context-driven architecture with clear separation of concerns. The system coordinates 11 distinct layers through a shared AppContext object that eliminates global state and enables testability.

Layer Responsibility Main File Paths
Entry Point CLI/daemon startup, argument parsing trendradar/__main__.py
Configuration & Context YAML/ENV loading, unified context object trendradar/core/loader.py, trendradar/context.py
Crawler Platform web scraping and RSS fetching trendradar/crawler/DataFetcher, trendradar/crawler/rss.py
Storage SQLite persistence, retention, remote sync trendradar/storage/manager.py, trendradar/storage/sqlite_mixin.py
Analysis Keyword matching, weight calculation, aggregation trendradar/core/analyzer.py, trendradar/core/frequency.py
AI Filter LLM-based news classification and tagging trendradar/ai/filter.py, trendradar/ai/client.py
Reporting HTML generation, data structure preparation trendradar/report/__init__.py, trendradar/report/generator.py
Notification Multi-channel dispatch with batching and translation trendradar/notification/dispatcher.py, trendradar/notification/renderer.py
Scheduler Timeline-based execution with "once" semantics trendradar/core/scheduler.py

Core Design Patterns

Context-Driven Dependency Injection

Every subsystem receives the same AppContext instance, making the codebase testable and free of global state. The AppContext class in trendradar/context.py exposes:

  • Configuration access
  • Time utilities (load_frequency_words, count_frequency)
  • Storage interface
  • High-level workflow methods (prepare_report, generate_html, run_ai_filter)

Pluggable Filtering Architecture

The FILTER.METHOD configuration toggles between keyword matching and AI-driven tagging without changing the surrounding pipeline. This is implemented through:

Both return compatible data structures that feed into the same reporting and notification layers.


Data Flow Through the Pipeline

Understanding the complete TrendRadar architecture requires tracing how data moves from source to notification:

1. Startup and Configuration

from trendradar.core.loader import load_config
from trendradar.context import AppContext

# Load configuration (uses CONFIG_PATH env var or defaults)

cfg = load_config()

# Create shared context

ctx = AppContext(cfg)

The load_config() function in trendradar/core/loader.py loads and validates the entire configuration tree from YAML files and environment variables.

2. Scheduling Resolution


# Create scheduler and check if we should run

scheduler = ctx.create_scheduler()
should_run = scheduler.should_crawl()  # or should_analyze(), should_push()

The trendradar/core/scheduler.py module implements timeline-based scheduling with cron-like periods and "once" flags to prevent duplicate executions.

3. Data Crawling


# Crawl platforms and RSS feeds

results, id_to_name, failed = ctx.create_scheduler().crawl_all()

The crawler layer in trendradar/crawler/DataFetcher and trendradar/crawler/rss.py fetches raw news from web platforms and RSS feeds, normalizing everything into a common data structure.

4. Storage and Deduplication

Crawled items persist through trendradar/storage/manager.py, which provides:

New-title detection happens through ctx.detect_new_titles() and ctx.detect_latest_new_titles(), which identify titles appearing for the first time today.

5. Analysis and Filtering

Keyword mode:

stats, total = ctx.count_frequency(
    results,
    *ctx.load_frequency_words(),
    id_to_name,
    title_info,
    new_titles=new_titles,
    mode="daily"
)

The trendradar/core/analyzer.py module handles keyword-group matching, weight calculation, and platform-wise aggregation. Weights combine rank, frequency, and hotness scores.

AI mode:


# Enable AI filtering

cfg["FILTER"]["METHOD"] = "ai"
ctx = AppContext(cfg)

# Run AI filter

ai_result = ctx.run_ai_filter(interests_file="my_interests.txt")

# Convert to report format

hotlist_stats, rss_stats = ctx.convert_ai_filter_to_report_data(ai_result, mode="daily")

The trendradar/ai/filter.py pipeline classifies items using large language models, handles tag extraction and versioning, and returns a unified result set compatible with the reporting layer.

6. Report Generation


# Generate HTML report

html_path = ctx.generate_html(
    stats,
    total,
    new_titles=new_titles,
    id_to_name=id_to_name,
    mode="daily"
)

# Prepare language-agnostic report data

report_data = ctx.prepare_report(hotlist_stats, id_to_name=id_to_name, mode="daily")

The reporting layer in trendradar/report/__init__.py and trendradar/report/generator.py builds data structures for HTML and text rendering. Templates reside in docs/.

7. Notification Dispatch

dispatcher = ctx.create_notification_dispatcher()
dispatcher.dispatch_all(
    report_data=report_data,
    report_type="热点分析报告",
    update_info=None,
    proxy_url=None,
    mode="daily",
    html_file_path=None,
    rss_items=rss_stats,
    rss_new_items=None,
    ai_analysis=None,
    standalone_data=None,
    skip_translation=False,
)

The trendradar/notification/dispatcher.py sends rendered reports to 10+ channels: Feishu, DingTalk, WeWork, Telegram, Email, Ntfy, Bark, Slack, and generic webhooks. Key supporting modules:


Key Design Decisions

Flexible Scheduling with Timeline Semantics

The timeline.yaml configuration defines daily/weekly "periods" (e.g., "morning", "evening") with "once" flags to prevent duplicate executions. This is implemented in trendradar/core/scheduler.py through the ResolvedSchedule class.

Storage Abstraction

StorageManager in trendradar/storage/manager.py hides implementation details behind a uniform API:

Backend Implementation
Local SQLite sqlite_mixin.py
Plain text Direct file I/O
Remote S3 trendradar/storage/remote.py
Pull sync Sync protocol in remote.py

This allows swapping backends without touching business logic.

Modular Notification System

Each notification channel has its own renderer and batch size configuration. The dispatcher can skip translation for already-translated data, reducing API costs when multiple channels share the same content.


Summary

  • TrendRadar architecture is a layered, modular Python application with 11 distinct functional layers coordinated through AppContext
  • Context-driven dependency injection eliminates global state and enables comprehensive testing
  • Pluggable filtering supports both keyword-based and AI-driven analysis without pipeline changes
  • Timeline-based scheduling with "once" semantics prevents duplicate executions
  • Storage abstraction unifies SQLite, file, and remote S3 backends behind a common API
  • Multi-channel notifications support 10+ platforms with batching, translation, and channel-specific rendering

Frequently Asked Questions

What programming pattern does TrendRadar use for dependency management?

TrendRadar uses context-driven dependency injection. Every subsystem receives the same AppContext instance passed through constructors or method parameters. This pattern eliminates global state, makes components individually testable with mocked contexts, and centralizes access to configuration, storage, and utilities in trendradar/context.py.

How does TrendRadar switch between keyword and AI filtering?

TrendRadar toggles filtering methods through the FILTER.METHOD configuration key. Setting "keyword" activates the trendradar/core/analyzer.py pipeline for word frequency and weight calculation. Setting "ai" activates trendradar/ai/filter.py for LLM-based classification. Both pipelines return compatible data structures, so trendradar/report/__init__.py and trendradar/notification/dispatcher.py work unchanged regardless of filtering method.

Can TrendRadar store data on remote servers instead of locally?

Yes. The StorageManager class in trendradar/storage/manager.py abstracts storage backends behind a uniform API. While local SQLite is the default via trendradar/storage/sqlite_mixin.py, the system also supports remote S3 storage and pull-sync protocols through trendradar/storage/remote.py. You can swap backends by changing configuration without modifying business logic in crawlers, analyzers, or notifiers.

What notification channels does TrendRadar support?

TrendRadar supports 10 notification channels: Feishu, DingTalk, WeWork, Telegram, Email, Ntfy, Bark, Slack, and generic webhooks. The trendradar/notification/dispatcher.py coordinates delivery, while trendradar/notification/renderer.py formats channel-specific payloads. The trendradar/notification/splitter.py handles batching for large reports, and optional AITranslator integration reduces API costs when multiple channels share content.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →